Hiroshi G. Okuno

dblp:09/2222 · also Hiroshi Gitchang Okuno · DBLP profile ↗
← Back
241ranked-venue papers
17as first author
1since 2021 · last 2021
0000-0002-8704-4318ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 192 · 12 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 100 · 10 first-authorSystems, architecture and hardware · 82 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 9Applied, interdisciplinary, general and emerging computing · 8Databases, data management, data science and information retrieval · 1 · 1 first-authorTheory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
15 papers
Audio and music processing · 100%
Artificial intelligence
24 papers
Speech recognition and synthesis · 50% Robot manipulation · 14% Language models and text generation · 7%
Human-computer interaction and pervasive computing
8 papers
Human-robot interaction · 59% Interaction techniques and input · 38% User interface design and tools · 3%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%

Topics — the 30 heaviest of 63, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing
source separation
0.742018
Speech Enhancement Based on Bayesian Low-Rank and Sparse Decomposition of Multichannel Magnitude Spectrograms · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Multichannel sound source dereverberation and separation for arbitrary number of sources based on Bayesian nonparametrics · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Bayesian Nonparametrics for Microphone Array Processing · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Audio and music processing › speech enhancement
dereverberation
0.532014
Multichannel sound source dereverberation and separation for arbitrary number of sources based on Bayesian nonparametrics · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Nonparametric Bayesian dereverberation of power spectrograms based on infinite-order autoregressive processes · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Blind Separation and Dereverberation of Speech Mixtures by Joint Optimization · IEEE Trans. Speech Audio Process. 2011
Audio and music processing
speech enhancement
0.522018
Speech Enhancement Based on Bayesian Low-Rank and Sparse Decomposition of Multichannel Magnitude Spectrograms · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Blind Separation and Dereverberation of Speech Mixtures by Joint Optimization · IEEE Trans. Speech Audio Process. 2011
Audio and music processing › speech enhancement
multichannel speech enhancement
0.312018
Speech Enhancement Based on Bayesian Low-Rank and Sparse Decomposition of Multichannel Magnitude Spectrograms · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Audio and music processing › source separation
nonnegative tensor factorization
0.312018
Speech Enhancement Based on Bayesian Low-Rank and Sparse Decomposition of Multichannel Magnitude Spectrograms · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
robust speech recognition
0.332015
Automatic Speech Recognition for Mixed Dialect Utterances by Mixing Dialect Language Models · IEEE ACM Trans. Audio Speech Lang. Process. 2015
Robust Recognition of Simultaneous Speech by a Mobile Robot · IEEE Trans. Robotics 2007
Enhanced Robot Speech Recognition Based on Microphone Array Source Separation and Missing Feature Theory · ICRA 2005
Audio and music processing
music information retrieval
0.332010
A Modeling of Singing Voice Robust to Accompaniment Sounds and Its Application to Singer Identification and Vocal-Timbre-Similarity-Based Music Information Retrieval · IEEE Trans. Speech Audio Process. 2010
Design and Implementation of Two-level Synchronization for Interactive Music Robot · AAAI 2010
Drum Sound Recognition for Polyphonic Audio Signals by Adaptation and Matching of Spectrogram Templates With Harmonic Structure Suppression · IEEE Trans. Speech Audio Process. 2007
Audio and music processing › source separation
blind source separation
0.222011
Blind Separation and Dereverberation of Speech Mixtures by Joint Optimization · IEEE Trans. Speech Audio Process. 2011
Upper-limit evaluation of robot audition based on ICA-BSS in multi-source, barge-in and highly reverberant conditions · ICRA 2010
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
dialectal speech recognition
0.212015
Automatic Speech Recognition for Mixed Dialect Utterances by Mixing Dialect Language Models · IEEE ACM Trans. Audio Speech Lang. Process. 2015
Natural language and speech › Language models and text generation
language modeling
0.212015
Automatic Speech Recognition for Mixed Dialect Utterances by Mixing Dialect Language Models · IEEE ACM Trans. Audio Speech Lang. Process. 2015
Natural language and speech › Speech recognition and synthesis › speech separation › computational auditory scene analysis
robot audition
0.242008
A robot referee for rock-paper-scissors sound games · ICRA 2008
Improvement of Robot Audition by Interfacing Sound Source Separation and Automatic Speech Recognition with Missing Feature Theory · ICRA 2004
Robot recognizes three simultaneous speech by active audition · ICRA 2003
Audio and music processing › source separation
multichannel source separation
0.212014
Multichannel sound source dereverberation and separation for arbitrary number of sources based on Bayesian nonparametrics · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Audio and music processing
sound source localization
0.212014
Bayesian Nonparametrics for Microphone Array Processing · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Audio and music processing
computational auditory scene analysis
0.232012
Bayesian Unification of Sound Source Localization and Separation with Permutation Resolution · AAAI 2012
Residue-Driven Architecture for Computational Auditory Scene Analysis · IJCAI 1995
Auditory Stream Segregation in Auditory Scene Analysis with a Multi-Agent System · AAAI 1994
Interaction techniques and input
voice interaction
0.212013
Hands-free human-robot communication robust to speaker's radial position · ICRA 2013
Robotics › Robot manipulation
object dynamics prediction
0.222008
Object dynamics prediction and motion generation based on reliable predictability · ICRA 2008
Predicting Object Dynamics from Visual Images through Active Sensing Experiences · ICRA 2007
Computer vision › 3D vision
3d scene reconstruction
0.112012
Incremental probabilistic geometry estimation for robot scene understanding · ICRA 2012
Natural language and speech › Speech recognition and synthesis
sound source separation
0.122008
A robot referee for rock-paper-scissors sound games · ICRA 2008
Improvement of Robot Audition by Interfacing Sound Source Separation and Automatic Speech Recognition with Missing Feature Theory · ICRA 2004
Robotics › Robot navigation and mapping
sound source localization
0.122008
Two-channel-based voice activity detection for humanoid robots in noisy home environments · ICRA 2008
Robot recognizes three simultaneous speech by active audition · ICRA 2003
Natural language and speech › Speech recognition and synthesis
speech separation
0.132007
Robust Recognition of Simultaneous Speech by a Mobile Robot · IEEE Trans. Robotics 2007
Real-Time Speaker Localization and Speech Separation by Audio-Visual Integration · ICRA 2002
Understanding Three Simultaneous Speeches · IJCAI (1) 1997
Audio and music processing › source separation › blind source separation
independent component analysis
0.112010
Upper-limit evaluation of robot audition based on ICA-BSS in multi-source, barge-in and highly reverberant conditions · ICRA 2010
Audio and music processing › music information retrieval › music alignment
score following
0.112010
Design and Implementation of Two-level Synchronization for Interactive Music Robot · AAAI 2010
Audio and music processing › music information retrieval
singer identification
0.112010
A Modeling of Singing Voice Robust to Accompaniment Sounds and Its Application to Singer Identification and Vocal-Timbre-Similarity-Based Music Information Retrieval · IEEE Trans. Speech Audio Process. 2010
Audio and music processing › source separation
speech separation
0.112010
Upper-limit evaluation of robot audition based on ICA-BSS in multi-source, barge-in and highly reverberant conditions · ICRA 2010
Machine learning › Generative modeling
motion generation
0.132014
Insertion of pause in drawing from babbling for robot's developmental imitation learning · ICRA 2014
Object dynamics prediction and motion generation based on reliable predictability · ICRA 2008
Predicting Object Dynamics from Visual Images through Active Sensing Experiences · ICRA 2007
Natural language and speech › Speech recognition and synthesis
speech production
0.112009
Continuous vocal imitation with self-organized vowel spaces in Recurrent Neural Network · ICRA 2009
Human-robot interaction
learning from demonstration
0.112009
Prediction and imitation of other's motions by reusing own forward-inverse model in robots · ICRA 2009
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › robust speech recognition
missing feature theory
0.122007
Robust Recognition of Simultaneous Speech by a Mobile Robot · IEEE Trans. Robotics 2007
Improvement of Robot Audition by Interfacing Sound Source Separation and Automatic Speech Recognition with Missing Feature Theory · ICRA 2004
Natural language and speech › Speech recognition and synthesis › speech analysis
voice activity detection
0.112008
Two-channel-based voice activity detection for humanoid robots in noisy home environments · ICRA 2008
Recommender systems
collaborative filtering
0.112008
An Efficient Hybrid Music Recommender System Using an Incrementally Trainable Probabilistic Generative Model · IEEE Trans. Speech Audio Process. 2008

Methods — techniques the papers use, named apart from their topics

recurrent neural network with parametric bias · 0.6dirichlet process · 0.4bayesian nonparametrics · 0.4hierarchical neural network · 0.3radial distance compensation · 0.3nonnegative tensor factorization · 0.3low-rank and sparse decomposition · 0.3curve fitting · 0.3bayesian modeling · 0.3acoustic model likelihood · 0.3geometric source separation · 0.3machine translation · 0.2language model interpolation · 0.2variational bayes · 0.2nonparametric bayesian model · 0.2motionese · 0.2minorization-maximization · 0.2body babbling · 0.2
YearPublicationVenuePosition
2021 Alternating Drive-and-Glide Flight Navigation of a Kiteplane for Sound Source Position Estimation
abstract
Drone audition, namely the hearing capability of a drone, is expected to compensate for the drawbacks of visual sensors in search-and-rescue missions. Current multi-rotor drones have limitations of flight duration and sound processing due to ego-noise generated by rotors and air-flow. Drone audition for a kiteplane, i.e., a fixed-wing drone that can fly slowly and stably, has not been investigated. This paper proposes "Alternating Drive-and-Glide Flight Navigation"(AltDGFNavi) of a kiteplane for sound source position estimation. AltDGFNavi consists of two functions: periodical switching rotor for driving and gliding to reduce ego-noise, and dynamic flight path generation to fly close to the target. AltDGFNavi was evaluated through numerical simulations, and the results of sound source position estimation demonstrated the effectiveness of AltDGFNavi.
Makoto Kumon, Hiroshi G. Okuno, Shuichi Tajima
IROS2
2020 Computational Design of Balanced Open Link Planar Mechanisms with Counterweights from User Sketches
abstract
We consider the design of under-actuated articulated mechanism that are able to maintain stable static balance. Our method augments an user-provided design with counter-weights whose mass and attachment locations are automatically computed. The optimized counterweights adjust the center of gravity such that, for bounded external perturbations, the mechanism returns to its original configuration. Using our sketch-based system, we present several examples illustrating a wide range of user-provided designs can be successfully converted into statically-balanced mechanisms. We further validate our results with a set of physical prototypes.
Takuto Takahashi, Hiroshi G. Okuno, Shigeki Sugano, Stelian Coros, Bernhard Thomaszewski
IROS2
2019 An Integrated Framework for Field Recording, Localization, Classification and Annotation of Birdsongs Using Robot Audition Techniques - Harkbird 2.0
abstract
Bird vocalizations are one of the important subjects in ecoacoustics because birds communicate diversely using various vocalizations such as songs and calls. We have developed a portable system, HARKBird to provide a basic function, i.e., birdsong localization, which automatically extracts sound sources and their direction of arrivals (DOA) using robot audition techniques based on HARK. In this paper, we introduce HARKBird 2.0 which is empowered for higher understanding of birdsongs. A new soundscape annotation tool for localization results is enhanced by an interactive interface for song classification based on an unsupervised feature mapping t-SNE. We show that HARKBird 2.0 provides bird researchers with an integrated framework to analyze spatio-spectro-temporal dynamics of birdsongs using the song analysis of Japanese bush warbler (Horornis diphone).
Shinji Sumitani, Reiji Suzuki, Naoaki Chiba, Shiho Matsubayashi, Takaya Arita, Kazuhiro Nakadai, Hiroshi G. Okuno
ICASSP7
2018 Extracting the Relationship between the Spatial Distribution and Types of Bird Vocalizations Using Robot Audition System HARK
abstract
For a deeper understanding of ecological functions and semantics of wild bird vocalizations (i.e., songs and calls), it is important to clarify the fine-scaled and detailed relationships among their characteristics of vocalizations and their behavioral contexts. However, it takes a lot of time and effort to obtain such data using conventional recordings or by human observation. Bringing out a robot to a field is our approach to solve this problem. We are developing a portable observation system called HARKBird using a robot audition HARK and microphone arrays to understand temporal patterns of vocalizations characteristics and their behavioral contexts. In this paper, we introduce a prototype system to 2D localize vocalizations of wild birds in real-time, and to classify their song types after recording. We show that the system can estimate the position of songs of a target individual and classify their songs with a reasonable quality to discuss their song - behavior relationships.
Shinji Sumitani, Reiji Suzuki, Shiho Matsubayashi, Takaya Arita, Kazuhiro Nakadai, Hiroshi G. Okuno
IROS6
2018 Design and Implementation of Programmable Drawing Automata based on Cam Mechanisms for Representing Spatial Trajectory
abstract
This paper presents the design and implementation of a preliminary version of a programmable drawing automaton (PDA-0) that draws a user-specified 3D trajectory. PDA-0 is strongly inspired by Jaquet Droz's a programmable drawing automaton built in the 1770s using 6,000 moving parts, which was hand-coded. PDA-0 consists of RSSR (Revolute-Spherical-S-R) linkage and cam mechanisms with three interchangeable cams. Interchangeable cams make PDA-0 programmable because a user-specified 2D/3D trajectory is encoded into the set of three cams. The user programs PDA-0 by specifying a trajectory via a GUI or 3D animation. Subsequently, the compiler estimates a 3D trajectory mathematically from the user-specified 2D/3D trajectory and calculates the shape of the three cams, i.e., a code for PDA-0 by solving kinematic constraints. Finally, PDA-0 with the 3D-printed cams executes the code to draw the user-specified trajectory. The current PDA-0 with three cams demonstrates drawing simple trajectories such as letters and symbols.
Takuto Takahashi, Hiroshi G. Okuno
IROS2
2018 Speech Enhancement Based on Bayesian Low-Rank and Sparse Decomposition of Multichannel Magnitude Spectrograms
abstract
This paper presents a blind multichannel speech enhancement method that can deal with the time-varying layout of microphones and sound sources. Since nonnegative tensor factorization (NTF) separates a multichannel magnitude (or power) spectrogram into source spectrograms without phase information, it is robust against the time-varying mixing system. This method, however, requires prior information such as the spectral bases (templates) of each source spectrogram in advance. To solve this problem, we develop a Bayesian model called robust NTF (Bayesian RNTF) that decomposes a multichannel magnitude spectrogram into target speech and noise spectrograms based on their sparseness and low rankness. Bayesian RNTF is applied to the challenging task of speech enhancement for a microphone array distributed on a hose-shaped rescue robot. When the robot searches for victims under collapsed buildings, the layout of the microphones changes over time and some of them often fail to capture target speech. Our method robustly works under such situations, thanks to its characteristic of time-varying mixing system. Experiments using a 3-m hose-shaped rescue robot with eight microphones show that the proposed method outperforms conventional blind methods in enhancement performance by the signal-to-noise ratio of 1.03 dB.
Yoshiaki Bando, Katsutoshi Itoyama, Masashi Konyo, Satoshi Tadokoro, Kazuhiro Nakadai, Kazuyoshi Yoshii, Tatsuya Kawahara, Hiroshi G. Okuno
IEEE ACM Trans. Audio Speech Lang. Process.8
2017 Development of microphone-array-embedded UAV for search and rescue task
abstract
This paper addresses online outdoor sound source localization using a microphone array embedded in an unmanned aerial vehicle (UAV). In addition to sound source localization, sound source enhancement and robust communication method are also described. This system is one instance of deployment of our continuously developing open source software for robot audition called HARK (Honda Research Institute Japan Audition for Robots with Kyoto University). To improve the robustness against outdoor acoustic noise, we propose to combine two sound source localization methods based on MUSIC (multiple signal classification) to cope with trade-off between latency and noise robustness. The standard Eigenvalue decomposition based MUSIC (SEVD-MUSIC) has smaller latency but less noise robustness, whereas the incremental generalized singular value decomposition based MUSIC (iGSVD-MUSIC) has higher noise robustness but larger latency. A UAV operator can use an appropriate method according to the situation. A sound enhancement method called online robust principal component analysis (ORPCA) enables the operator to detect a target sound source more easily. To improve the stability of wireless communication, and robustness of the UAV system against weather changes, we developed data compression based on free lossless audio codec (FLAC) extended to support a 16 ch audio data stream via UDP, and developed a water-resistant microphone array. The resulting system successfully worked in an outdoor search and rescue task in ImPACT Tough Robotics Challenge in November 2016.
Kazuhiro Nakadai, Makoto Kumon, Hiroshi G. Okuno, Kotaro Hoshiba, Mizuho Wakabayashi, Kai Washizaki, Takahiro Ishiki, Daniel Gabriel, Yoshiaki Bando, Takayuki Morito, Ryosuke Kojima, Osamu Sugiyama
IROS3
2016 Call Alternation Between Specific Pairs of Male Frogs Revealed by a Sound-Imaging Method in Their Natural Habitat
Ikkyu Aihara, Takeshi Mizumoto, Hiromitsu Awano, Hiroshi G. Okuno
INTERSPEECH4
2016 Localizing Bird Songs Using an Open Source Robot Audition System with a Microphone Array
Reiji Suzuki, Shiho Matsubayashi, Kazuhiro Nakadai, Hiroshi G. Okuno
INTERSPEECH4
2016 Parallel Speech Corpora of Japanese Dialects
Koichiro Yoshino, Naoki Hirayama, Shinsuke Mori, Fumihiko Takahashi, Katsutoshi Itoyama, Hiroshi G. Okuno
LREC6
2015 Challenges in deploying a microphone array to localize and separate sound sources in real auditory scenes
abstract
Analyzing the auditory scene of real environments is challenging partly because an unknown number and type of sound sources are observed at the same time and partly because these sounds are observed on a significantly different sound pressure level at the microphone. These are difficult problems even with state-of-the-art sound source localization and separation methods. In this paper, we exploit two such methods using a microphone array: (1) Bayesian nonparametric microphone array processing (BNP-MAP), which is capable of separating and localizing sound sources when the number of sound sources is unspecified, and (2) robot audition software “HARK” is capable of separating and localizing in real time. Through experimentation, we found that BNP-MAP is more robust against differences in the sound pressure levels of the source signals and in the spatial closeness of source positions. Experiments analyzing real scenes of human conversations recorded in a big exhibition hall and bird calling recorded at a natural park demonstrate the efficacy and applicability of BNP-MAP.
Yoshiaki Bando, Takuma Otsuka, Katsutoshi Itoyama, Kazuyoshi Yoshii, Yoko Sasaki, Satoshi Kagami, Hiroshi G. Okuno
ICASSP7
2015 Robot audition: Its rise and perspectives
abstract
The ability of robots to listen to several things at once with their own “ears”, that is, robot audition, is an important factor in improving interaction and symbiosis between humans and robots. The critical issue in robot audition is real-time processing and robustness against noisy environments with high flexibility to support various kinds of robots and hardware configurations. This paper first overviews activities and issues related to robot audition. Then, it presents the “HARK” robot audition software, which provides three primary functions for robot audition, sound source localization, sound source separation, and separated sound recognition, and then reports their performance. Finally, it discusses future directions in new promising areas as well as robotics.
Hiroshi G. Okuno, Kazuhiro Nakadai
ICASSP1
2015 Microphone-accelerometer based 3D posture estimation for a hose-shaped rescue robot
abstract
3D posture estimation for a hose-shaped robot is critical in rescue activities due to complex physical environments. Conventional sound-based posture estimation assumes rather flat physical environments and focuses only on 2D, resulting in poor performance in real world environments with rubble. This paper presents novel 3D posture estimation by exploiting microphones and accelerometers. The idea of our method is to compensate the lack of posture information obtained by sound-based time-difference-of arrival (TDOA) with the tilt information obtained from accelerometers. This compensation is formulated as a nonlinear state-space model and solved by the unscented Kalman filter. Experiments are conducted by using a 3m hose-shaped robot with eight units of a microphone and an accelerometer and seven units of a loudspeaker and a vibration motor deployed in a simple 3D structure. Experimental results demonstrate that our method reduces the errors of initial states to about 20 cm in the 3D space. If the initial errors of initial states are less than 20 %, our method can estimate the correct 3D posture in real-time.
Yoshiaki Bando, Katsutoshi Itoyama, Masashi Konyo, Satoshi Tadokoro, Kazuhiro Nakadai, Kazuyoshi Yoshii, Hiroshi G. Okuno
IROS7
2015 Improved sound source localization in horizontal plane for binaural robot audition
Ui-Hyun Kim, Kazuhiro Nakadai, Hiroshi G. Okuno
Appl. Intell.3
2015 Audio-visual speech recognition using deep learning
abstract
Audio-visual speech recognition (AVSR) system is thought to be one of the most promising solutions for reliable speech recognition, particularly when the audio is corrupted by noise. However, cautious selection of sensory features is crucial for attaining high recognition performance. In the machine-learning community, deep learning approaches have recently attracted increasing attention because deep neural networks can effectively extract robust latent features that enable various recognition algorithms to demonstrate revolutionary generalization capabilities under diverse application conditions. This study introduces a connectionist-hidden Markov model (HMM) system for noise-robust AVSR. First, a deep denoising autoencoder is utilized for acquiring noise-robust audio features. By preparing the training data for the network with pairs of consecutive multiple steps of deteriorated audio features and the corresponding clean features, the network is trained to output denoised audio features from the corresponding features deteriorated by noise. Second, a convolutional neural network (CNN) is utilized to extract visual features from raw mouth area images. By preparing the training data for the CNN as pairs of raw images and the corresponding phoneme label outputs, the network is trained to predict phoneme labels from the corresponding mouth area input images. Finally, a multi-stream HMM (MSHMM) is applied for integrating the acquired audio and visual HMMs independently trained with the respective features. By comparing the cases when normal and denoised mel-frequency cepstral coefficients (MFCCs) are utilized as audio features to the HMM, our unimodal isolated word recognition results demonstrate that approximately 65 % word recognition rate gain is attained with denoised MFCCs under 10 dB signal-to-noise-ratio (SNR) for the audio signal input. Moreover, our multimodal isolated word recognition results utilizing MSHMM with denoised MFCCs and acquired visual features demonstrate that an additional word recognition rate gain is attained for the SNR conditions below 10 dB.
Kuniaki Noda, Yuki Yamaguchi, Kazuhiro Nakadai, Hiroshi G. Okuno, Tetsuya Ogata
Appl. Intell.4
2015 Automatic Speech Recognition for Mixed Dialect Utterances by Mixing Dialect Language Models
abstract
This paper presents an automatic speech recognition (ASR) system that accepts a mixture of various kinds of dialects. The system recognizes dialect utterances on the basis of the statistical simulation of vocabulary transformation and combinations of several dialect models. Previous dialect ASR systems were based on handcrafted dictionaries for several dialects, which involved costly processes. The proposed system statistically trains transformation rules between a common language and dialects, and simulates a dialect corpus for ASR on the basis of a machine translation technique. The rules are trained with small sets of parallel corpora to make up for the lack of linguistic resources on dialects. The proposed system also accepts mixed dialect utterances that contain a variety of vocabularies. In fact, spoken language is not a single dialect but a mixed dialect that is affected by the circumstances of speakers’ backgrounds (e.g., native dialects of their parents or where they live). We addressed two methods to combine several dialects appropriately for each speaker. The first was recognition with language models of mixed dialects with automatically estimated weights that maximized the recognition likelihood. This method performed the best, but calculation was very expensive because it conducted grid searches of combinations of dialect mixing proportions that maximized the recognition likelihood. The second was integration of results of recognition from each single dialect language model. The improvements with this model were slightly smaller than those with the first method. Its calculation cost was, however, inexpensive and it worked in real-time on general workstations. Both methods achieved higher recognition accuracies for all speakers than those with the single dialect models and the common language model, and we could choose a suitable model for use in ASR that took into consideration the computational costs and recognition accuracies.
Naoki Hirayama, Koichiro Yoshino, Katsutoshi Itoyama, Shinsuke Mori, Hiroshi G. Okuno
IEEE ACM Trans. Audio Speech Lang. Process.5
2014 Transcribing vocal expression from polyphonic music
abstract
A method for transcribing vocal expressions such as vibrato, glissando, and kobushi separately from polyphonic music is described. The expressions appear as fluctuation in the fundamental frequency contour of the singing voice. They can be used for search and retrieval of music and for expressive singing voice synthesis based on singing style since they strongly reflect the individuality of the singer. The fundamental frequency contour of the singing voice is estimated using the Viterbi algorithm with limitation from a corresponding note sequence. Next, the notes are aligned with the fundamental frequency sequence temporally. Finally, each expression is identified and parameterized in accordance with designed rules. Experiments demonstrated that this method can transcribe expressions in the singing voice from commercial recordings.
Yukara Ikemiya, Katsutoshi Itoyama, Hiroshi G. Okuno
ICASSP3
2014 Audio part mixture alignment based on hierarchical nonparametric Bayesian model of musical audio sequence collection
abstract
This paper proposes “audio part mixture alignment,” a method for temporally aligning multiple audio signals, each of which is a rendition of a non-disjoint subset of a common piece of music. The method decomposes each audio signal into shared components and components unique to each rendition. At the same time, it aligns each audio signal based on the shared component. Decomposition of audio signal is modeled using a hierarchical Dirichlet process (Hierarchical DP, HDP), and sequence alignment is modeled as a left-to-right hidden Markov model (HMM). Variational Bayesian inference is used to jointly infer the alignment and component decomposition. The proposed method is compared with a classic audio-to-audio alignment method, and it is found that the proposed method is more robust to the discrepancy of parts between two audio signals.
Akira Maezawa, Hiroshi G. Okuno
ICASSP2
2014 Automatic transcription of guitar tablature from audio signals in accordance with player's proficiency
abstract
We describe a method for automatically transcribing guitar tablatures from audio signals in accordance with the player's proficiency for use as support for a guitar player's practice. The system estimates the multiple pitches in each time frame and the optimal fingering considering playability and player's proficiency. It combines a conventional multipitch estimation method with a basic dynamic programming method. The difficulty of the fingerings can be changed by tuning the parameter representing the relative weights of the acoustical reproducibility and the fingering easiness. Experiments conducted using synthesized guitar audio signals to evaluate the transcribed tablatures in terms of the multipitch estimation accuracy and fingering easiness demonstrated that the system can simplify the fingering with higher precision of multipitch estimation results than the conventional method.
Kazuki Yazawa, Katsutoshi Itoyama, Hiroshi G. Okuno
ICASSP3
2014 Insertion of pause in drawing from babbling for robot's developmental imitation learning
abstract
In this paper, we present a method to improve a robot's imitation performance in a drawing scenario by inserting pauses in motion. Human's drawing skills are said to develop through five stages: 1) Scribbling, 2) Fortuitous Realism, 3) Failed Realism, 4) Intellectual Realism, and 5) Visual Realism. We focus on stages 1) and 3) for creating our system, each corresponding to body babbling and imitation learning, respectively. For stage 1), the robot randomly moves its arm to associate robot's arm dynamics with the drawing result. Presuming that the robot has no knowledge about its own dynamics, the robot learns its body dynamics in this stage. For stage 3), we consider a scenario where a robot would imitate a human's drawing motion. Upon creating the system, we focus on the motionese phenomenon, which is one of the key factors for discussing acquisition of a skill through a human parent-child interaction. In motionese, the parent would first show each action elaborately to the child, when teaching a skill. As the child starts to improve, the parent's actions would be simplified. Likewise in our scenario, the human would first insert pauses during the drawing motions where the direction of drawing changes (i.e. corners). As the robot's imitation learning of drawing converges, the human would change to drawing without pauses. The experimental results show that insertion of pause in drawing imitation scenarios greatly improves the robot's drawing performance.
Shun Nishide, Keita Mochizuki, Hiroshi G. Okuno, Tetsuya Ogata
ICRA3
2014 Transferring Vocal Expression of F0 Contour Using Singing Voice Synthesizer
Yukara Ikemiya, Katsutoshi Itoyama, Hiroshi G. Okuno
IEA/AIE (2)3
2014 Lipreading using convolutional neural network
Kuniaki Noda, Yuki Yamaguchi, Kazuhiro Nakadai, Hiroshi G. Okuno, Tetsuya Ogata
INTERSPEECH4
2014 Visualization of auditory awareness based on sound source positions estimated by depth sensor and microphone array
abstract
We have developed a system for visualizing auditory awareness on the basis of sound source locations estimated using a depth sensor and microphone array. Previous studies on visualizing the acoustic environment viewed the level of sound pressures directly on the captured image, so the visualization was often based on a mixture of several sound sources. As a result, which targets to focus on was not intuitive. To help users selectively to find the targets and focus on the target analysis, we should extract the captured acoustic information and selectively propose it with the user demand. We have designed a three-layer visualization model for auditory awareness consisting of a sound source distribution layer, a sound location layer, and a sound saliency layer. The model extracts acoustic information by using the depth image and multi-directional sound sources captured with a depth sensor and microphone array. This model is used in the system we developed for visualizing auditory awareness.
Takahiro Iyama, Osamu Sugiyama, Takuma Otsuka, Katsutoshi Itoyama, Hiroshi G. Okuno
IROS5
2014 Making a robot dance to diverse musical genre in noisy environments
abstract
In this paper we address the problem of musical genre recognition for a dancing robot with embedded microphones capable of distinguishing the genre of a musical piece while moving in a real-world scenario. For this purpose, we assess and compare two state-of-the-art musical genre recognition systems, based on Support Vector Machines and Markov Models, in the context of different real-world acoustic environments. In addition, we compare different preprocessing robot audition variants (single channel and separated signal from multiple channels) and test different acoustic models, learned a priori, to tackle multiple noise conditions of increasing complexity in the presence of noises of different natures (e.g., robot motion, speech). The results with six different musical genres suggest improved results, in the order of 43.6pp for the most complex conditions, when recurring to Sound Source Separation and acoustic models trained in similar conditions to the testing scenarios. A robot dance demonstration session confirms the applicability of the proposed integration for genre-adaptive dancing robots in real-world noisy environments.
João Lobato Oliveira, Keisuke Nakamura, Thibault Langlois, Fabien Gouyon, Kazuhiro Nakadai, Angelica Lim, Luís Paulo Reis, Hiroshi G. Okuno
IROS8
2014 Sound annotation tool for multidirectional sounds based on spatial information extracted by HARK robot audition software
abstract
With the rise of inexpensive microphone array products and the robot audition software called HARK, we can record and analyze multidirectional sound sources easily. The combination of microphone array and the software enables us to separate, localize, and track multidirectional sound sources. Most of the solutions for accessing these separated sound source information provide clients for interpreting simplified information about the separated sources, but not to directly execute the semantic annotations. Since the multidirectional sound annotation requires simultaneous labeling of separated sound sources and a multidirectional overview of the sources, it is essential to have an efficient way of annotation and an intuitive view of multidirectional sounds. Our proposed sound annotation tool provides drag & drop operation of annotation with a 3D sound source view and also provides annotation autocompletion with a SVM trained with the user's annotation history. The proposed features enable users to do the annotation task intuitively and confirm its result. We also conducted an evaluation demonstrating the efficiency of annotation done using the tool.
Osamu Sugiyama, Katsutoshi Itoyama, Kazuhiro Nakadai, Hiroshi G. Okuno
SMC4
2014 Nonparametric Bayesian dereverberation of power spectrograms based on infinite-order autoregressive processes
abstract
This paper describes a monaural audio dereverberation method that operates in the power spectrogram domain. The method is robust to different kinds of source signals such as speech or music. Moreover, it requires little manual intervention, including the complexity of room acoustics. The method is based on a non-conjugate Bayesian model of the power spectrogram. It extends the idea of multi-channel linear prediction to the power spectrogram domain, and formulates a model of reverberation as a non-negative, infinite-order autoregressive process. To this end, the power spectrogram is interpreted as a histogram count data, which allows a nonparametric Bayesian model to be used as the prior for the autoregressive process, allowing the effective number of active components to grow, without bound, with the complexity of data. In order to determine the marginal posterior distribution, a convergent algorithm, inspired by the variational Bayes method, is formulated. It employs the minorization-maximization technique to arrive at an iterative, convergent algorithm that approximates the marginal posterior distribution. Both objective and subjective evaluations show advantage over other methods based on the power spectrum. We also apply the method to a music information retrieval task and demonstrate its effectiveness.
Akira Maezawa, Katsutoshi Itoyama, Kazuyoshi Yoshii, Hiroshi G. Okuno
IEEE ACM Trans. Audio Speech Lang. Process.4
2014 Bayesian Nonparametrics for Microphone Array Processing
abstract
Sound source localization and separation from a mixture of sounds are essential functions for computational auditory scene analysis. The main challenges are designing a unified framework for joint optimization and estimating the sound sources under auditory uncertainties such as reverberation or unknown number of sounds. Since sound source localization and separation are mutually dependent, their simultaneous estimation is required for better and more robust performance. A unified model is presented for sound source localization and separation based on Bayesian nonparametrics. Experiments using simulated and recorded audio mixtures show that a method based on this model achieves state-of-the-art sound source separation quality and has more robust performance on the source number estimation under reverberant environments.
Takuma Otsuka, Katsuhiko Ishiguro, Hiroshi Sawada, Hiroshi G. Okuno
IEEE ACM Trans. Audio Speech Lang. Process.4
2014 Multichannel sound source dereverberation and separation for arbitrary number of sources based on Bayesian nonparametrics
abstract
Multichannel signal processing using a microphone array provides fundamental functions for coping with multi-source situations, such as sound source localization and separation, that are needed to extract the auditory information for each source. Auditory uncertainties about the degree of reverberation and the number of sources are known to degrade performance or limit the practical application of microphone array processing. Such uncertainties must therefore be overcome to realize general and robust microphone array processing. These uncertainty issues have been partly addressed-existing methods focus on either source number uncertainty or the reverberation issue, where joint separation and dereverberation has been achieved only for the overdetermined conditions. This paper presents an all-round method that achieves source separation and dereverberation for an arbitrary number of sources including underdetermined conditions. Our method uses Bayesian nonparametrics that realize an infinitely extensible modeling flexibility so as to bypass the model selection in the separation and dereverberation problem, which is caused by the source number uncertainty. Evaluation using a dereverberation and separation task with various numbers of sources including underdetermined conditions demonstrates that (1) our method is applicable to the separation and dereverberation of underdetermined mixtures, and that (2) the source extraction performance is comparable to that of a state-of-the-art method suitable only for overdetermined conditions.
Takuma Otsuka, Katsuhiko Ishiguro, Takuya Yoshioka, Hiroshi Sawada, Hiroshi G. Okuno
IEEE ACM Trans. Audio Speech Lang. Process.5
2013 Multiple index combination for Japanese spoken term detection with optimum index selection based on OOV-region classifier
Naoyuki Kanda, Katsutoshi Itoyama, Hiroshi G. Okuno
ICASSP3
2013 Initialization-robust Bayesian multipitch analyzer based on psychoacoustical and musical criteria
abstract
We present a new Bayesian multipitch analyzer that dispenses with a precise optimization of parameter initialization or hyperparameters. Our method uses a new family of prior distribution, characteristic prior; it efficiently restricts the existence region of the latent variables, that is, the product of a conjugate prior and a characteristic function. The update formulas become a simple form that is actually suitable for Gibbs sampling. We construct characteristic priors of harmonic structures based on psychoacoustical and musical knowledge and apply them to nonnegative harmonic factorization. Experimental results improve 5.2 points in F-measure under a tough condition, random initialization with no hyperparameter optimization.
Daichi Sakaue, Takuma Otsuka, Katsutoshi Itoyama, Hiroshi G. Okuno
ICASSP4
2013 Audio-based guitar tablature transcription using multipitch analysis and playability constraints
abstract
This paper proposes a method of guitar tablature transcription from audio signals. Multipitch estimation and fingering configuration estimation are essential for transcribing tablatures. Conventional multipitch estimation methods, including latent harmonic allocation (LHA), often estimate combinations of pitches that people cannot play due to inherent physical constraints. Unplayable combinations of pitches are eliminated by filtering the results of LHA with three constraints. We first enumerate playable fingering configurations, and use them to suppress any undesirable combination of pitches. The optimal fingering configuration in each time frame is optimized to satisfy the need for temporal continuity by using dynamic programming. We use synthesized guitar sounds from MIDI data (ground truth) for evaluation. Experiments with them demonstrate the improvement of multipitch estimation by 5.9 points on average in F-measure and the transcribed tablatures are playable.
Kazuki Yazawa, Daichi Sakaue, Kohei Nagira, Katsutoshi Itoyama, Hiroshi G. Okuno
ICASSP5
2013 Hands-free human-robot communication robust to speaker's radial position
abstract
In this paper we present a method in room transfer function (RTF) estimation, employed specifically for dereverberation in hands-free human-robot communication.We introduce a radial distance compensation scheme which significantly improved the RTF estimate robust to the speech power variation due to changes in speaker's radial position. The proposed method is implemented in two levels; first, waveform-level compensation is executed to reflect the change in power caused by the change of radial position to the RTF. We generated possible RTF estimates within a close neighbourhood based on curve fitting. Then, we select among these estimates the optimal RTF based on acoustic model likelihood criterion, the same criterion employed in automatic speech recognition (ASR) systems. The latter is referred to as acoustic model-level compensation, which links the generated RTF to the ASR. We note that in ASR application, both waveform and acoustic models play an important role in achieving optimal performance. Thus, the synergistic effect of the two processes guarantee ASR performance improvement when used in conjunction with our ASR-based dereverberation scheme. Experimental evaluation show robustness in recognition performance when used in hands-free human-robot communication environment.
Randy Gomez, Keisuke Nakamura, Kazuhiro Nakadai, Ui-Hyun Kim, Hiroshi G. Okuno, Tatsuya Kawahara
ICRA5
2013 Improved Sound Source Localization and Front-Back Disambiguation for Humanoid Robots with Two Ears
Ui-Hyun Kim, Kazuhiro Nakadai, Hiroshi G. Okuno
IEA/AIE3
2013 Automatic estimation of dialect mixing ratio for dialect speech recognition
abstract
This paper proposes methods for determining an appropriate mixing ratio of dialects in automatic speech recognition (ASR) for dialects. To handle ASR for various dialects, it has been reported to be effective to train a language model using a dialectmixed corpus. One reason behind this is geographical continuity of spoken dialect; we regard spoken dialect as a mixture of various dialects. This mixing ratio changes at every moment as well as depends on a speaker. We can improve recognition accuracybygivingan appropriatedialectmixingratio foraspeaker’s dialect. The mixing ratio is generally unknown and requires to be estimated and updated referring to input utterances. We handle two methods for updating it based on recognition results; one is to compute contribution of dialects for each recognized word, and the other is to predict mixture information referring to a whole recognized sentence based on topic modeling. The experimental result shows that the mixing ratio estimated by these methods realized higher recognition accuracy than a fixed mixing ratio. Index Terms: dialect, supervised latent Dirichlet allocation (sLDA), mixing ratio.
Naoki Hirayama, Koichiro Yoshino, Katsutoshi Itoyama, Shinsuke Mori, Hiroshi G. Okuno
INTERSPEECH5
2013 Posture estimation of hose-shaped robot using microphone array localization
abstract
This paper presents a posture estimation of hose-shaped robot using microphone array localization. The hose-shaped robots, one of major rescue robots, have problems with navigation because their posture is too flexible for a remote operator to control to go as far as desired. For navigational and mission usability, the posture estimation of the hose-shaped robot is essential. We developed a posture estimation method with a microphone array and small loudspeakers equipped on the hose-shaped robot. Our method consists of two steps: (1) playing a known sound from the loudspeaker one-by-one, and (2) estimating the microphone positions on the hose-shaped robot instead of estimating the posture directly. We designed a time difference of arrival (TDOA) estimation method to be robust against directional noise and implemented a prototype system using a posture model of the hose-shaped robot and an Extended Kalman Filter (EKF). The validity of our approach is evaluated by the experiments with both signals recorded in an anechoic chamber and simulated data.
Yoshiaki Bando, Takeshi Mizumoto, Katsutoshi Itoyama, Kazuhiro Nakadai, Hiroshi G. Okuno
IROS5
2013 Noise correlation matrix estimation for improving sound source localization by multirotor UAV
abstract
A method has been developed for improving sound source localization (SSL) using a microphone array from an unmanned aerial vehicle with multiple rotors, a “multirotor UAV”. One of the main problems in SSL from a multirotor UAV is that the ego noise of the rotors on the UAV interferes with the audio observation and degrades the SSL performance. We employ a generalized eigenvalue decomposition-based multiple signal classification (GEVD-MUSIC) algorithm to reduce the effect of ego noise. While GEVD-MUSIC algorithm requires a noise correlation matrix corresponding to the auto-correlation of the multichannel observation of the rotor noise, the noise correlation is nonstationary due to the aerodynamic control of the UAV. Therefore, we need an adaptive estimation method of the noise correlation matrix for a robust SSL using GEVD-MUSIC algorithm. Our method uses a Gaussian process regression to estimate the noise correlation matrix in each time period from the measurements of self-monitoring sensors attached to the UAV such as the pitch-roll-yaw tilt angles, xyz speeds, and motor control values. Experiments compare our method with existing SSL methods in terms of precision and recall rates of SSL. The results demonstrate that our method outperforms existing methods, especially under high signal-to-noise-ratio conditions.
Koutarou Furukawa, Keita Okutani, Kohei Nagira, Takuma Otsuka, Katsutoshi Itoyama, Kazuhiro Nakadai, Hiroshi G. Okuno
IROS7
2013 Developmental Human-Robot Imitation Learning of Drawing with a Neuro Dynamical System
abstract
This paper mainly deals with robot developmental learning on drawing and discusses the influences of physical embodiment to the task. Humans are said to develop their drawing skills through five phases: 1) Scribbling, 2) Fortuitous Realism, 3) Failed Realism, 4) Intellectual Realism, 5) Visual Realism. We implement phases 1) and 3) into the humanoid robot NAO, holding a pen, using a neuro dynamical model, namely Multiple Timescales Recurrent Neural Network (MTRNN). For phase 1), we used random arm motion of the robot as body babbling to associate motor dynamics with pen position dynamics. For phase 3), we developed incremental imitation learning to imitate and develop the robot's drawing skill using basic shapes: circle, triangle, and rectangle. We confirmed two notable features from the experiment. First, the drawing was better performed for shapes requiring arm motions used in babbling. Second, performance of clockwise drawing of circle was good from beginning, which is a similar phenomenon that can be observed in human development. The results imply the capability of the model to create a developmental robot relating to human development.
Keita Mochizuki, Shun Nishide, Hiroshi G. Okuno, Tetsuya Ogata
SMC3
2012 Bayesian Unification of Sound Source Localization and Separation with Permutation Resolution
abstract
Sound source localization and separation with permutation resolution are essential for achieving a computational auditory scene analysis system that can extract useful information from a mixture of various sounds. Because existing methods cope separately with these problems despite their mutual dependence, the overall result with these approaches can be degraded by any failure in one of these components. This paper presents a unified Bayesian framework to solve these problems simultaneously where localization and separation are regarded as a clustering problem. Experimental results confirm that our method outperforms state-of-the-art methods in terms of the separation quality with various setups including practical reverberant environments.
Takuma Otsuka, Katsuhiko Ishiguro, Hiroshi Sawada, Hiroshi G. Okuno
AAAI4
2012 Statistical Method of Building Dialect Language Models for ASR Systems
Naoki Hirayama, Shinsuke Mori, Hiroshi G. Okuno
COLING3
2012 Initialization-robust multipitch estimation based on latent harmonic allocation using overtone corpus
abstract
We present a new method for modeling the overtone structures of musical instruments that uses an overtone corpus generated using a MIDI synthesizer. Since multipitch estimation requires a joint estimation of F0's and their overtone structures, one of the most important problems is the overtone structure modeling. Latent harmonic allocation (LHA), a promising multipitch estimation method, is difficult to use for various applications because it requires appropriate prior distributions of the overtone structures, which cannot be determined from statistical evidence. Our method uses an overtone corpus to avoid the problem of setting prior distributions and instead restricts the lower and upper bounds of each overtone weight. The bounds are determined from reference signals generated by a MIDI synthesizer. Experimental results demonstrated that the overtone structures were stably and accurately estimated for a wide variety of initial settings.
Daichi Sakaue, Katsutoshi Itoyama, Tetsuya Ogata, Hiroshi G. Okuno
ICASSP4
2012 Incremental probabilistic geometry estimation for robot scene understanding
abstract
Our goal is to give mobile robots a rich representation of their environment as fast as possible. Current mapping methods such as SLAM are often sparse, and scene reconstruction methods using tilting laser scanners are relatively slow. In this paper, we outline a new method for iterative construction of a geometric mesh using streaming time-of-flight range data. Our results show that our algorithm can produce a stable representation after 6 frames, with higher accuracy than raw time-of-flight data.
Louis-Kenzo Cahier, Tetsuya Ogata, Hiroshi G. Okuno
ICRA3
2012 Automatic Chord Recognition Based on Probabilistic Integration of Acoustic Features, Bass Sounds, and Chord Transition
Katsutoshi Itoyama, Tetsuya Ogata, Hiroshi G. Okuno
IEA/AIE3
2012 Self-organization of object features representing motion using Multiple Timescales Recurrent Neural Network
abstract
Affordance theory suggests that humans recognize the environment based on invariants. Invariants are features that describe the environment offering behavioral information to humans. Two types of invariants exist, structural invariants and transformational invariants. In our previous paper, we developed a method that self-organizes transformational invariants, or motion features, from camera images based on robot's experiences. The model used a bi-directional technique combining a recurrent neural network for dynamics learning, namely Recurrent Neural Network with Parametric Bias (RNNPB), and a hierarchical neural network for feature extraction. The bi-directional training method developed in the previous work was effective in clustering the motion of objects, but the analysis did not give good segregation results of the self-organized features (transformational invariants) among different motion types. In this paper, we present a refined model which integrates dynamics learning and feature extraction in a single model. The refined model is comprised of Multiple Timescales Recurrent Neural Network (MTRNN), which possesses better learning capability than RNNPB. Self-organization result of four types of motions have proved the model's capability to create clusters of object motions. The analysis showed that the model extracted feature sequences with different characteristics for four object motion types.
Shun Nishide, Jun Tani, Hiroshi G. Okuno, Tetsuya Ogata
IJCNN3
2012 Body area segmentation from visual scene based on predictability of neuro-dynamical system
abstract
We propose neural models for segmenting the area of a body from visual scene based on predictability. Neuroscience has shown that a prediction model in brain, which predicts sensory-feedback from motor command, can divide the sensory-feedback into the self-motion derived feedback and other derived feedback. The prediction model is important for prediction control of the body. Previous studies in robotics of the prediction model assumed that a robot can recognize the position of its body (e.g. its hand) and that the view contains only that body part. In our models, motor commands and visual feedback (pixel image that includes not only a hand but also object and background) are input into a neural network model and then the body area is segmented and prediction model of body is acquired. Our model contains two parts: 1) An object detection model obtains a conversion system between object positions and the pixel image. 2) A movement prediction model predicts hand-object positions from motor commands and identifies the body. We confirmed that our models can segment the body/object area based on their pixel textures and discriminate between them by using prediction error.
Harumitsu Nobuta, Kenta Kawamoto, Kuniaki Noda, Kohtaro Sabe, Shun Nishide, Hiroshi G. Okuno, Tetsuya Ogata
IJCNN6
2012 Who is the leader in a multiperson ensemble? - Multiperson human-robot ensemble model with leaderness -
abstract
This paper presents a state space model for a multiperson ensemble and an estimation method of the onset timings, tempos, and leaders. In a multiperson ensemble, determining one explicit leader is difficult because (1) participants' rhythms are mutually influenced and (2) they compete with each other. Most ensemble studies however assumed that one leader exists at a time and the others just follow the leader. To deal with the multiple and time-varying leaders, we define leaderness indicating the power to influence the others as the product of the tempo stability and the distance from the ensemble tempo. This definition means that a leader should have a strong desire to change the current tempo. Using the leaderness, we present a state space model of a multiperson ensemble and an unscented Kalman filter based estimation method. The model consists of the leaderness update, the ensemble tempo update, the individual tempo update, and the onset timing adaptation, each of which has a relationship to psychological results of an ensemble. We evaluate our method using simulation and human behavior. The simulation results show that our model is stable for various initial tempos and the number of participants. For the human behavior, pairs and triads of participants are asked to tap keys in synchronization with the others. The results show that the leaderness successfully indicate the dynamics of the leaders, and the onset errors are 181msec and 241msec for pairs and triads on average, respectively, which are comparable to those of humans (153msec and 227msec for pairs and triads, respectively.)
Takeshi Mizumoto, Tetsuya Ogata, Hiroshi G. Okuno
IROS3
2012 Live assessment of beat tracking for robot audition
abstract
In this paper we propose the integration of an online audio beat tracking system into the general framework of robot audition, to enable its application in musically-interactive robotic scenarios. To this purpose, we introduced a staterecovery mechanism into our beat tracking algorithm, for handling continuous musical stimuli, and applied different multi-channel preprocessing algorithms (e.g., beamforming, ego noise suppression) to enhance noisy auditory signals lively captured in a real environment. We assessed and compared the robustness of our audio beat tracker through a set of experimental setups, under different live acoustic conditions of incremental complexity. These included the presence of continuous musical stimuli, built of a set of concatenated musical pieces; the presence of noises of different natures (e.g., robot motion, speech); and the simultaneous processing of different audio sources on-the-fly, for music and speech. We successfully tackled all these challenging acoustic conditions and improved the beat tracking accuracy and reaction time to music transitions while simultaneously achieving robust automatic speech recognition.
João Lobato Oliveira, Gökhan Ince, Keisuke Nakamura, Kazuhiro Nakadai, Hiroshi G. Okuno, Luís Paulo Reis, Fabien Gouyon
IROS5
2012 Unified auditory functions based on Bayesian topic model
abstract
Existing auditory functions for robots such as sound source localization and separation have been implemented in a cascaded framework whose overall performance may be degraded by any failure in its subsystems. These approaches often require a careful and environment-dependent tuning for each subsystems to achieve better performance. This paper presents a unified framework for sound source localization and separation where the whole system is integrated as a Bayesian topic model. This method improves both localization and separation with a common configuration under various environments by iterative inference using Gibbs sampling. Experimental results from three environments of different reverberation times confirm that our method outperforms state-of-the-art sound source separation methods, especially in the reverberant environments, and shows localization performance comparable to that of the existing robot audition system.
Takuma Otsuka, Katsuhiko Ishiguro, Hiroshi Sawada, Hiroshi G. Okuno
IROS4
2012 Sound sources selection system by using onomatopoeic querries from multiple sound sources
abstract
Our motivation is to develop a robot that treats auditory information in real environment because auditory information is useful for animated communications or understanding our surroundings. Interactions by using sound information need an aquisition of it and a proper sound source reference between a user and a robot leads to it. Such sound source reference is difficult due to multiple sound sources generating in real environemnt, and we use onomatopoeic representations as a representation for the reference. This paper shows a system that selects a sound source specified by a user from multiple sound sources. Users use onomatopoeias in the specification, and our system separates a mixed sound and converts separated sounds into onomatopoeias for the selection. Onomatopoeais have the ambiguity that each user gives each expression to a certain sound and we create an original similarity based on Minimum Edit Distance and acoustic features for solving its problem. In experiments, our system receives a mixed sound consisting of three sounds and a user's query as inputs, and checks a count of a consistency of a sound source selected by a system and a sound source specified by a user in 100 tests. The result shows.
Yusuke Yamamura, Toru Takahashi 0001, Tetsuya Ogata, Hiroshi G. Okuno
IROS4
2012 An active audition framework for auditory-driven HRI: Application to interactive robot dancing
abstract
In this paper we propose a general active audition framework for auditory-driven Human-Robot Interaction (HRI). The proposed framework simultaneously processes speech and music on-the-fly, integrates perceptual models for robot audition, and supports verbal and non-verbal interactive communication by means of (pro)active behaviors. To ensure a reliable interaction, on top of the framework a behavior decision mechanism based on active audition policies the robot's actions according to the reliability of the acoustic signals for auditory processing. To validate the framework's application to general auditory-driven HRI, we propose the implementation of an interactive robot dancing system. This system integrates three preprocessing robot audition modules: sound source localization, sound source separation, and ego noise suppression; two modules for auditory perception: live audio beat tracking and automatic speech recognition; and multi-modal behaviors for verbal and non-verbal interaction: music-driven dancing and speech-driven dialoguing. To fully assess the system, we set up experimental and interactive real-world scenarios with highly dynamic acoustic conditions, and defined a set of evaluation criteria. The experimental tests revealed accurate and robust beat tracking and speech recognition, and convincing dance beat-synchrony. The interactive sessions confirmed the fundamental role of the behavior decision mechanism for actively maintaining a robust and natural human-robot interaction.
João Lobato Oliveira, Gökhan Ince, Keisuke Nakamura, Kazuhiro Nakadai, Hiroshi G. Okuno, Luís Paulo Reis, Fabien Gouyon
RO-MAN5
2012 Efficient Blind Dereverberation and Echo Cancellation Based on Independent Component Analysis for Actual Acoustic Signals
abstract
This letter presents a new algorithm for blind dereverberation and echo cancellation based on independent component analysis (ICA) for actual acoustic signals. We focus on frequency domain ICA (FD-ICA) because its computational cost and speed of learning convergence are sufficiently reasonable for practical applications such as hands-free speech recognition. In applying conventional FD-ICA as a preprocessing of automatic speech recognition in noisy environments, one of the most critical problems is how to cope with reverberations. To extract a clean signal from the reverberant observation, we model the separation process in the short-time Fourier transform domain and apply the multiple input/output inverse-filtering theorem (MINT) to the FD-ICA separation model. A naive implementation of this method is computationally expensive, because its time complexity is the second order of reverberation time. Therefore, the main issue in dereverberation is to reduce the high computational cost of ICA. In this letter, we reduce the computational complexity to the linear order of the reverberation time by using two techniques: (1) a separation model based on the independence of delayed observed signals with MINT and (2) spatial sphering for preprocessing. Experiments show that the computational cost grows in proportion to the linear order of the reverberation time and that our method improves the word correctness of automatic speech recognition by 10 to 20 points in a RT₂₀= 670 ms reverberant environment.
Ryu Takeda, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
Neural Comput.6
2011 Cluster Self-organization of Known and Unknown Environmental Sounds Using Recurrent Neural Network
Shun Nishide, Toru Takahashi 0001, Hiroshi G. Okuno, Tetsuya Ogata
ICANN (1)4
2011 Simultaneous processing of sound source separation and musical instrument identification using Bayesian spectral modeling
abstract
This paper presents a method of both separating audio mixtures into sound sources and identifying the musical instruments of the sources. A statistical tone model of the power spectrogram, called an integrated model, is defined and source separation and instrument identification are carried out on the basis of Bayesian inference. Since, the parameter distributions of the integrated model depend on each instrument, the instrument name is identified by selecting the one that has the maximum relative instrument weight. Experimental results showed correct instrument identification enables precise source separation even when many overtones overlap.
Katsutoshi Itoyama, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ICASSP5
2011 Polyphonic audio-to-score alignment based on Bayesian Latent Harmonic Allocation Hidden Markov Model
abstract
This paper presents a Bayesian method for temporally aligning a music score and an audio rendition. A critical problem in audio-to-score alignment is in dealing with the wide variety of timbre and volume of the audio rendition. In contrast with existing works that achieve this through ad-hoc feature design or careful training of tone models, we propose a Bayesian audio-to-score alignment method by modeling music performance as a Bayesian Hidden Markov Model, each state of which emits a Bayesian signal model based on Latent Harmonic Allocation. After attenuating reverberation, variational Bayes method is used to iteratively adapt the alignment, instrument tone model and the volume balance at each position of the score. The method is evaluated using sixty works of classical music of a variety of instrumentation ranging from solo piano to full orchestra. We verify that our method improves the alignment accuracy compared to dynamic time warping based on chroma vector for orchestral music, or our method employed in a maximum likelihood setting.
Akira Maezawa, Hiroshi G. Okuno, Tetsuya Ogata, Masataka Goto
ICASSP2
2011 I-Divergence-based dereverberation method with auxiliary function approach
abstract
This paper presents a dereverberation method based on I-divergence minimization, which is particularly suitable for music signals. Existing dereverberation methods, including one designed for music, sometimes distort instrument sounds and make staccato-like tones. The problems with the Itakura-Saito-divergence-based formulation of the existing methods are their tendency to excessive suppression of direct sound and the difficulty of incorporating and optimizing sophisticated source models suitable for music signals. The proposed I-divergence-based method can mitigate these problems. Employing the I-divergence measure enables us to avoid the direct sound suppression problem and to use powerful music spectrum models without complicating its optimization. We develop a convergence-guaranteed parameter estimation algorithm based on the auxiliary function approach. Experimental results reveal the effectiveness of the proposed dereverberation method.
Naoki Yasuraoka, Hirokazu Kameoka, Takuya Yoshioka, Hiroshi G. Okuno
ICASSP4
2011 Use of a Sparse Structure to Improve Learning Performance of Recurrent Neural Networks
Hiromitsu Awano, Shun Nishide, Hiroaki Arie, Jun Tani, Toru Takahashi 0001, Hiroshi G. Okuno, Tetsuya Ogata
ICONIP (3)6
2011 Design and implementation of selectable sound separation on the Texai telepresence system using HARK
abstract
This paper presents the design and implementation of selectable sound separation functions on the telepresence system "Texai" using the robot audition software "HARK." An operator of Texai can "walk" around a faraway office to attend a meeting or talk with people through video-conference instead of meeting in person. With a normal microphone, the operator has difficulty recognizing the auditory scene of the Texai, e.g., he/she cannot know the number and the locations of sounds. To solve this problem, we design selectable sound separation functions with 8 microphones in two modes, overview and filter modes, and implement them using HARK's sound source localization and separation. The overview mode visualizes the direction-of-arrival of surrounding sounds, while the filter mode provides sounds that originate from the range of directions he/she specifies. The functions enable the operator to be aware of a sound even if it comes from behind the Texai, and to concentrate on a particular sound. The design and implementation was completed in five days due to the portability of HARK. Experimental evaluations with actual and simulated data show that the resulting system localizes sound sources with a tolerance of 5 degrees.
Takeshi Mizumoto, Kazuhiro Nakadai, Takami Yoshida, Ryu Takeda, Takuma Otsuka, Toru Takahashi 0001, Hiroshi G. Okuno
ICRA7
2011 Robot with Two Ears Listens to More than Two Simultaneous Utterances by Exploiting Harmonic Structures
Yasuharu Hirasawa, Toru Takahashi 0001, Tetsuya Ogata, Hiroshi G. Okuno
IEA/AIE (1)4
2011 Environmental Sound Recognition for Robot Audition Using Matching-Pursuit
Nobuhide Yamakawa, Toru Takahashi 0001, Tetsuro Kitahara, Tetsuya Ogata, Hiroshi G. Okuno
IEA/AIE (2)5
2011 Fast and Simple Iterative Algorithm of Lp-Norm Minimization for Under-Determined Speech Separation
abstract
This paper presents an efficient algorithm to solve Lp-norm minimization problem for under-determined speech separation; that is, for the case that there are more sound sources than microphones. We employ an auxiliary function method in order to derive update rules under the assumption that the amplitude of each sound source follows generalized Gaussian distribution. Experiments reveal that our method solves the L1-norm minimization problem ten times faster than a general solver, and also solves Lp-norm minimization problem efficiently, especially when the parameter p is small; when p is not more than 0.7, it runs in real-time without loss of separation quality. Index Terms: speech separation, under-determined condition, Lp-norm minimization, auxiliary function method
Yasuharu Hirasawa, Naoki Yasuraoka, Toru Takahashi 0001, Tetsuya Ogata, Hiroshi G. Okuno
INTERSPEECH5
2011 Bayesian Extension of MUSIC for Sound Source Localization and Tracking
abstract
This paper presents a Bayesian extension of MUSIC-based sound source localization (SSL) and tracking method. SSL is important for distant speech enhancement and simultaneous speech separation for improving speech recognition, as well as for auditory scene analysis by mobile robots. One of the drawbacks of existing SSL methods is the necessity of careful parameter tunings, e.g., the sound source detection threshold depending on the reverberation time and the number of sources. Our contribution consists of (1) automatic parameter estimation in the variational Bayesian framework and (2) tracking of sound sources with reliability. Experimental results demonstrate our method robustly tracks multiple sound sources in a reverberant environment with RT20 = 840 (ms). Index Terms: simultaneous sound source localization, MUSIC algorithm, variational Bayes, particle filter
Takuma Otsuka, Kazuhiro Nakadai, Tetsuya Ogata, Hiroshi G. Okuno
INTERSPEECH4
2011 Particle-filter based audio-visual beat-tracking for music robot ensemble with human guitarist
abstract
This paper presents an audio-visual beat-tracking method for ensemble robots with a human guitarist. Beat-tracking, or estimation of tempo and beat times of music, is critical to the high quality of musical ensemble performance. Since a human plays the guitar in out-beat in back beat and syncopation, the main problems of beat-tracking of a human's guitar playing are twofold: tempo changes and varying note lengths. Most conventional methods have not addressed human's guitar playing. Therefore, they lack the adaptation of either of the problems. To solve the problems simultaneously, our method uses not only audio but visual features. We extract audio features with Spectro-Temporal Pattern Matching (STPM) and visual features with optical flow, mean shift and Hough transform. Our beat-tracking estimates tempo and beat time using a particle filter; both acoustic feature of guitar sounds and visual features of arm motions are represented as particles. The particle is determined based on prior distribution of audio and visual features, respectively Experimental results confirm that our integrated audio-visual approach is robust against tempo changes and varying note lengths. In addition, they also show that estimation convergence rate depends only a little on the number of particles. The real-time factor is 0.88 when the number of particles is 200, and this shows out method works in real-time.
Tatsuhiko Itohara, Takuma Otsuka, Takeshi Mizumoto, Tetsuya Ogata, Hiroshi G. Okuno
IROS5
2011 Improvement of speaker localization by considering multipath interference of sound wave for binaural robot audition
abstract
This paper presents an improved speaker localization method based on the generalized cross-correlation (GCC) method weighted by the phase transform (PHAT) for binaural robot audition. The problem with the conventional direction-of-arrival (DOA) estimation based on the GCC-PHAT method is a multipath interference whereby a sound wave travels to microphones via the front-head path and the back-head path in binaural robot audition. This paper describes a new time delay factor for the GCC-PHAT method to compensate multipath interference on the assumption of spherical robot head. In addition, the restriction of the time difference of arrival (TDOA) estimation by the sampling frequency is also solved by applying the maximum likelihood (ML) estimation in frequency domain. Experiments conducted in the SIG-2 humanoid robot show that the proposed method reduces localization errors by 17.8 degrees on average and by over 35 degrees in side directions comparing to the conventional DOA estimation.
Ui-Hyun Kim, Takeshi Mizumoto, Tetsuya Ogata, Hiroshi G. Okuno
IROS4
2011 A Two-Stage Domain Selection Framework for Extensible Multi-Domain Spoken Dialogue Systems
Mikio Nakano, Kazunori Komatani, Kyoko Matsuyama, Kotaro Funakoshi, Hiroshi G. Okuno
SIGDIAL Conference6
2011 Handwriting prediction based character recognition using recurrent neural network
abstract
Humans are said to unintentionally trace handwriting sequences in their brains based on handwriting experiences when recognizing written text. In this paper, we propose a model for predicting handwriting sequence for written text recognition based on handwriting experiences. The model is first trained using image sequences acquired while writing text. The image features of sequences are self-organized from the images using Self-Organizing Map. The feature sequences are used to train a neuro-dynamics learning model. For recognition, the text image is input into the model for predicting the handwriting sequence and recognition of the text. We conducted two experiments using ten Japanese characters. The results of the experiments show the effectivity of the model.
Shun Nishide, Hiroshi G. Okuno, Tetsuya Ogata, Jun Tani
SMC2
2011 A multi-expert model for dialogue and behavior control of conversational robots and agents
Mikio Nakano, Yuji Hasegawa, Kotaro Funakoshi, Johane Takeuchi, Toyotaka Torii, Kazuhiro Nakadai, Naoyuki Kanda, Kazunori Komatani, Hiroshi G. Okuno, Hiroshi Tsujino
Knowl. Based Syst.9
2011 Emergence of hierarchical structure mirroring linguistic composition in a recurrent neural network
Wataru Hinoshita, Hiroaki Arie, Jun Tani, Hiroshi G. Okuno, Tetsuya Ogata
Neural Networks4
2011 Blind Separation and Dereverberation of Speech Mixtures by Joint Optimization
abstract
This paper proposes a method for performing blind source separation (BSS) and blind dereverberation (BD) at the same time for speech mixtures. In most previous studies, BSS and BD have been investigated separately. The separation performance of conventional BSS methods deteriorates as the reverberation time increases while many existing BD methods rely on the assumption that there is only one sound source in a room. Therefore, it has been difficult to perform both BSS and BD when the reverberation time is long. The proposed method uses a network, in which dereverberation and separation networks are connected in tandem, to estimate source signals. The parameters for the dereverberation network (prediction matrices) and those for the separation network (separation matrices) are jointly optimized. This enables a BD process to take a BSS process into account. The prediction and separation matrices are alternately optimized with each depending on the other; hence, we call the proposed method the conditional separation and dereverberation (CSD) method. Comprehensive evaluation results are reported, where all the speech materials contained in the complete test set of the TIMIT corpus are used. The CSD method improves the signal-to-interference ratio by an average of about 4 dB over the conventional frequency-domain BSS approach for reverberation times of 0.3 and 0.5 s. The direct-to-reverberation ratio is also improved by about 10 dB.
Takuya Yoshioka, Tomohiro Nakatani, Masato Miyoshi, Hiroshi G. Okuno
IEEE Trans. Speech Audio Process.4
2010 Design and Implementation of Two-level Synchronization for Interactive Music Robot
abstract
Our goal is to develop an interactive music robot, i.e., a robot that presents a musical expression together with humans. A music interaction requires two important functions: synchronization with the music and musical expression, such as singing and dancing. Many instrument-performing robots are only capable of the latter function, they may have difficulty in playing live with human performers. The synchronization function is critical for the interaction. We classify synchronization and musical expression into two levels: (1) the rhythm level and (2) the melody level. Two issues in achieving two-layer synchronization and musical expression are: (1) simultaneous estimation of the rhythm structure and the current part of the music and (2) derivation of the estimation confidence to switch behavior between the rhythm level and the melody level. This paper presents a score following algorithm, incremental audio to score alignment, that conforms to the two-level synchronization design using a particle filter. Our method estimates the score position for the melody level and the tempo for the rhythm level. The reliability of the score position estimation is extracted from the probability distribution of the score position. Experiments are carried out using polyphonic jazz songs. The results confirm that our method switches levels in accordance with the difficulty of the score estimation. When the tempo of the music is less than 120 (beats per minute; bpm), the estimated score positions are accurate and reported; when the tempo is over 120 (bpm), the system tends to report only the tempo to suppress the error in the reported score position predictions.
Takuma Otsuka, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
AAAI6
2010 Music dereverberation using harmonic structure source model and Wiener filter
abstract
This paper proposes a dereverberation method for musical audio signals. Existing dereverberation methods are designed for speech signals and are not necessarily effective for suppressing long and dense reverberation in musical audio signals because: 1) an all-pole model and a non-parametric model, which are used to represent source spectra, do not match musical tones, and 2) the conventional inverse-filter-based dereverberation is not effective for suppressing long and dense reverberation. To overcome the two problems, an appropriate dereverberation approach for musical audio signals is established. The first problem is resolved by using a harmonic Gaussian mixture model (GMM) to accurately model the harmonic structure of a source spectrum. The second problem is resolved by performing dereverberation with a Wiener filter based on both an estimated inverse filter and an estimated source spectrum model. Experimental results reveal the effectiveness of the proposed dereverberation method using these two solutions.
Naoki Yasuraoka, Takuya Yoshioka, Tomohiro Nakatani, Atsushi Nakamura, Hiroshi G. Okuno
ICASSP5
2010 Noisy speech enhancement based on prior knowledge about spectral envelope and harmonic structure
abstract
This paper considers the enhancement of noisy speech. Earlier studies have revealed that an approach that enhances spectral envelopes by using prior knowledge about the all-pole (AP) model parameters of clean speech learnt from speech corpora is advantageous in terms of the amount of musical noise and speech distortion. This paper proposes a new speech enhancement method, in which harmonic structure enhancement is incorporated in learning-based spectral envelope enhancement to further improve performance. The harmonic structure is represented by using a harmonic Gaussian mixture model (GMM), which is parameterized by a voicing indicator and a fundamental frequency. The parameters of the AP model and the harmonic GMM are jointly estimated by maximum a posteriori estimation, thus enabling the enhancement of spectral envelopes and harmonic structures in a unified framework. The proposed method outperforms the spectral envelope enhancement approach by 0.85 dB in cepstral distance.
Takuya Yoshioka, Tomohiro Nakatani, Hiroshi G. Okuno
ICASSP3
2010 Improvement in listening capability for humanoid robot HRP-2
abstract
This paper describes improvement of sound source separation for a simultaneous automatic speech recognition (ASR) system of a humanoid robot. A recognition error in the system is caused by a separation error and interferences of other sources. In separability, an original geometric source separation (GSS) is improved. Our GSS uses a measured robot's head related transfer function (HRTF) to estimate a separation matrix. As an original GSS uses a simulated HRTF calculated based on a distance between microphone and sound source, there is a large mismatch between the simulated and the measured transfer functions. The mismatch causes a severe degradation of recognition performance. Faster convergence speed of separation matrix reduces separation error. Our approach gives a nearer initial separation matrix based on a measured transfer function from an optimal separation matrix than a simulated one. As a result, we expect that our GSS improves the convergence speed. Our GSS is also able to handle an adaptive step-size parameter. These new features are added into open source robot audition software (OSS) called "HARK" which is newly updated as version 1.0.0. The HARK has been installed on a HRP-2 humanoid with an 8-element microphone array. The listening capability of HRP-2 is evaluated by recognizing a target speech signal which is separated from a simultaneous speech signal by three talkers. The word correct rate (WCR) of ASR improves by 5 points under normal acoustic environments and by 10 points under noisy environments. Experimental results show that HARK 1.0.0 improves the robustness against noises.
Toru Takahashi 0001, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ICRA5
2010 Upper-limit evaluation of robot audition based on ICA-BSS in multi-source, barge-in and highly reverberant conditions
abstract
This paper presents the upper-limit evaluation of robot audition based on ICA-BSS in multi-source, barge-in and highly reverberant conditions. The goal is that the robot can automatically distinguish a target speech from its own speech and other sound sources in a reverberant environment. We focus on the multi-channel semi-blind ICA (MCSB-ICA), which is one of the sound source separation methods with a microphone array, to achieve such an audition system because it can separate sound source signals including reverberations with few assumptions on environments. The evaluation of MCSB-ICA has been limited to robot's speech separation and reverberation separation. In this paper, we evaluate MCSB-ICA extensively by applying it to multi-source separation problems under common reverberant environments. Experimental results prove that MCSB-ICA outperforms conventional ICA by 30 points in automatic speech recognition performance.
Ryu Takeda, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ICRA6
2010 Recognition and Generation of Sentences through Self-organizing Linguistic Hierarchy Using MTRNN
Wataru Hinoshita, Hiroaki Arie, Jun Tani, Tetsuya Ogata, Hiroshi G. Okuno
IEA/AIE (3)5
2010 Violin Fingering Estimation Based on Violin Pedagogical Fingering Model Constrained by Bowed Sequence Estimation from Audio Input
Akira Maezawa, Katsutoshi Itoyama, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IEA/AIE (3)6
2010 Improving Identification Accuracy by Extending Acceptable Utterances in Spoken Dialogue System Using Barge-in Timing
Kyoko Matsuyama, Kazunori Komatani, Toru Takahashi 0001, Tetsuya Ogata, Hiroshi G. Okuno
IEA/AIE (2)5
2010 Music-Ensemble Robot That Is Capable of Playing the Theremin While Listening to the Accompanied Music
Takuma Otsuka, Takeshi Mizumoto, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IEA/AIE (1)7
2010 System for Supporting Web-based Public Debate Using Transcripts of Face-to-Face Meeting
Shun Shiramatsu, Jun Takasaki, Tatiana Zidrasco, Tadachika Ozono, Toramatsu Shintani, Hiroshi G. Okuno
IEA/AIE (3)6
2010 An Improvement in Audio-Visual Voice Activity Detection for Automatic Speech Recognition
Takami Yoshida, Kazuhiro Nakadai, Hiroshi G. Okuno
IEA/AIE (1)3
2010 Analyzing user utterances in barge-in-able spoken dialogue system for improving identification accuracy
abstract
In our barge-in-able spoken dialogue system, the user’s behaviors such as barge-in timing and utterance expressions vary according to his/her characteristics and situations. The system adapts to the behaviors by modeling them. We analyzed 1584 utterances collected by our systems of quiz and news-listing tasks and showed that ratio of using referential expressions depends on individual users and average lengths of listed items. This tendency was incorporated as a prior probability into our method and improved the identification accuracy of the user’s intended items. Index Terms: barge-in, spoken dialogue systems, utterance timing, user characteristics
Kyoko Matsuyama, Kazunori Komatani, Ryu Takeda, Toru Takahashi 0001, Tetsuya Ogata, Hiroshi G. Okuno
INTERSPEECH6
2010 Effects of modelling within- and between-frame temporal variations in power spectra on non-verbal sound recognition
abstract
Research on environmental sound recognition has not shown great development in comparison with that on speech and musical signals. One of the reasons is that the sound category of environmental sounds covers a broad range of acoustical natures. We classified them in order to explore suitable recognition techniques for each characteristic. We focus on impulsive sounds and their non-stationary feature within and between analytic frames. We used matching-pursuit as a framework to use wavelet analysis for extracting temporal variation of audio features inside a frame. We also investigated the validity of modeling decaying patterns of sounds using Hidden markov models. Experimental results indicate that sounds with multiple impulsive signals are recognized better by using time-frequency analyzing bases than by frequency domain analysis. Classification of sound classes with a long and clear decaying pattern improves when HMMs with multiple number of hidden states are applied. Index Terms: audio signal classification, non-speech sound recognition, environmental sound recognition, time-frequency analysis, Matching-Pursuit
Nobuhide Yamakawa, Tetsuro Kitahara, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
INTERSPEECH6
2010 Exploiting harmonic structures to improve separating simultaneous speech in under-determined conditions
abstract
In real-world situations, a robot may often encounter “under-determined” situation, where there are more sound sources than microphones. This paper presents a speech separation method using a new constraint on the harmonic structure for a simultaneous speech-recognition system in under-determined conditions. The requirements for a speech separation method in a simultaneous speech-recognition system are (1) ability to handle a large number of talkers, and (2) reduction of distortion in acoustic features. Conventional methods use a maximum likelihood estimation in sound source separation, which fulfills requirement (1). Since it is a general approach, the performance is limited when separating speech. This paper presents a two-stage method to improve the separation. The first stage uses maximum likelihood estimation and extracts the harmonic structure, and the second stage exploits the harmonic structure as a new constraint to achieve requirement (2). We carried out an experiment that simulated three simultaneous utterances using impulse responses recorded by two microphones in an anechoic chamber. The experimental results revealed that our method could improve speech recognition correctness by about four points.
Yasuharu Hirasawa, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS5
2010 Robot musical accompaniment: integrating audio and visual cues for real-time synchronization with a human flutist
abstract
Musicians often have the following problem: they have a music score that requires 2 or more players, but they have no one with whom to practice. So far, score-playing music robots exist, but they lack adaptive abilities to synchronize with fellow players' tempo variations. In other words, if the human speeds up their play, the robot should also increase its speed. However, computer accompaniment systems allow exactly this kind of adaptive ability. We present a first step towards giving these accompaniment abilities to a music robot. We introduce a new paradigm of beat tracking using 2 types of sensory input - visual and audio - using our own visual cue recognition system and state-of-the-art acoustic onset detection techniques. Preliminary experiments suggest that by coupling these two modalities, a robot accompanist can start and stop a performance in synchrony with a flutist, and detect tempo changes within half a second.
Angelica Lim, Takeshi Mizumoto, Louis-Kenzo Cahier, Takuma Otsuka, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS8
2010 Human-robot ensemble between robot thereminist and human percussionist using coupled oscillator model
abstract
This paper presents a novel synchronizing method for a human-robot ensemble using coupled oscillators. We define an ensemble as a synchronized performance produced through interactions between independent players. To attain better synchronized performance, the robot should predict the human's behavior to reduce the difference between the human's and robot's onset timings. Existing studies in such synchronization only adapts to onset intervals, thus, need a considerable time to synchronize. We use a coupled oscillator model to predict the human's behavior. Experimental results show that our method reduces the average of onset time errors; when we use a metronome, a tempo-varying metronome or a human drummer, errors are reduced by 38%, 10% or 14% on the average, respectively. These results mean that the prediction of human's behaviors is effective for the synchronized performance.
Takeshi Mizumoto, Takuma Otsuka, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS7
2010 Motion generation based on reliable predictability using self-organized object features
abstract
Predictability is an important factor for determining robot motions. This paper presents a model to generate robot motions based on reliable predictability evaluated through a dynamics learning model which self-organizes object features. The model is composed of a dynamics learning module, namely Recurrent Neural Network with Parametric Bias (RNNPB), and a hierarchical neural network as a feature extraction module. The model inputs raw object images and robot motions. Through bi-directional training of the two models, object features which describe the object motion are self-organized in the output of the hierarchical neural network, which is linked to the input of RNNPB. After training, the model searches for the robot motion with high reliable predictability of object motion. Experiments were performed with the robot's pushing motion with a variety of objects to generate sliding, falling over, bouncing, and rolling motions. For objects with single motion possibility, the robot tended to generate motions that induce the object motion. For objects with two motion possibilities, the robot evenly generated motions that induce the two object motions.
Shun Nishide, Tetsuya Ogata, Jun Tani, Toru Takahashi 0001, Kazunori Komatani, Hiroshi G. Okuno
IROS6
2010 An improvement in automatic speech recognition using soft missing feature masks for robot audition
abstract
We describe integration of preprocessing and automatic speech recognition based on Missing-Feature-Theory (MFT) to recognize a highly interfered speech signal, such as the signal in a narrow angle between a desired and interfered speakers. As a speech signal separated from a mixture of speech signals includes the leakage from other speech signals, recognition performance of the separated speech degrades. An important problem is estimating the leakage in time-frequency components. Once the leakage is estimated, we can generate missing feature masks (MFM) automatically by using our method. A new weighted sigmoid function is introduced for our MFM generation method. An experiment shows that a word correct rate improves from 66 % to 74 % by using our MFM generation method tuned by a search base approach in the parameter space.
Toru Takahashi 0001, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS5
2010 Speedup and performance improvement of ICA-based robot audition by parallel and resampling-based block-wise processing
abstract
This paper describes a speedup and performance improvement of multi-channel semi-blind ICA (MCSB-ICA) with parallel and resampling-based block-wise processing. MCSB-ICA is an integrated method of sound source separation that accomplishes blind source separation, blind dereverberation, and echo cancellation. This method enables robots to separate user's speech signals from observed signals including the robot's own speech, other speech and their reverberations without a priori information. The main problem when MCSB-ICA is applied to robot audition is its high computational cost. We tackle this by multi-threading programming, and the two main issues are 1) the design of parallel processing and 2) incremental implementation. These are solved by a) multiple-stack-based parallel implementation, and b) resampling-based overlaps and block-wise separation. The experimental results proved that our method reduced the real-time factor to less than 0.5 with an eight-core CPU, and it improves the performance of automatic speech recognition by 2-10 points compared with the single-stack-based parallel implementation without the resampling technique.
Ryu Takeda, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS6
2010 Two-layered audio-visual speech recognition for robots in noisy environments
abstract
Audio-visual (AV) integration is one of the key ideas to improve perception in noisy real-world environments. This paper describes automatic speech recognition (ASR) to improve human-robot interaction based on AV integration. We developed AV-integrated ASR, which has two AV integration layers, that is, voice activity detection (VAD) and ASR. However, the system has three difficulties: 1) VAD and ASR have been separately studied although these processes are mutually dependent, 2) VAD and ASR assumed that high resolution images are available although this assumption never holds in the real world, and 3) an optimal weight between audio and visual stream was fixed while their reliabilities change according to environmental changes. To solve these problems, we propose a new VAD algorithm taking ASR characteristics into account, and a linear-regression-based optimal weight estimation method. We evaluate the algorithm for auditory-and/or visually-contaminated data. Preliminary results show that the robustness of VAD improved even when the resolution of the images is low, and the AVSR using estimated stream weight shows the effectiveness of AV integration.
Takami Yoshida, Kazuhiro Nakadai, Hiroshi G. Okuno
IROS3
2010 Online Error Detection of Barge-In Utterances by Using Individual Users' Utterance Histories in Spoken Dialogue System
Kazunori Komatani, Hiroshi G. Okuno
SIGDIAL Conference2
2010 Human-robot cooperation in arrangement of objects using confidence measure of neuro-dynamical system
abstract
The objective of our study was to develop dynamic collaboration between a human and a robot. Most conventional studies have created pre-designed rule-based collaboration systems to determine the timing and behavior of robots to participate in tasks. Our aim is to introduce the confidence of the task as a criterion for robots to determine their timing and behavior. In this paper, we report the effectiveness of applying reproduction accuracy as a measure for quantitatively evaluating confidence in an object arrangement task. Our method is comprised of three phases. First, we obtain human-robot interaction data through the Wizard of OZ method. Second, the obtained data are trained using a neuro-dynamical system, namely, the Multiple Time-scales Recurrent Neural Network (MTRNN). Finally, the prediction error in MTRNN is applied as a confidence measure to determine the robot's behavior. The robot participated in the task when its confidence was high, while it just observed when its confidence was low. Training data were acquired using an actual robot platform, Hiro. The method was evaluated using a robot simulator. The results revealed that motion trajectories could be precisely reproduced with a high degree of confidence, demonstrating the effectiveness of the method.
Hiromitsu Awano, Tetsuya Ogata, Shun Nishide, Toru Takahashi 0001, Kazunori Komatani, Hiroshi G. Okuno
SMC6
2010 Inter-modality mapping in robot with recurrent neural network
Tetsuya Ogata, Shun Nishide, Hideki Kozima, Kazunori Komatani, Hiroshi G. Okuno
Pattern Recognit. Lett.5
2010 A Modeling of Singing Voice Robust to Accompaniment Sounds and Its Application to Singer Identification and Vocal-Timbre-Similarity-Based Music Information Retrieval
abstract
This paper describes a method of modeling the characteristics of a singing voice from polyphonic musical audio signals including sounds of various musical instruments. Because singing voices play an important role in musical pieces with vocals, such representation is useful for music information retrieval systems. The main problem in modeling the characteristics of a singing voice is the negative influences caused by accompaniment sounds. To solve this problem, we developed two methods,accompaniment sound reduction and reliable frame selection.The former makes it possible to calculate feature vectors that represent a spectral envelope of a singing voice after reducing accompaniment sounds. It first extracts the harmonic components of the predominant melody from sound mixtures and then resynthesizes the melody by using a sinusoidal model driven by these components. The latter method then estimates the reliability of frame of the obtained melody (i.e., the influence of accompaniment sound) by using two Gaussian mixture models (GMMs) for vocal and nonvocal frames to select the reliable vocal portions of musical pieces. Finally, each song is represented by its GMM consisting of the reliable frames. This new representation of the singing voice is demonstrated to improve the performance of an automatic singer identification system and to achieve an MIR system based on vocal timbre similarity.
Hiromasa Fujihara, Masataka Goto, Tetsuro Kitahara, Hiroshi G. Okuno
IEEE Trans. Speech Audio Process.4
2009 ICA-based efficient blind dereverberation and echo cancellation method for barge-in-able robot audition
abstract
This paper describes a new method that allows ldquoBarge-Inrdquo in various environments for robot audition. ldquoBarge-inrdquo means that a user begins to speak simultaneously while a robot is speaking. To achieve the function, we must deal with problems on blind dereverberation and echo cancellation at the same time. We adopt Independent Component Analysis (ICA) because it essentially provides a natural framework for these two problems. To deal with reverberation, we apply a Multiple Input/Output INverse-filtering Theorem-based model of observation to the frequency domain ICA. The main problem is its high-computational cost of ICA. We reduce the computational complexity to the linear order of reverberation time by using two techniques: 1) a separation modelbased on observed signal independence, and 2) enforced spatial sphering for preprocessing. The experimental results revealed that our method improved word correctness of reverberant speech by 10-20 points.
Ryu Takeda, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ICASSP6
2009 Continuous vocal imitation with self-organized vowel spaces in Recurrent Neural Network
abstract
A continuous vocal imitation system was developed using a computational model that explains the process of phoneme acquisition by infants. Human infants perceive speech sounds not as discrete phoneme sequences but as continuous acoustic signals. One of critical problems in phoneme acquisition is the design for segmenting these continuous speech sounds. The key idea to solve this problem is that articulatory mechanisms such as the vocal tract help human beings to perceive speech sound units corresponding to phonemes. To segment acoustic signal with articulatory movement, we apply the segmenting method to our system by Recurrent Neural Network with Parametric Bias (RNNPB). This method determines the multiple segmentation boundaries in a temporal sequence using the prediction error of the RNNPB model, and the PB values obtained by the method can be encoded as kind of phonemes. Our system was implemented by using a physical vocal tract model, called the Maeda model. Experimental results demonstrated that our system can self-organize the same phonemes in different continuous sounds, and can imitate vocal sound involving arbitrary numbers of vowels using the vowel space in the RNNPB. This suggests that our model reflects the process of phoneme acquisition.
Hisashi Kanda, Tetsuya Ogata, Toru Takahashi 0001, Kazunori Komatani, Hiroshi G. Okuno
ICRA5
2009 Prediction and imitation of other's motions by reusing own forward-inverse model in robots
abstract
This paper proposes a model that enables a robot to predict and imitate the motions of another by reusing its body forward-inverse model. Our model includes three approaches: (i) projection of a self-forward model for predicting phenomena in the external environment (other individuals), (ii) ldquotriadic relationrdquo that is mediation by a physical object between self and others, (iii) introduction of infant imitation by a parent. The recurrent neural network with parametric bias (RNNPB) model is used as the robot's self forward-inverse model. A group of hierarchical neural networks are attached to the RNNPB model as ldquoconversion modulesrdquo. Experiments demonstrated that a robot with our model could imitate a human's motions by translating the viewpoint. It could also discriminate known/unknown motions appropriately, and associate whole motion dynamics from only one motion snap image.
Tetsuya Ogata, Ryunosuke Yokoya, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno
ICRA5
2009 Adjusting Occurrence Probabilities of Automatically-Generated Abbreviated Words in Spoken Dialogue Systems
Masaki Katsumaru, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IEA/AIE4
2009 Improving speech understanding accuracy with limited training data using multiple language models and multiple understanding models
abstract
We aim to improve a speech understanding module with a small amount of training data. A speech understanding module uses a language model (LM) and a language understanding model (LUM). A lot of training data are needed to improve the models. Such data collection is, however, difficult in an actual process of development. We therefore design and develop a new framework that uses multiple LMs and LUMs to improve speech understanding accuracy under various amounts of training data. Even if the amount of available training data is small, each LM and each LUM can deal well with different types of utterances and more utterances are understood by using multiple LM and LUM. As one implementation of the framework, we develop a method for selecting the most appropriate speech understanding result from several candidates. The selection is based on probabilities of correctness calculated by logistic regressions. We evaluate our framework with various amounts of training data. Index Terms: speech understanding, multiple language models and language understanding models, limited training data
Masaki Katsumaru, Mikio Nakano, Kazunori Komatani, Kotaro Funakoshi, Tetsuya Ogata, Hiroshi G. Okuno
INTERSPEECH6
2009 Enabling a user to specify an item at any time during system enumeration - item identification for barge-in-able conversational dialogue systems
abstract
In conversational dialogue systems, users prefer to speak at any time and to use natural expressions. We have developed an Independent Component Analysis (ICA) based semi-blind source separation method, which allows users to barge-in over system utterances at any time. We created a novel method from timing information derived from barge-in utterances to identify one item that a user indicates during system enumeration. First, we determine the timing distribution of user utterances containing referential expressions and then approximate it using a gamma distribution. Second, we represent both the utterance timing and automatic speech recognition (ASR) results as probabilities of the desired selection from the system’s enumeration. We then integrate these two probabilities to identify the item having the maximum likelihood of selection. Experimental results using 400 utterances indicated that our method outperformed two methods used as a baseline (one of ASR results only and one of utterance timing only) in identification accuracy. Index Terms: spoken dialogue system, conversational interaction, barge-in, utterance timing
Kyoko Matsuyama, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
INTERSPEECH4
2009 Emergence of evolutionary interaction with voice and motion between two robots using RNN
abstract
We propose a model of evolutionary interaction between two robots where signs used for communication emerge through mutual adaptation. Signs used in human interaction, e.g., language, gestures and eye contact change and evolve in form and meaning through repeated use. To create flexible human-like interaction systems, it is necessary to deal with signs as a dynamic property and to construct a framework in which signs emerge from mutual adaptation by agents. Our target is multi-modal interaction using voice and motion between two robots where a voice/motion pattern is used as a sign referring to a motion/voice pattern. To enable evolutionary signs (voice and motion patterns) to be recognized and generated, we utilized a dynamics model: Multiple Timescale Recurrent Neural Network (MTRNN). To enable the robots to interpret signs, we utilized hierarchical neural networks, which transform dynamics model parameters of voice/motion into those of motion/voice. In our experiment, two robots modified their own interpretation of signs constantly through mutual adaptation in interaction where they responded to the other's voice with motion one after the other. As a result of the experiment, we found that the interaction kept evolving through the robots' repeated and alternate miscommunications and re-adaptations, and this induced the emergence of diverse new signs that depended on the robots' body dynamics through the generalization capability of MTRNN.
Wataru Hinoshita, Tetsuya Ogata, Hideki Kozima, Hisashi Kanda, Toru Takahashi 0001, Hiroshi G. Okuno
IROS6
2009 Phoneme acquisition model based on vowel imitation using Recurrent Neural Network
abstract
A phoneme-acquisition system was developed using a computational model that explains the developmental process of human infants in the early period of acquiring language. There are two important findings in constructing an infant's acquisition of phonemes: (1) an infant's vowel like cooing tends to invoke utterances that are imitated by its caregiver, and (2) maternal imitation effectively reinforces infant vocalization. Therefore, we hypothesized that infants can acquire phonemes to imitate their caregivers' voices by trial and error, i. e., infants use self-vocalization experience to search for imitable and unimitable elements in their caregivers' voices. On the basis of this hypothesis, we constructed a phoneme acquisition process using interaction involving vowel imitation between a human and an infant model. Our infant model had a vocal tract system, called the Maeda model, and an auditory system implemented by using mel-frequency cepstral coefficients (MFCCs) through STRAIGHT analysis. We applied recurrent neural network with parametric bias (RNNPB) to learn the experience of self-vocalization, to recognize the human voice, and to produce the sound imitated by the infant model. To evaluate imitable and unimitable sounds, we used the prediction error of the RNNPB model. The experimental results revealed that as imitation interactions were repeated, the formants of sounds imitated by our system moved closer to those of human voices, and our system could self-organize the same vowels in different continuous sounds. This suggests that our system can reflect the process of phoneme acquisition.
Hisashi Kanda, Tetsuya Ogata, Toru Takahashi 0001, Kazunori Komatani, Hiroshi G. Okuno
IROS5
2009 Thereminist robot: Development of a robot theremin player with feedforward and feedback arm control based on a Theremin's pitch model
abstract
We propose a Thereminist robot system that plays the Theremin based on a Theremin's pitch model. The theremin, which is a 1920s electronic musical instrument, is played by moving a player's hand position in the air without touching it. It is difficult to play the Theremin because the relationship between the hand position and Theremin's pitch (pitch characteristics) is non-linear and varies according to the electromagnetic field (hereafter called environment). These characteristics cause two problems: (1) adapting to the environment change is required and (2) a nai¿ve design tends to depend on robot's particular hardware. We implement the coarse-to-fine control system on the Thereminist robot using newly proposed two pitch models: parametric and nonparametric ones. The Thereminist robot works as below: first, the robot calibrates the pitch model by parameter fitting with the Levenberg-Marquardt method. Second, the robot moves its hand in a coarse manner by feedforward control based on the pitch model. Finally, the robot adjusts its position by feedback control (proportional-integral control). In these steps, the robot can play a required pitch quickly, because the robot moves its hand using the pitch model without listening to the Theremin's sound Thus, the time to play the exact pitch is shorter than when only feedback control is used. Three experiments were conducted to evaluate the robustness against the number of samples, environment change, and types of robots. The results revealed that our pitch model describes using only 12 samples of pitches for estimation of the parameters, and adapts if the environment changes. In addition, our system works on two different robots: HRP-2 and ASIMO.
Takeshi Mizumoto, Hiroshi Tsujino, Toru Takahashi 0001, Tetsuya Ogata, Hiroshi G. Okuno
IROS5
2009 Modeling tool-body assimilation using second-order Recurrent Neural Network
abstract
Tool-body assimilation is one of the intelligent human abilities. Through trial and experience, humans are capable of using tools as if they are part of their own bodies. This paper presents a method to apply a robot's active sensing experience for creating the tool-body assimilation model. The model is composed of a feature extraction module, dynamics learning module, and a tool recognition module. Self-Organizing Map (SOM) is used for the feature extraction module to extract object features from raw images. Multiple Time-scales Recurrent Neural Network (MTRNN) is used as the dynamics learning module. Parametric Bias (PB) nodes are attached to the weights of MTRNN as second-order network to modulate the behavior of MTRNN based on the tool. The generalization capability of neural networks provide the model the ability to deal with unknown tools. Experiments are performed with HRP-2 using no tool, I-shaped, T-shaped, and L-shaped tools. The distribution of PB values have shown that the model has learned that the robot's dynamic properties change when holding a tool. The results of the experiment show that the tool-body assimilation model is capable of applying to unknown objects to generate goal-oriented motions.
Shun Nishide, Tatsuhiro Nakagawa, Tetsuya Ogata, Jun Tani, Toru Takahashi 0001, Hiroshi G. Okuno
IROS6
2009 Incremental polyphonic audio to score alignment using beat tracking for singer robots
abstract
We aim at developing a singer robot capable of listening to music with its own ¿ears¿ and interacting with a human's musical performance. Such a singer robot requires at least three functions: listening to the music, understanding what position in the music is being performed, and generating a singing voice. In this paper, we focus on the second function, that is, the capability to align an audio signal to its musical score represented symbolically. Issues underlying the score alignment problem are: (1) diversity in the sounds of various musical instruments, (2) difference between the audio signal and the musical score, (3) fluctuation in tempo of the musical performance. Our solutions to these issues are as follows: (1) the design of features based on a chroma vector in the 12-tone model and onset of the sound, (2) defining the rareness for each tone based on the idea that scarcely used tone is salient in the audio signal, and (3) the use of a switching Kalman filter for robust tempo estimation. The experimental result shows that our score alignment method improves the average of cumulative absolute errors in score alignment by 29% using 100 popular music tunes compared to the beat tracking without score alignment.
Takuma Otsuka, Toru Takahashi 0001, Hiroshi G. Okuno, Kazunori Komatani, Tetsuya Ogata, Kazumasa Murata, Kazuhiro Nakadai
IROS3
2009 Missing-feature-theory-based robust simultaneous speech recognition system with non-clean speech acoustic model
abstract
A humanoid robot must recognize a target speech signal while people around the robot chat with them in real-world. To recognize the target speech signal, robot has to separate the target speech signal among other speech signals and recognize the separated speech signal. As separated signal includes distortion, automatic speech recognition (ASR) performance degrades. To avoid the degradation, we trained an acoustic model from non-clean speech signals to adapt acoustic feature of distorted signal and adding white noise to separated speech signal before extracting acoustic feature. The issues are (1) To determine optimal noise level to add the training speech signals, and (2) To determine optimal noise level to add the separated signal. In this paper, we investigate how much noises should be added to clean speech data for training and how speech recognition performance improves for different positions of three talkers with soft masking. Experimental results show that the best performance is obtained by adding white noises of 30 dB. The ASR with the acoustic model outperforms with ASR with the clean acoustic model by 4 points.
Toru Takahashi 0001, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS5
2009 Step-size parameter adaptation of multi-channel semi-blind ICA with piecewise linear model for barge-in-able robot audition
abstract
This paper describes a step-size parameter adaptation technique of multi-channel semi-blind independent component analysis (MCSB-ICA) for a ¿barge-in-able¿ robot audition system. By ¿barge-in¿, we mean that the user can speak simultaneously when the robot is speaking.We focused on MCSB-ICA to achieve such an audition system because it can separate a user's and a robot's speech under reverberant environments. The problem with MCSB-ICA for robot audition is the slow speed of convergence in estimating a separation filter due to its step-size parameters. Many optimization methods cannot be adopted because their computational costs are proportional to the 2nd order of the reverberation time. Our method yields adaptive step-size parameters with MCSB-ICA at low computational costs. It is based on three techniques; (1) recursive expression of the separation process, (2) a piecewise linear model of the step-size of the separation filter, and (3) adaptive step-size parameters with a sub-ICA-filter. Experimental results show that our approach attains faster convergence speed and lower computational costs than those with a fixed step-size parameter.
Ryu Takeda, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS6
2009 Bowed String Sequence Estimation of a Violin Based on Adaptive Audio Signal Classification and Context-Dependent Error Correction
abstract
The sequence of strings played on a bowed string instrument is essential to understanding of the fingering. Thus, its estimation is required for machine understanding of violin playing. Audio-based identification is the only viable way to realize this goal for existing music recordings. A naive implementation using audio classification alone, however, is inaccurate and is not robust against variations in string or instruments. We develop a bowed string sequence estimation method by combining audio-based bowed string classification and context-dependent error correction. The robustness against different setups of instruments improves by normalizing the F0-dependent features using the average feature of a recording. The performance of error correction is evaluated using an electric violin with two different brands of strings and an acoustic violin. By incorporating mean normalization, the recognition error of recognition accuracy due to changing the string alleviates by 8 points, and that due to change of instrument by 12 points. Error correction decreases the error due to change of string by 8 points and that due to different instrument by 9 points.
Akira Maezawa, Katsutoshi Itoyama, Toru Takahashi 0001, Tetsuya Ogata, Hiroshi G. Okuno
ISM5
2009 Robot Audition: Missing Feature Theory Approach and Active Audition
Hiroshi G. Okuno, Kazuhiro Nakadai, Hyun-Don Kim
ISRR1
2009 Changing timbre and phrase in existing musical performances as you like: manipulations of single part using harmonic and inharmonic models
abstract
This paper presents a new music manipulation method that can change the timbre and phrases of an existing instrumental performance in a polyphonic sound mixture. This method consists of three primitive functions: 1) extracting and analyzing of a single instrumental part from polyphonic music signals, 2) mixing the instrument timbre with another, and 3) rendering a new phrase expression for another given score. The resulting customized part is re-mixed with the remaining parts of the original performance to generate new polyphonic music signals. A single instrumental part is extracted by using an integrated tone model that consists of harmonic and inharmonic tone models with the aid of the score of the single instrumental part. The extraction incorporates a residual model for the single instrumental part in order to avoid crosstalk between instrumental parts. The extracted model parameters are classified into their averages and deviations. The former is treated as instrument timbre and is customized by mixing, while the latter is treated as phrase expression and is customized by rendering. We evaluated our method in three experiments. The first experiment focused on introduction of the residual model, and it showed that the model parameters are estimated more accurately by 35.0 points. The second focused on timbral customization, and it showed that our method is more robust by 42.9 points in spectral distance compared with a conventional sound analysis-synthesis method, STRAIGHT. The third focused on the acoustic fidelity of customizing performance, and it showed that rendering phrase expression according to the note sequence leads to more accurate performance by 9.2 points in spectral distance in comparison with a rendering method that ignores the note sequence.
Naoki Yasuraoka, Takehiro Abe, Katsutoshi Itoyama, Toru Takahashi 0001, Tetsuya Ogata, Hiroshi G. Okuno
ACM Multimedia6
2009 Ranking Help Message Candidates Based on Robust Grammar Verification Results and Utterance History in Spoken Dialogue Systems
Kazunori Komatani, Satoshi Ikeda, Yuichiro Fukubayashi, Tetsuya Ogata, Hiroshi G. Okuno
SIGDIAL Conference5
2009 A Model of Temporally Changing User Behaviors in a Deployed Spoken Dialogue System
Kazunori Komatani, Tatsuya Kawahara, Hiroshi G. Okuno
UMAP3
2008 Two-channel-based voice activity detection for humanoid robots in noisy home environments
abstract
The purpose of this research is to accurately classify the speech signals originating from the front even in noisy home environments. This ability can help robots to improve speech recognition and to spot keywords. We therefore developed a new voice activity detection (VAD) based on the complex spectrum circle centroid (CSCC) method. It can classify the speech signals that are received at the front of two microphones by comparing the spectral energy of observed signals with that of target signals estimated by CSCC. Also, it can work in real time without training filter coefficients beforehand even in noisy environments (SNR ≫ 0 dB) and can cope with speech noises generated by audio-visual equipments such as televisions and audio devices. Since the CSCC method requires the directions of the noise signals, we also developed a sound source localization system integrated with cross-power spectrum phase (CSP) analysis and an expectation-maximization (EM) algorithm. This system was demonstrated to enable a robot to cope with multiple sound sources using two microphones.
Hyun-Don Kim, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ICRA4
2008 A robot referee for rock-paper-scissors sound games
abstract
This paper describes a robot referee for “rockpaper-scissors (RPS)” sound games; the robot decides the winner from a combination of rock, paper and scissors uttered by two or three people simultaneously without using any visual information. In this referee task, the robot has to cope with speech with low signal-to-noise ratio (SNR) due to a mixture of speeches, robot motor noises, and ambient noises. Our robot referee system, thus, consists of two subsystems - a real-time robot audition subsystem and a dialog subsystem focusing on RPS sound games. The robot audition subsystem can recognize simultaneous speeches by exploiting two key ideas; preprocessing consisting of sound source localization and separation with a microphone array, and system integration based on missing feature theory (MFT). Preprocessing improves the SNR of a target sound signal using geometric source separation with a multi-channel post-filter. MFT uses only reliable acoustic features in speech recognition and masks out unreliable parts caused by interfering sounds and preprocessing. MFT thus provides smooth integration between preprocessing and automatic speech recognition. The dialog subsystem is implemented as a system-initiative dialog system for multiple players based on deterministic finite automata. It first waits for a trigger command to start an RPS sound game, controls the dialog with players in the game, and finally decides the winner of the game. The referee system is constructed for Honda ASIMO with an 8-ch microphone array. In the case with two players, we attained a 70% task completion rate for the games on average.
Kazuhiro Nakadai, Shun'ichi Yamamoto, Hiroshi G. Okuno, Hirofumi Nakajima, Yuji Hasegawa, Hiroshi Tsujino
ICRA3
2008 Object dynamics prediction and motion generation based on reliable predictability
abstract
Consistency of object dynamics, which is related to reliable predictability, is an important factor for generating object manipulation motions. This paper proposes a technique to generate autonomous motions based on consistency of object dynamics. The technique resolves two issues: construction of an object dynamics prediction model and evaluation of consistency. The authors utilize Recurrent Neural Network with Parametric Bias to self-organize the dynamics, and link static images to the self-organized dynamics using a hierarchical neural network to deal with the first issue. For evaluation of consistency, the authors have set an evaluation function based on object dynamics relative to robot motor dynamics. Experiments have shown that the method is capable of predicting 90% of unknown object dynamics. Motion generation experiments have proved that the technique is capable of generating autonomous pushing motions that generate consistent rolling motions.
Shun Nishide, Tetsuya Ogata, Ryunosuke Yokoya, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno
ICRA6
2008 Integrating Topic Estimation and Dialogue History for Domain Selection in Multi-domain Spoken Dialogue Systems
Satoshi Ikeda, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IEA/AIE4
2008 Rapid Prototyping of Robust Language Understanding Modules for Spoken Dialogue Systems
Yuichiro Fukubayashi, Kazunori Komatani, Mikio Nakano, Kotaro Funakoshi, Hiroshi Tsujino, Tetsuya Ogata, Hiroshi G. Okuno
IJCNLP7
2008 Extensibility verification of robust domain selection against out-of-grammar utterances in multi-domain spoken dialogue system
abstract
We developed a robust domain selection method and verified its extensibility. An issue in domain selection is its robustness against out-of-grammar utterances. It is essential to generate correct system responses because such utterances often cause domain selection errors. We therefore integrated the topic estimation results and the dialogue history to construct a robust domain classifier. Another issue is that domain selection should be performed within an extensible framework, because the system is often modified and extended. That is, the classifier should still have high performance without reconstructing it after adding new domains. The extensibility of our method was not experimentally verified yet, because it requires a lot of effort to collect new dialogue data after extending the system. Therefore, we verified extensibility without collecting new data. We constructed the classifier by leaving out some domains in the dialogue data and then evaluated its accuracy as the classifier for the data where the left-out domains were virtually added. Index Terms: multi-domain spoken dialogue system, domain selection, out-of-grammar utterance
Satoshi Ikeda, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
INTERSPEECH4
2008 Expanding vocabulary for recognizing user's abbreviations of proper nouns without increasing ASR error rates in spoken dialogue systems
abstract
Users often abbreviate long words when using spoken dialogue systems, which results in automatic speech recognition (ASR) errors. We define abbreviated words as sub-words of the original word, and add them into an ASR dictionary. The first problem is that proper nouns cannot be correctly segmented by general morphological analyzers, although long and compounded words need to be segmented in agglutinative languages such as Japanese. The second is that, as vocabulary increases, adding many abbreviated words degrades the ASR accuracy. We develop two methods, (1) to segment words by using conjunction probabilities between characters, and (2) to manipulate occurrence probabilities of generated abbreviated words on the basis of the phonological similarities between abbreviated and original words. By our method, the ASR accuracy is improved by 24.2 points for utterances containing abbreviated words, and degraded by only a 0.1 point for those containing original words. Index Terms: spoken dialogue systems, abbreviated words, proper nouns, vocabulary expansion
Masaki Katsumaru, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
INTERSPEECH4
2008 Predicting ASR errors by exploiting barge-in rate of individual users for spoken dialogue systems
abstract
We exploit the barge-in rate of individual users to predict automatic speech recognition (ASR) errors. A barge-in is a situation in which a user starts speaking during a system prompt, and it can be detected even when ASR results are not reliable. Such features not using ASR results can be a clue for managing a situation in which user utterances cannot be successfully recognized. Since individual users in our system can be identified by their phone numbers, we accumulate how often each user barges in and use this rate as a user profile for determining whether a current “barge-in ” utterance should be accepted or not. We furthermore set a window that reflects the temporal transition of the user’s behavior as they get accustomed to the system. Experimental results show that setting the window improves the prediction accuracy of whether the utterance should be accepted or not. The experiments also clarify the minimum window width for improving accuracy. Index Terms: spoken dialogue system, user modeling, barge-in 1.
Kazunori Komatani, Tatsuya Kawahara, Hiroshi G. Okuno
INTERSPEECH3
2008 Soft missing-feature mask generation for simultaneous speech recognition system in robots
Toru Takahashi 0001, Shun'ichi Yamamoto, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
INTERSPEECH6
2008 Segmenting acoustic signal with articulatory movement using Recurrent Neural Network for phoneme acquisition
abstract
This paper proposes a computational model for phoneme acquisition by infants. Human infants perceive speech sounds not as discrete phoneme sequences but as continuous acoustic signals. One of critical problems in phoneme acquisition is the design for segmenting these continuous speech sounds. The key idea to solve this problem is that articulatory mechanisms such as the vocal tract help human beings to perceive speech sound units corresponding to phonemes. That is, the ability to distinguish phonemes is learned by recognizing unstable points in the dynamics of continuous sound with articulatory movement. We have developed a vocal imitation system embodying the relationship between articulatory movements and sounds produced by the movements. To segment acoustic signal with articulatory movement, we apply the segmenting method to our system by recurrent neural network with parametric bias (RNNPB). This method determines the multiple segmentation boundaries in a temporal sequence using the prediction error of the RNNPB model, and the PB values obtained by the method can be encoded as kind of phonemes. Our system was implemented by using a physical vocal tract model, called the Maeda model. Experimental results demonstrated that our system can self-organize the same phonemes in different continuous sounds. This suggests that our model reflects the process of phoneme acquisition.
Hisashi Kanda, Tetsuya Ogata, Kazunori Komatani, Hiroshi G. Okuno
IROS4
2008 Target speech detection and separation for humanoid robots in sparse dialogue with noisy home environments
abstract
In normal human communication, people face the speaker when listening and usually pay attention to the speaker’ face. Therefore, in robot audition, the recognition of the front talker is critical for smooth interactions. This paper presents an enhanced speech detection method for a humanoid robot that can separate and recognize speech signals originating from the front even in noisy home environments. The robot audition system consists of a new type of voice activity detection (VAD) based on the complex spectrum circle centroid (CSCC) method and a maximum signal-to-noise (Max-SNR) beamformer. This VAD based on CSCC can classify speech signals that are retrieved at the frontal region of two microphones embedded on the robot. The system works in real-time without needing training filter coefficients given in advance even in a noisy environment (SNR ≫ 0 dB). It can cope with speech noise generated from televisions and audio devices that does not originate from the center. Experiments using a humanoid robot, SIG2, with two microphones showed that our system enhanced extracted target speech signals more than 12 dB (SNR) and the success rate of automatic speech recognition for Japanese words was increased about 17 points.
Hyun-Don Kim, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS5
2008 Design and evaluation of two-channel-based sound source localization over entire azimuth range for moving talkers
abstract
We propose a way to evaluate various sound localization systems for moving sounds under the same conditions. To construct a database for moving sounds, we developed a moving sound creation tool using the API library developed by the ARINIS Company. We developed a two-channel-based sound source localization system integrated with a cross-power spectrum phase (CSP) analysis and EM algorithm. The CSP of sound signals obtained with only two microphones is used to localize the sound source without having to use prior information such as impulse response data. The EM algorithm helps the system cope with several moving sound sources and reduce localization error. We evaluated our sound localization method using artificial moving sounds and confirmed that it can well localize moving sounds slower than 1.125 rad/sec. Finally, we solve the problem of distinguishing whether sounds are coming from the front or back by rotating a robotpsilas head equipped with only two microphones. Our system was applied to a humanoid robot called SIG2, and we confirmed its ability to localize sounds over the entire azimuth range.
Hyun-Don Kim, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS4
2008 A robot listens to music and counts its beats aloud by separating music from counting voice
abstract
This paper presents a beat-counting robot that can count musical beats aloud, i.e., speak ldquoone, two, three, four, one, two, ...rdquo along music, while listening to music by using its own ears. Music-understanding robots that interact with humans should be able not only to recognize music internally, but also to express their own internal states. To develop our beat-counting robot, we have tackled three issues: (1) recognition of hierarchical beat structures, (2) expression of these structures by counting beats, and (3) suppression of counting voice (self-generated sound) in sound mixtures recorded by ears. The main issue is (3) because the interference of counting voice in music causes the decrease of the beat recognition accuracy. So we designed the architecture for music-understanding robot that is capable of dealing with the issue of self-generated sounds. To solve these issues, we took the following approaches: (1) beat structure prediction based on musical knowledge on chords and drums, (2) speed control of counting voice according to music tempo via a vocoder called STRAIGHT, and (3) semi-blind separation of sound mixtures into music and counting voice via an adaptive filter based on ICA (independent component analysis) that uses the waveform of the counting voice as a prior knowledge. Experimental result showed that suppressing robotpsilas own voice improved music recognition capability.
Takeshi Mizumoto, Ryu Takeda, Kazuyoshi Yoshii, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS6
2008 A robot uses its own microphone to synchronize its steps to musical beats while scatting and singing
abstract
Musical beat tracking is one of the effective technologies for human-robot interaction such as musical sessions. Since such interaction should be performed in various environments in a natural way, musical beat tracking for a robot should cope with noise sources such as environmental noise, its own motor noises, and self voices, by using its own microphone. This paper addresses a musical beat tracking robot which can step, scat and sing according to musical beats by using its own microphone. To realize such a robot, we propose a robust beat tracking method by introducing two key techniques, that is, spectro-temporal pattern matching and echo cancellation. The former realizes robust tempo estimation with a shorter window length, thus, it can quickly adapt to tempo changes. The latter is effective to cancel self noises such as stepping, scatting, and singing. We implemented the proposed beat tracking method for Honda ASIMO. Experimental results showed ten times faster adaptation to tempo changes and high robustness in beat tracking for stepping, scatting and singing noises. We also demonstrated the robot times its steps while scatting or singing to musical beats.
Kazumasa Murata, Kazuhiro Nakadai, Kazuyoshi Yoshii, Ryu Takeda, Toyotaka Torii, Hiroshi G. Okuno, Yuji Hasegawa, Hiroshi Tsujino
IROS6
2008 Active sensing based dynamical object feature extraction
abstract
This paper presents a method to autonomously extract object features that describe their dynamics from active sensing experiences. The model is composed of a dynamics learning module and a feature extraction module. Recurrent Neural Network with Parametric Bias (RNNPB) is utilized for the dynamics learning module, learning and self-organizing the sequences of robot and object motions. A hierarchical neural network is linked to the input of RNNPB as the feature extraction module for extracting object features that describe the object motions. The two modules are simultaneously trained using image and motion sequences acquired from the robotpsilas active sensing with objects. Experiments are performed with the robotpsilas pushing motion with a variety of objects to generate sliding, falling over, bouncing, and rolling motions. The results have shown that the model is capable of extracting features that distinguish the characteristics of object dynamics.
Shun Nishide, Tetsuya Ogata, Ryunosuke Yokoya, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno
IROS6
2008 Barge-in-able robot audition based on ICA and missing feature theory under semi-blind situation
abstract
This paper describes a robot audition system that allows the user to barge-in; that is, the user can speak simultaneously when the robot is speaking. Our ldquobarge-in-ablerdquo system consists of two stages: (1) cancellation of robot speech and (2) recognition of the separated user speech under the ldquosemi-blind situationrdquo. The semi-blind situation is where a robotpsilas speech signal is known but a userpsilas speech signal is not. The first stage is achieved by using an adaptive filter based on time-frequency domain Independent Component Analysis, because that can separate robot speech more robustly against noise than conventional echo cancellers. To improve performance in online processing, we utilized known source normalization and the exponentially weighted stepsize method. The second stage is achieved by automatic speech recognition (ASR) based on the missing feature theory which provides robust recognition by exploiting the reliability of speech features distorted due to noise and/or separation. The semi-blind situation simplifies the estimation of such reliabilities. Experiments demonstrated that our system improved word correctness of ASR by 10.0%.
Ryu Takeda, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS5
2008 Design and Implementation of 3D Auditory Scene Visualizer towards Auditory Awareness with Face Tracking
abstract
If machine audition can recognize an auditory scene containing simultaneous and moving talkers, what kinds of awareness will people gain from an auditory scene visualizer? This paper presents the design and implementation of 3D Auditory Scene Visualizer based on the visual information seeking mantra, i.e., ldquooverview first, zoom and filter, then details on demandrdquo. The machine audition system called HARK captures 3D sounds with a microphone array, localizes and separates sounds, and recognizes separated sounds by automatic speech recognition (ASR). The 3D visualizer implemented in Java 3D displays each sound stream as a beam originating from the center of the microphones (overview mode), shows temporal snapshots with/without specifying focusing areas (zoom and filter mode), and shows detailed information about a particular sound stream (details on demand). In the details-ondemand mode, ASR results are displayed in a ldquokaraokerdquo manner, i.e., character-by-character. This three-mode visualization will give the user auditory awareness enhanced by HARK. In addition, a face-tracking system automatically changes the focus of attention by tracking the userpsilas face. The resulting system is portable and can be deployed in any place, so it is expected to give more vivid awareness than expensive high-fidelity auditory scene reproduction systems.
Yuji Kubota, Masatoshi Yoshida, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ISM5
2008 SalienceGraph: Visualizing Salience Dynamics of Written Discourse by Using Reference Probability and PLSA
Shun Shiramatsu, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
PRICAI4
2008 Managing out-of-grammar utterances by topic estimation with domain extensibility in multi-domain spoken dialogue systems
Kazunori Komatani, Satoshi Ikeda, Tetsuya Ogata, Hiroshi G. Okuno
Speech Commun.4
2008 An Efficient Hybrid Music Recommender System Using an Incrementally Trainable Probabilistic Generative Model
abstract
This paper presents a hybrid music recommender system that ranks musical pieces while efficiently maintaining collaborative and content-based data, i.e., rating scores given by users and acoustic features of audio signals. This hybrid approach overcomes the conventional tradeoff between recommendation accuracy and variety of recommended artists. Collaborative filtering, which is used on e-commerce sites, cannot recommend nonbrated pieces and provides a narrow variety of artists. Content-based filtering does not have satisfactory accuracy because it is based on the heuristics that the user's favorite pieces will have similar musical content despite there being exceptions. To attain a higher recommendation accuracy along with a wider variety of artists, we use a probabilistic generative model that unifies the collaborative and content-based data in a principled way. This model can explain the generative mechanism of the observed data in the probability theory. The probability distribution over users, pieces, and features is decomposed into three conditionally independent ones by introducing latent variables. This decomposition enables us to efficiently and incrementally adapt the model for increasing numbers of users and rating scores. We evaluated our system by using audio signals of commercial CDs and their corresponding rating scores obtained from an e-commerce site. The results revealed that our system accurately recommended pieces including nonrated ones from a wide variety of artists and maintained a high degree of accuracy even when new users and rating scores were added.
Kazuyoshi Yoshii, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IEEE Trans. Speech Audio Process.5
2007 Design and implementation of a robot audition system for automatic speech recognition of simultaneous speech
abstract
This paper addresses robot audition that can cope with speech that has a low signal-to-noise ratio (SNR) in real time by using robot-embedded microphones. To cope with such a noise, we exploited two key ideas; Preprocessing consisting of sound source localization and separation with a microphone array, and system integration based on missing feature theory (MFT). Preprocessing improves the SNR of a target sound signal using geometric source separation with multichannel post-filter. MFT uses only reliable acoustic features in speech recognition and masks unreliable parts caused by errors in preprocessing. MFT thus provides smooth integration between preprocessing and automatic speech recognition. A real-time robot audition system based on these two key ideas is constructed for Honda ASIMO and Humanoid SIG2 with 8-ch microphone arrays. The paper also reports the improvement of ASR performance by using two and three simultaneous speech signals.
Shun'ichi Yamamoto, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ASRU8
2007 Integration and Adaptation of Harmonic and Inharmonic Models for Separating Polyphonic Musical Signals
abstract
This paper describes a sound source separation method for polyphonic sound mixtures of music to build an instrument equalizer for remixing multiple tracks separated from compact-disc recordings by changing the volume level of each track. Although such mixtures usually include both harmonic and inharmonic sounds, the difficulties in dealing with both types of sounds together have not been addressed in most previous methods that have focused on either of the two types separately. We therefore developed an integrated weighted-mixture model consisting of both harmonic-structure and inharmonic-structure tone models (generative models for the power spectrogram). On the basis of the MAP estimation using the EM algorithm, we estimated all model parameters of this integrated model under several original constraints for preventing over-training and maintaining intra-instrument consistency. Using standard MIDI files as prior information of the model parameters, we applied this model to compact-disc recordings and achieved the instrument equalizer.
Katsutoshi Itoyama, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ICASSP (1)5
2007 Vowel Imitation Using Vocal Tract Model and Recurrent Neural Network
Hisashi Kanda, Tetsuya Ogata, Kazunori Komatani, Hiroshi G. Okuno
ICONIP (2)4
2007 Predicting Object Dynamics from Visual Images through Active Sensing Experiences
abstract
Prediction of dynamic features is an important task for determining the manipulation strategies of an object. This paper presents a technique for predicting dynamics of objects relative to the robot's motion from visual images. During the learning phase, the authors use recurrent neural network with parametric bias (RNNPB) to self-organize the dynamics of objects manipulated by the robot into the PB space. The acquired PB values, static images of objects, and robot motor values are input into a hierarchical neural network to link the static images to dynamic features (PB values). The neural network extracts prominent features that induce each object dynamics. For prediction of the motion sequence of an unknown object, the static image of the object and robot motor value are input into the neural network to calculate the PB values. By inputting the PB values into the closed loop RNNPB, the predicted movements of the object relative to the robot motion are calculated sequentially. Experiments were conducted with the humanoid robot Robovie-IIs pushing objects at different heights. Reducted grayscale images and shoulder pitch angles were input into the neural network to predict the dynamics of target objects. The results of the experiment proved that the technique is efficient for predicting the dynamics of the objects.
Shun Nishide, Tetsuya Ogata, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno
ICRA5
2007 Distance Estimation of Hidden Objects Based on Acoustical Holography by applying Acoustic Diffraction of Audible Sound
abstract
Occlusion is a problem for range finders; ranging systems using cameras or lasers cannot be used to estimate distance to an object (hidden object) that is occluded by another (obstacle). We developed a method to estimate the distance to the hidden object by applying acoustic diffraction of audible sound. Our method is based on time-of-flight (TOF), which has been used in ultrasound ranging systems. We determined the best frequency of audible sound and designed its optimal modulated signal for our system. We determined that the system estimates the distance to the hidden object as well as the obstacle. However, the measurement signal obtained from the hidden object was weak. Thus, interference from sound signals reflected from other objects or walls was not negligible. Therefore, we combined acoustical holography (AH) and TOF, which enabled a partial analysis of the reflection sound intensity field around the obstacle and hidden object. Our method was effective for ranging two objects of the same size within a 1.2 m depth range. The accuracy of our method was 3 cm for the obstacle, and 6 cm for the hidden object.
Haruhiko Niwa, Tetsuya Ogata, Kazunori Komatani, Hiroshi G. Okuno
ICRA4
2007 Human-Robot Cooperation using Quasi-symbols Generated by RNNPB Model
abstract
We describe a means of human robot interaction based not on natural language but on "quasi symbols," which represent sensory-motor dynamics in the task and/or environment. It thus overcomes a key problem of using natural language for human-robot interaction - the need to understand the dynamic context. The quasi-symbols used are motion primitives corresponding to the attractor dynamics of the sensory-motor flow. These primitives are extracted from the observed data using the recurrent neural network with parametric bias (RNNPB) model. Binary representations based on the model parameters were implemented as quasi symbols in a humanoid robot, Robovie. The experiment task was robot-arm operation on a table. The quasi-symbols acquired by learning enabled the robot to perform novel motions. A person was able to control the arm through speech interaction using these quasi-symbols. These quasi symbols formed a hierarchical structure corresponding to the number of nodes in the model. The meaning of some of the quasi-symbols depended on the context, indicating that they are useful for human-robot interaction.
Tetsuya Ogata, Shohei Matsumoto, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno
ICRA5
2007 Real-Time Auditory and Visual Talker Tracking Through Integrating EM Algorithm and Particle Filter
Hyun-Don Kim, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IEA/AIE4
2007 Evaluation of Two Simultaneous Continuous Speech Recognition with ICA BSS and MFT-Based ASR
Ryu Takeda, Shun'ichi Yamamoto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IEA/AIE5
2007 Topic estimation with domain extensibility for guiding user's out-of-grammar utterances in multi-domain spoken dialogue systems
abstract
In a multi-domain spoken dialogue system, a user’s utterances are more prone to be out-of-grammar, because this kind of system deals with more tasks than a single-domain system. We defined a topic as a domain about which users want to find more information, and we developed a method of recovering out-ofgrammar utterances based on topic estimation, i.e., by providing a help message in the estimated domain. Moreover, the domain extensibility, that is, to facilitate adding new domains, should be inherently retained in multi-domain systems. We therefore collected documents from the Web as training data for topic estimation. Because the data contained not a few noises, we used Latent Semantic Mapping (LSM), which enables robust topic estimation by removing the effect of noise from the data. The experimental results based on using 272 utterances collected with a Woz-like method showed that our method increased the topic estimation accuracy by 23.1 points from the baseline. Index Terms: multi-domain spoken dialogue system, topic estimation, out-of-grammar utterance 1.
Satoshi Ikeda, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
INTERSPEECH4
2007 Analyzing temporal transition of real user's behaviors in a spoken dialogue system
abstract
Managing various behaviors of real users is indispensable for spoken dialogue systems to operate adequately in real environments. We have analyzed various users ’ behaviors using data collected over 34 months from the Kyoto City Bus Information System. We focused on “barge-in ” and added barge-in rates to our analysis. Temporal transitions of users ’ behaviors, such as automatic speech recognition (ASR) accuracy, task success rates and barge-in rates, were initially investigated. We then examined the relationship between ASR accuracy and barge-in rates. Analysis revealed that the ASR accuracy of utterances inputted with barge-ins was lower because many novices, who were not accustomed to the timing when to utter, used the system. We also observed that the ASR accuracy of utterances with barge-ins differed based on the barge-in rates of individual users. The results indicate that the barge-in rate can be used as a novel user profile for detecting ASR errors. Index Terms: spoken dialogue system, real user behavior, barge-in
Kazunori Komatani, Tatsuya Kawahara, Hiroshi G. Okuno
INTERSPEECH3
2007 Vocal imitation using physical vocal tract model
abstract
A vocal imitation system was developed using a computational model that supports the motor theory of speech perception. A critical problem in vocal imitation is how to generate speech sounds produced by adults, whose vocal tracts have physical properties (i.e., articulatory motions) differing from those of infants’ vocal tracts. To solve this problem, a model based on the motor theory of speech perception, was constructed. This model suggests that infants simulate the speech generation by estimating their own articulatory motions in order to interpret the speech sounds of adults. Applying this model enables the vocal imitation system to estimate articulatory motions for unexperienced speech sounds that have not actually been generated by the system. The system was implemented by using Recurrent Neural Network with Parametric Bias (RNNPB) and a physical vocal tract model, called the Maeda model. Experimental results demonstrated that the system was sufficiently robust with respect to individual differences in speech sounds and could imitate unexperienced vowel sounds.
Hisashi Kanda, Tetsuya Ogata, Kazunori Komatani, Hiroshi G. Okuno
IROS4
2007 Auditory and visual integration based localization and tracking of humans in daily-life environments
abstract
The purpose of this research is to develop techniques that enable robots to choose and track a desired person for interaction in daily-life environments. Therefore, localizing multiple moving sounds and human faces is necessary so that robots can locate a desired person. For sound source localization, we used a cross-power spectrum phase analysis (CSP) method and showed that CSP can localize sound sources only using two microphones and does not need impulse response data. An expectation-maximization (EM) algorithm was shown to enable a robot to cope with multiple moving sound sources. For face localization, we developed a method that can reliably detect several faces using the skin color classification obtained by using the EM algorithm. To deal with a change in color state according to illumination condition and various skin colors, the robot can obtain new skin color features of faces detected by OpenCV, an open vision library, for detecting human faces. Finally, we developed a probability based method to integrate auditory and visual information and to produce a reliable tracking path in real time. Furthermore, the developed system chose and tracked people while dealing with various background noises that are considered loud, even in the daily-life environments.
Hyun-Don Kim, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS4
2007 Two-way translation of compound sentences and arm motions by recurrent neural networks
abstract
We present a connectionist model that combines motions and language based on the behavioral experiences of a real robot. Two models of recurrent neural network with parametric bias (RNNPB) were trained using motion sequences and linguistic sequences. These sequences were combined using their respective parameters so that the robot could handle many-to-many relationships between motion sequences and linguistic sequences. Motion sequences were articulated into some primitives corresponding to given linguistic sequences using the prediction error of the RNNPB model. The experimental task in which a humanoid robot moved its arm on a table demonstrated that the robot could generate a motion sequence corresponding to given linguistic sequence even if the motions or sequences were not included in the training data, and vice versa.
Tetsuya Ogata, Masamitsu Murase, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno
IROS5
2007 Exploiting known sound source signals to improve ICA-based robot audition in speech separation and recognition
abstract
This paper describes a new semi-blind source separation (semi-BSS) technique with independent component analysis (ICA) for enhancing a target source of interest and for suppressing other known interference sources. The semi BSS technique is necessary for double-talk free robot audition systems in order to utilize known sound source signals such as self speech, music, or TV-sound, through a line-in or ubiquitous network. Unlike the conventional semi-BSS with ICA, we use the time-frequency domain convolution model to describe the reflection of the sound and a new mixing process of sounds for ICA. In other words, we consider that reflected sounds during some delay time are different from the original. ICA then separates the reflections as other interference sources. The model enables us to eliminate the frame size limitations of the frequency-domain ICA, and ICA can separate the known sources under a highly reverberative environment. Experimental results show that our method outperformed the conventional semi-BSS using ICA under simulated normal and highly reverberative environments.
Ryu Takeda, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS5
2007 Discovery of other individuals by projecting a self-model through imitation
abstract
This paper proposes a novel model which enables a humanoid robot infant to discover other individual (e.g. human parent). In this work, the authors define “other individual” as an actor which can be predicted by a self-model. For modeling the developmental process of discovering ability, the following three approaches are employed. (i) Projection of a selfmodel for predicting other individual’s actions. (ii) Mediation by a physical object between self and other individual. (iii) Introduction of infant imitation by parent. For creating the self-model of a robot, we apply Recurrent Neural Network with Parametric Bias (RNNPB) model which can learn the robot’s body dynamics. For the other-model of a human, conventional hierarchical neural networks are attached to the RNNPB model as “conversion modules”. Our target task is a moving an object. For evaluation of our model, human discovery experiments by the robot projecting its self-model were conducted. The results demonstrated that our method enabled the robot to predict the human’s motions, and to estimate the human’s position fairly accurately, which proved its adequacy.
Ryunosuke Yokoya, Tetsuya Ogata, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno
IROS5
2007 A biped robot that keeps steps in time with musical beats while listening to music with its own ears
abstract
We aim at enabling a biped robot to interact with humans through real-world music in daily-life environments, e.g., to autonomously keep its steps (stamps) in time with musical beats. To achieve this, the robot should be able to robustly predict the beat times in real time while listening to musical performance with its own ears (head-embedded microphones). However, this has not previously been addressed in most studies on music-synchronized robots due to the difficulty in predicting the beat times in real-world music. To solve this problem, we implemented a beat-tracking method developed in the field of music information processing. The predicted beat times are then used by a feedback-control method that adjusts the robot's step intervals to synchronize its steps in time with the beats. The experimental results show that the robot can adjust its steps in time with the beat times as the tempo changes. The resulting robot needed about 25 [s] to recognize the tempo change after it and then synchronize its steps.
Kazuyoshi Yoshii, Kazuhiro Nakadai, Toyotaka Torii, Yuji Hasegawa, Hiroshi Tsujino, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS8
2007 Auditory and Visual Integration based Localization and Tracking of Multiple Moving Sounds in Daily-life Environments
abstract
This paper presents techniques that enable talker tracking for effective human-robot interaction. To track moving people in daily-life environments, localizing multiple moving sounds is necessary so that robots can locate talkers. However, the conventional method requires an array of microphones and impulse response data. Therefore, we propose a way to integrate a cross-power spectrum phase analysis (CSP) method and an expectation-maximization (EM) algorithm. The CSP can localize sound sources using only two microphones and does not need impulse response data. Moreover, the EM algorithm increases the system's effectiveness and allows it to cope with multiple sound sources. We confirmed that the proposed method performs better than the conventional method. In addition, we added a particle filter to the tracking process to produce a reliable tracking path and the particle filter is able to integrate audio-visual information effectively. Furthermore, the applied particle filter is able to track people while dealing with various noises that are even loud sounds in the daily-life environments.
Hyun-Don Kim, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
RO-MAN4
2007 Drum Sound Recognition for Polyphonic Audio Signals by Adaptation and Matching of Spectrogram Templates With Harmonic Structure Suppression
abstract
This paper describes a system that detects onsets of the bass drum, snare drum, and hi-hat cymbals in polyphonic audio signals of popular songs. Our system is based on a template-matching method that uses power spectrograms of drum sounds as templates. This method calculates the distance between a template and each spectrogram segment extracted from a song spectrogram, using Goto's distance measure originally designed to detect the onsets in drums-only signals. However, there are two main problems. The first problem is that appropriate templates are unknown for each song. The second problem is that it is more difficult to detect drum-sound onsets in sound mixtures including various sounds other than drum sounds. To solve these problems, we propose template-adaptation and harmonic-structure-suppression methods. First of all, an initial template of each drum sound, called a seed template, is prepared. The former method adapts it to actual drum-sound spectrograms appearing in the song spectrogram. To make our system robust to the overlapping of harmonic sounds with drum sounds, the latter method suppresses harmonic components in the song spectrogram before the adaptation and matching. Experimental results with 70 popular songs showed that our template-adaptation and harmonic-structure-suppression methods improved the recognition accuracy and achieved 83%, 58%, and 46% in detecting onsets of the bass drum, snare drum, and hi-hat cymbals, respectively.
Kazuyoshi Yoshii, Masataka Goto, Hiroshi G. Okuno
IEEE Trans. Speech Audio Process.3
2007 Robust Recognition of Simultaneous Speech by a Mobile Robot
abstract
This paper describes a system that gives a mobile robot the ability to perform automatic speech recognition with simultaneous speakers. A microphone array is used along with a real-time implementation of geometric source separation (GSS) and a postfilter that gives a further reduction of interference from other sources. The postfllter is also used to estimate the reliability of spectral features and compute a missing feature mask. The mask is used in a missing feature theory-based speech recognition system to recognize the speech from simultaneous Japanese speakers in the context of a humanoid robot. Recognition rates are presented for three simultaneous speakers located at 2 m from the robot. The system was evaluated on a 200-word vocabulary at different azimuths between sources, ranging from 10deg to 90deg. Compared to the use of the microphone array source separation alone, we demonstrate an average reduction in relative recognition error rate of 24% with the postfllter and of 42% when the missing features approach is combined with the postfllter. We demonstrate the effectiveness of our multisource microphone array postfilter and the improvement it provides when used in conjunction with the missing features theory.
Jean-Marc Valin, Seiichi Yamamoto, Jean Rouat, François Michaud, Kazuhiro Nakadai, Hiroshi G. Okuno
IEEE Trans. Robotics6
2006 F0 Estimation Method for Singing Voice in Polyphonic Audio Signal Based on Statistical Vocal Model and Viterbi Search
abstract
This paper describes a method for estimating F0s of vocal from polyphonic audio signals. Because melody is sung by a singer in many musical pieces, the estimation of F0s of the vocal part is useful for many applications. Based on existing multiple-F0 estimation method, we evaluate the vocal probabilities of the harmonic structure of each F0 candidate. In order to calculate the vocal probabilities of the harmonic structure, we extract and resynthesize the harmonic structure by using a sinusoidal model and extract feature vectors. Then, we evaluate the vocal probability by using vocal and non-vocal Gaussian mixture models (GMMs). Finally, we track F0 trajectories using these probabilities based on Viterbi search. Experimental results show that our method improves estimation accuracy from 78.1% to 84.3%, which is 28.3% reduction of misestimation
Hiromasa Fujihara, Tetsuro Kitahara, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ICASSP (5)6
2006 Instrogram: A New Musical Instrument Recognition Technique Without Using Onset Detection NOR F0 Estimation
abstract
This paper describes a new technique for recognizing musical instruments in polyphonic music. Because the conventional framework for musical instrument recognition in polyphonic music had to estimate the onset time and fundamental frequency (F0) of each note, instrument recognition strictly suffered from errors of onset detection and F0 estimation. Unlike such a note-based processing framework, our technique calculates the temporal trajectory of instrument existence probabilities for every possible F0, and the results are visualized with a spectrogram-like graphical representation called instrogram. The instrument existence probability is defined as the product of a nonspecific instrument existence probability calculated using PreFEst and a conditional instrument existence probability calculated using the hidden Markov model. Experimental results show that the obtained instrograms reflect the actual instrumentations and facilitate instrument recognition
Tetsuro Kitahara, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ICASSP (5)5
2006 Robust Tracking of Multiple Sound Sources by Spatial Integration of Room And Robot Microphone Arrays
abstract
Sound source tracking is an important function for a robot operating in a daily environment, because the robot should recognize where a sound event such as speech, music and other environmental sounds originates from. This paper addresses sound source tracking by integrating a room and a robot microphone array. The room microphone array consists of 64 microphones attached to the walls. It provides 2D (x-y) sound source localization based on a weighted delay-and-sum beamforming method. The robot microphone array consists of eight microphones installed on a robot head, and localizes multiple sound sources in azimuth. The localization results are integrated to track sound sources by using a particle filter for multiple sound sources. The experimental results show that particle filter based integration reduces localization errors and provides accurate and robust 2D sound source tracking.
Kazuhiro Nakadai, Hirofumi Nakajima, Masamitsu Murase, Satoshi Kaijiri, Kentaro Yamada, Takahiro Nakamura, Yuji Hasegawa, Hiroshi G. Okuno, Hiroshi Tsujino
ICASSP (4)8
2006 An Error Correction Framework Based on Drum Pattern Periodicity for Improving Drum Sound Detection
abstract
This paper presents a framework for correcting errors of automatic drum sound detection focusing on the periodicity of drum patterns. We define drum patterns as periodic structures found in onset sequences of bass and snare drum sounds. Our framework extracts periodic drum patterns from imperfect onset sequences of detected drum sounds (bottom-up processing) and corrects errors using the periodicity of the drum patterns (top-down processing). We implemented this framework on our drum-sound detection system. We first obtained onset sequences of the drum sounds with our system and extracted drum patterns. On the basis of our observation that the same drum patterns tend to be repeated, we detected time points which deviate from the periodicity as error candidates. Finally, we verified each error candidate to judge whether it is an actual onset or not. Experiments of drum sound detection for polyphonic audio signals of popular CD recordings showed that our correction framework improved the average detection accuracy from 77.4% to 80.7%
Kazuyoshi Yoshii, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ICASSP (5)5
2006 Genetic Algorithm-Based Improvement of Robot Hearing Capabilities in Separating and Recognizing Simultaneous Speech Signals
Shun'ichi Yamamoto, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Ryu Takeda, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IEA/AIE9
2006 Speaker identification under noisy environments by using harmonic structure extraction and reliable frame weighting
abstract
We present methods for automatic speaker identification in noisy environments. To improve noise robustness of speaker identification, we developed two methods, theharmonic structure extraction method and the reliable frame weighting method. The harmonic structure extraction method enables the speaker of input speech signals to be identified after environmental noise has been reduced. This method first extracts harmonic components of the speech from the sound mixtures and then resynthesizes a clean speech signal by using a sinusoidal model driven by harmonic components. The reliable frame weighting method then determines how each frame of the resynthesized speech is reliable (i.e. little influenced by environmental noises) by using two Gaussian mixture models for the speech and noise. The speaker can be robustly identified by attaching importance to reliable frames. Experimental results with thirty speakers showed that our method was able to reduce the influences of environmental noise and achieved an error rate of 10.7%, while the error rate for a conventional method was 18.9%. Index Terms: speaker identification, noise robustness, voice extraction, voice reliability, Gaussian mixture model.
Hiromasa Fujihara, Tetsuro Kitahara, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
INTERSPEECH6
2006 Dynamic help generation by estimating user²s mental model in spoken dialogue systems
abstract
In a speech interface, a gap between a user’s mental model and actual structures of systems tends to be large because the amount of information conveyed by speech is limited. We address dynamic help generation adapted to users, to decrease the gap between them. We defined a domain concept tree as an expression of a system’s actual structure. We estimated and maintained user’s knowledge about the system on the tree. Every node in the tree has values representing the degree to which a user understands the concepts corresponding to the nodes. The values are updated based on the content of user’s utterances and help messages the system gives. Help messages provided for users are determined by referring to the domain concept tree and identifying concepts the user does not understand. We evaluated our method by testing twelve novice subjects. Both the average time to complete tasks and the number of utterances significantly decreased because of the help messages provided by our method. Index Terms: spoken dialogue system, adaptive help generation, novice user
Yuichiro Fukubayashi, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
INTERSPEECH4
2006 Improving speech recognition of two simultaneous speech signals by integrating ICA BSS and automatic missing feature mask generation
Ryu Takeda, Shun'ichi Yamamoto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
INTERSPEECH5
2006 Real-Time Tracking of Multiple Sound Sources by Integration of In-Room and Robot-Embedded Microphone Arrays
abstract
Real-time and robust sound source tracking is an important function for a robot operating in a daily environment, because the robot should recognize where a sound event such as speech, music and other environmental sounds originate from. This paper addresses real-time sound source tracking by real-time integration of an in-room microphone array (IRMA) and a robot-embedded microphone array (REMA). The IRMA system consists of 64 ch microphones attached to the walls. It localizes multiple sound sources based on weighted delay-and-sum beam-forming on a 2D plane. The REMA system localizes multiple sound sources in azimuth using eight microphones attached to a robot's head on a rotational table. The localization results are integrated to track multiple sound sources by using a particle filter in real-time. The experimental results show that particle filter based integration improved accuracy and robustness in multiple sound source tracking even when the robot's head was in rotation
Kazuhiro Nakadai, Hirofumi Nakajima, Masamitsu Murase, Hiroshi G. Okuno, Yuji Hasegawa, Hiroshi Tsujino
IROS4
2006 Multiple Acoustical Holography Method for Localization of Objects in Broad Range using Audible Sound
abstract
This paper describes a new acoustic localization method using audible sound, which can be applied over a broader range of search directions. In the field of robotics, most conventional indoor localization systems based on sonar range finders use ultrasound to obtain a highly accurate distance. Because ultrasound has high directivity, many measurements are required to localize objects in a large space. To achieve localization with one-time measurement, we use audible sounds. We then calculate an intensity field of the reflection sound to estimate object positions. Although acoustical holography (AH) is a well-known technique to do this, it has problems in that it generates false images. We propose multiple AH (MAH) to solve this problem. The method is used to divide a measurement plane into sub-planes and to apply AH to each sub-plane. By integrating the results of applying AH to the sub-planes, false images can be suppressed because the positions of the false images differ depending on the position of the sub-plane. In addition, we use multiple frequencies to advance an accuracy of the localization based on MAH in a real environment. We constructed a localization system with only one speaker and a microphone array. In both simulation and actual experiments, we confirmed that MAH was effective method for the suppression of false image and could be used within the range of an angle view of 120 deg
Haruhiko Niwa, Tetsuya Ogata, Kazunori Komatani, Hiroshi G. Okuno
IROS4
2006 Missing-Feature based Speech Recognition for Two Simultaneous Speech Signals Separated by ICA with a pair of Humanoid Ears
abstract
Robot audition is a critical technology in making robots symbiosis with people. Since we hear a mixture of sounds in our daily lives, sound source localization and separation, and recognition of separated sounds are three essential capabilities. Sound source localization has been recently studied well for robots, while the other capabilities still need extensive studies. This paper reports the robot audition system with a pair of omni-directional microphones embedded in a humanoid to recognize two simultaneous talkers. It first separates sound sources by independent component analysis (ICA) with single-input multiple-output (SIMO) model. Then, spectral distortion for separated sounds is estimated to identify reliable and unreliable components of the spectrogram. This estimation generates the missing feature masks as spectrographic masks. These masks are then used to avoid influences caused by spectral distortion in automatic speech recognition based on missing-feature method. The novel ideas of our system reside in estimates of spectral distortion of temporal-frequency domain in terms of feature vectors. In addition, we point out that the voice-activity detection (VAD) is effective to overcome the weak point of ICA against the changing number of talkers. The resulting system outperformed the baseline robot audition system by 15%
Ryu Takeda, Shun'ichi Yamamoto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS5
2006 Real-Time Robot Audition System That Recognizes Simultaneous Speech in The Real World
abstract
This paper presents a robot audition system that recognizes simultaneous speech in the real world by using robot-embedded microphones. We have previously reported missing feature theory (MFT) based integration of sound source separation (SSS) and automatic speech recognition (ASR) for building robust robot audition. We demonstrated that a MFT-based prototype system drastically improved the performance of speech recognition even when three speakers talked to a robot simultaneously. However, the prototype system had three problems; being offline, hand-tuning of system parameters, and failure in voice activity detection (VAD). To attain online processing, we introduced FlowDesigner-based architecture to integrate sound source localization (SSL), SSS and ASR. This architecture brings fast processing and easy implementation because it provides a simple framework of shared-object-based integration. To optimize the parameters, we developed genetic algorithm (GA) based parameter optimization, because it is difficult to build an analytical optimization model for mutually dependent system parameters. To improve VAD, we integrated new VAD based on a power spectrum and location of a sound source into the system, since conventional VAD relying only on power often fails due to low signal-to-noise ratio of simultaneous speech. We, then, constructed a robot audition system for Honda ASIMO. As a result, we showed that the system worked online and fast, and had a better performance in robustness and accuracy through experiments on recognition of simultaneous speech in a noisy and echoic environment
Shun'ichi Yamamoto, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS8
2006 Experience Based Imitation Using RNNPB
abstract
Robot imitation is a useful and promising alternative to robot programming. Robot imitation involves two crucial issues. The first is how a robot can imitate a human whose physical structure and properties differ greatly from its own. The second is how the robot can generate various motions from finite programmable patterns (generalization). This paper describes a novel approach to robot imitation based on its own physical experiences. Let us consider a target task of moving an object on a table. For imitation, we focused on an active sensing process in which the robot acquires the relation between the object's motion and its own arm motion. For generalization, we applied a recurrent neural network with parametric bias (RNNPB) model to enable recognition/generation of imitation motions. The robot associates the arm motion which reproduces the observed object's motion presented by a human operator. Experimental results demonstrated that our method enabled the robot to imitate not only motion it has experienced but also unknown motion, which proved its capability for generalization
Ryunosuke Yokoya, Tetsuya Ogata, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno
IROS5
2006 Automatic Synchronization between Lyrics and Music CD Recordings Based on Viterbi Alignment of Segregated Vocal Signals
abstract
This paper describes a system that can automatically synchronize between polyphonic musical audio signals and corresponding lyrics. Although there were methods that can synchronize between monophonic speech signals and corresponding text transcriptions by using Viterbi alignment techniques, they cannot be applied to vocals in CD recordings because accompaniment sounds often overlap with vocals. To align lyrics with such vocals, we therefore developed three methods: a method for segregating vocals from polyphonic sound mixtures, a method for detecting vocal sections, and a method for adapting a speech-recognizer phone model to segregated vocal signals. Experimental results for 10 Japanese popular-music songs showed that our system can synchronize between music and lyrics with satisfactory accuracy for 8 songs
Hiromasa Fujihara, Masataka Goto, Jun Ogata, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ISM6
2006 Musical Instrument Recognizer "Instrogram" and Its Application to Music Retrieval Based on Instrumentation Similarity
abstract
Instrumentation is an important cue in retrieving musical content. Conventional methods for instrument recognition performing notewise require accurate estimation of the onset time and fundamental frequency (FO) for each note, which is not easy in polyphonic music. This paper presents a non-notewise method for instrument recognition in polyphonic musical audio signals. Instead of such note-wise estimation, our method calculates the temporal trajectory of instrument existence probabilities for every FO and visualizes it as a spectrogram-like graphical representation, called an instrogram. This method can avoid the influence by errors of onset detection and FO estimation because it does not use them. We also present methods for MPEG-7-based instrument annotation and music information retrieval based on the similarity between instrograms. Experimental results with realistic music show the average accuracy of 76.2% for the instrument annotation and that the instrogram-based similarity measure represents the actual instrumentation similarity better than an MFCC-based one
Tetsuro Kitahara, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ISM5
2006 Recognition of Simultaneous Speech by Estimating Reliability of Separated Signals for Robot Audition
Shun'ichi Yamamoto, Ryu Takeda, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
PRICAI9
2006 Using multiple edit distances to automatically grade outputs from Machine translation systems
abstract
This paper addresses the challenging problem of automatically evaluating output from machine translation (MT) systems that are subsystems of speech-to-speech MT (SSMT) systems. Conventional automatic MT evaluation methods include BLEU, which MT researchers have frequently used. However, BLEU has two drawbacks in SSMT evaluation. First, BLEU assesses errors lightly at the beginning of translations and heavily in the middle, even though its assessments should be independent of position. Second, BLEU lacks tolerance in accepting colloquial sentences with small errors, although such errors do not prevent us from continuing an SSMT-mediated conversation. In this paper, the authors report a new evaluation method called “g Rader based on Edit Distances (RED)” that automatically grades each MT output by using a decision tree (DT). The DT is learned from training data that are encoded by using multiple edit distances, that is, normal edit distance (ED) defined by insertion, deletion, and replacement, as well as its extensions. The use of multiple edit distances allows more tolerance than either ED or BLEU. Each evaluated MT output is assigned a grade by using the DT. RED and BLEU were compared for the task of evaluating MT systems of varying quality on ATR's Basic Travel Expression Corpus (BTEC). Experimental results show that RED significantly outperforms BLEU.
Yasuhiro Akiba, Kenji Imamura, Eiichiro Sumita, Hiromi Nakaiwa, Shun'ichi Yamamoto, Hiroshi G. Okuno
IEEE Trans. Speech Audio Process.6
2005 Enhanced Robot Speech Recognition Based on Microphone Array Source Separation and Missing Feature Theory
abstract
A humanoid robot under real-world environments usually hears mixtures of sounds, and thus three capabilities are essential for robot audition; sound source localization, separation, and recognition of separated sounds. While the first two are frequently addressed, the last one has not been studied so much. We present a system that gives a humanoid robot the ability to localize, separate and recognize simultaneous sound sources. A microphone array is used along with a real-time dedicated implementation of Geometric Source Separation (GSS) and a multi-channel post-filter that gives us a further reduction of interferences from other sources. An automatic speech recognizer (ASR) based on the Missing Feature Theory (MFT) recognizes separated sounds in real-time by generating missing feature masks automatically from the post-filtering step. The main advantage of this approach for humanoid robots resides in the fact that the ASR with a clean acoustic model can adapt the distortion of separated sound by consulting the post-filter feature masks. Recognition rates are presented for three simultaneous speakers located at 2m from the robot. Use of both the post-filter and the missing feature mask results in an average reduction in error rate of 42% (relative).
Shun'ichi Yamamoto, Jean-Marc Valin, Kazuhiro Nakadai, Jean Rouat, François Michaud, Tetsuya Ogata, Hiroshi G. Okuno
ICRA7
2005 Distance-Based Dynamic Interaction of Humanoid Robot with Multiple People
Tsuyoshi Tasaki, Shohei Matsumoto, Hayato Ohba, Mitsuhiko Toda, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IEA/AIE7
2005 Contextual constraints based on dialogue models in database search task for spoken dialogue systems
abstract
This paper describes the incorporation of contextual information into spoken dialogue systems in the database search task. Appropriatedialoguemodeling is requiredto manageautomatic speech recognition (ASR) errors using dialogue-level information. We define two dialogue models: a model for dialogue flow and a model of structured dialogue history. The model for dialogueflowassumesdialoguesin the databasesearchtaskconsist of only two modes. In the structured dialogue history model, query conditions are maintained as a tree structure, taking into consideration their inputted order. The constraints derived from these models are integrated by using a decision tree learning, so that the system candeterminea dialogueact of the utteranceand whether each content word should be accepted or rejected, even when it contains ASR errors. The experimental result showed that our method could interpret content words better than conventional one without the contextual information. Furthermore, it was also shown that our method was domain-independent because it achieved equivalent accuracy in another domain without any more training.
Kazunori Komatani, Naoyuki Kanda, Tetsuya Ogata, Hiroshi G. Okuno
INTERSPEECH4
2005 Multiple moving speaker tracking by microphone array on mobile robot
abstract
Real-world applications often require tracking multiple moving speakers for improving human-robot interactions and/or sound source separation. This paper presents multiple moving speaker tracking using an 8ch microphone array system installed on a mobile robot. This problem is difficult because the system does not assume that sound sources and/or the microphone array are fixed. Our solutions consist of two key ideas – time delay of arrival estimation, and multiple Kalman filters. The former localizes multiple sound sources based on beamforming in real time. Non-linear movements are tracked by using a set of Kalman filters with different history lengths in order to reduce errors in tracking multiple moving speakers under noisy and echoic environments. For quantitative evaluation of the tracking, motion references of sound sources and a mobile robot, called SIG2, were measured accurately by ultrasonic 3D tag sensors. As a result, we showed that the system tracked three simultaneous sound sources even when SIG2 moved in a room with large reverberation due to glass walls. 1.
Masamitsu Murase, Shun'ichi Yamamoto, Jean-Marc Valin, Kazuhiro Nakadai, Kentaro Yamada, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
INTERSPEECH8
2005 Implementation of active direction-pass filter on dynamically reconfigurable processor
abstract
In this paper, we report the design and implementation of a sound source separation system using a dynamically reconfigurable device. A robot in real-world environments should have an ability to treat a mixture of multiple sound signals. Active direction-pass filter (ADPF) which extracts sound from a specific direction by using a pair of microphones has been developed as such a method of sound source separation. The ADPF was used as a front-end for an automatic speech recognition system, and recognition of three simultaneous speech signals has been reported. The ADPF, however, requires a lot of computational power, while the battery capacity and the physical size of the robot are limited. To reduce the power consumption and the size of the system, we adopted the dynamically reconfigurable device, DRP developed by NEC Electronics. We implemented the ADPF on DRP, and investigated the effectiveness of dynamically reconfigurable device for these applications. The preliminary experiment shows that ADPF on DRP separates a mixture of sound sources in real-time with practical accuracy.
Shunsuke Kurotaki, Noriaki Suzuki, Kazuhiro Nakadai, Hiroshi G. Okuno, Hideharu Amano
IROS4
2005 A two-layer model for behavior and dialogue planning in conversational service robots
abstract
This paper presents a model for the behavior and dialogue planning module of conversational service robots. Most of the previously built conversational robots cannot perform dialogue management necessary for accurately recognizing human intentions and providing information to humans. This model integrates robot behavior planning models with spoken dialogue management that is robust enough to engage in mixed-initiative dialogues in specific domains. It has two layers; the upper layer is responsible for global task planning using hierarchical planning and the lower layer engages in local planning by utilizing modules called experts, which are specialized for performing certain kind of tasks by performing physical actions and engaging in dialogues. This model enables switching and canceling tasks based on recognized human intentions. A preliminary implementation of the model, which has been integrated with Honda ASIMO, has shown its effectiveness.
Mikio Nakano, Yuji Hasegawa, Kazuhiro Nakadai, Takahiro Nakamura, Johane Takeuchi, Toyotaka Torii, Hiroshi Tsujino, Naoyuki Kanda, Hiroshi G. Okuno
IROS9
2005 Extracting multi-modal dynamics of objects using RNNPB
abstract
Dynamic features play an important role in recognizing objects that have similar static features in colors and or shapes. This paper focuses on active sensing that exploits dynamic feature of an object. An extended version of the robot, Robovie-IIs, moves an object by its arm to obtain its dynamic features. Its issue is how to extract symbols from various kinds of temporal states of the object. We use the recurrent neural network with parametric bias (RNNPB) that generates self-organized nodes in the parametric bias space. The RNNPB with 42 neurons was trained with the data of sounds, trajectories, and tactile sensors generated while the robot was moving/hitting an object with its own arm. The clusters of 20 kinds of objects were successfully self-organized. The experiments with unknown (not trained) objects demonstrated that our method configured them in the PB space appropriately, which proves its generalization capability.
Tetsuya Ogata, Hayato Ohba, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno
IROS5
2005 Spatially mapping of friendliness for human-robot interaction
abstract
It is important that robots interact with multiple people. However, most research has dealt with only interaction between one robot and one person and assumed that the distance between them does not change. This paper focuses on the spatial relationships between a robot and multiple people during interaction. Based on the distance between them, our robot selects appropriate functions to use. It does this using a method we developed for spatially mapping the friendliness of each space around the robot. The robot interacts with the highest friendliness spaces (people) selectively, thereby enabling interaction between the robot and multiple people. Our humanoid robot, SIG2 which the proposed method was implemented into, interacted with about 30 visitors, at the Kyoto University Museum. The results obtained using questionnaires after interaction showed that the actions of SIG2 were easy to understand even when it interacted with multiple people at the same time and that SIG2 behaved in a friendly manner.
Tsuyoshi Tasaki, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS4
2005 Making a robot recognize three simultaneous sentences in real-time
abstract
A humanoid robot under real-world environments usually hears mixtures of sounds, and thus three capabilities are essential for robot audition; sound source localization, separation, and recognition of separated sounds. We have adopted the missing feature theory (MFT) for automatic recognition of separated speech, and developed the robot audition system. A microphone array is used along with a real-time dedicated implementation of geometric source separation (GSS) and a multi-channel post-filter that gives us a further reduction of interferences from other sources. The automatic speech recognition based on MFT recognizes separated sounds by generating missing feature masks automatically from the post-filtering step. The main advantage of this approach for humanoid robots resides in the fact that the ASR with a clean acoustic model can adapt the distortion of separated sound by consulting the post-filter feature masks. In this paper, we used the improved Julius as an MFT-based automatic speech recognizer (ASR). The Julius is a real-time large vocabulary continuous speech recognition (LVCSR) system. We performed the experiment to evaluate our robot audition system. In this experiment, the system recognizes a sentence, not an isolated word. We showed the improvement in the system performance through three simultaneous speech recognition on the humanoid SIG2.
Shun'ichi Yamamoto, Kazuhiro Nakadai, Jean-Marc Valin, Jean Rouat, François Michaud, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS8
2005 Empirical Verification of Meaning-Game-based Generalization of Centering Theory with Large Japanese Corpus
Shun Shiramatsu, Kazunori Komatani, Takashi Miyata, Koichi Hashida, Hiroshi G. Okuno
PACLIC5
2005 Walking with body-sense in virtual space using the nonlinear oscillator
abstract
This paper presents a novel construction of a locomotion system that compensates for walking-sense with simple interfaces using hands or fingers. The realization of moving with body-sense in virtual space has required high-quality system designs including walking cancellations, haptic feedback, and high-resolution displays. However, such approaches result in increasing costs of calculation and space, which obstruct the spread of VR technology. We therefore propose a new framework of locomotion systems with simple interfaces that give users "passivity and restraint", which are essential components of walking-sense. They are realized with mutual entrainment between a nonlinear oscillator and users' input. Two experiments were conducted for evaluation. The first one showed that our system gives users sense of distance based on a body standard and sense of rhythm with stable input. The second one demonstrated that the users of our system can experience a subjective body-sense and a sense of velocity.
Kenri Kodaka, Tetsuya Ogata, Hiroshi G. Okuno
SMC3
2005 Pitch-Dependent Identification of Musical Instrument Sounds
Tetsuro Kitahara, Masataka Goto, Hiroshi G. Okuno
Appl. Intell.3
2005 User Modeling in Spoken Dialogue Systems to Generate Flexible Guidance
Kazunori Komatani, Shinichi Ueno, Tatsuya Kawahara, Hiroshi G. Okuno
User Model. User Adapt. Interact.4
2004 Using a Mixture of N-Best Lists from Multiple MT Systems in Rank-Sum-Based Confidence Measure for MT Outputs
Yasuhiro Akiba, Eiichiro Sumita, Hiromi Nakaiwa, Seiichi Yamamoto, Hiroshi G. Okuno
COLING5
2004 Efficient Confirmation Strategy for Large-scale Text Retrieval Systems with Spoken Dialogue Interface
Kazunori Komatani, Teruhisa Misu, Tatsuya Kawahara, Hiroshi G. Okuno
COLING4
2004 Category-level identification of non-registered musical instrument sounds
abstract
This paper describes a method that identifies sounds of non-registered musical instruments (i.e., musical instruments that are not contained in the training data) at a category level. Although the problem of how to deal with non-registered musical instruments is essential in musical instrument identification, it has not been dealt with in previous studies. Our method solves this problem by distinguishing between registered and non-registered instruments and identifying the category name of the non-registered instruments. When a given sound is registered, its instrument name, e.g. violin, is identified. Even if it is not registered, its category name, e.g. strings, can be identified. The important issue in achieving such identification is to adopt a musical instrument hierarchy reflecting the acoustical similarity. We present a method for acquiring such a hierarchy from a musical instrument sound database. Experimental results show that around 77% of non-registered instrument sounds, on average, were correctly identified at the category level.
Tetsuro Kitahara, Masataka Goto, Hiroshi G. Okuno
ICASSP (4)3
2004 Comparing features for forming music streams in automatic music transcription
abstract
In formating temporal sequences of notes played by the same instrument (referred to as music streams), timbre of musical instruments may be a predominant feature. In polyphonic music, the performance of timbre extraction based on power-related features deteriorates, because such features are blurred when two or more frequency components are superimposed in the same frequency. To cope with this problem, we integrated timbre similarity and direction proximity with success, but left using other features as future work. In this paper, we investigate four features. timbre similarity, direction proximity, pitch transition and pitch relation consistency to clarify the precedence among them in music stream formation. Experimental results with quartet music show that direction proximity is the most dominant feature, and pitch transition is the secondary. In addition, the performance of music stream formation was improved from 63.3% by only timbre similarity to 84.9% by integrating four features.
Yohei Sakuraba, Tetsuro Kitahara, Hiroshi G. Okuno
ICASSP (4)3
2004 Improvement of Robot Audition by Interfacing Sound Source Separation and Automatic Speech Recognition with Missing Feature Theory
abstract
We have been developed robot audition system using the active direction-pass filter (ADPF) with the Scattering Theory, and demonstrated that the humanoid SIG could separate and recognize three simultaneous speeches originating from different directions. This is the first result that a robot can listen to several things simultaneously. However, its general applicability to other robots is not yet confirmed. Since automatic speech recognition (ASR) requires direction- and speaker-dependent acoustic models, it is difficult to adapt various kinds of environments. In addition ASR with lots of acoustic models causes slow processing. In this paper, these three problems are resolved. First, we confirmed the generality of the ADPF by applying it to two humanoids, SIG2 and Replie, under different environments. Next, we present the new interface between ADPF and ASR based on the Missing Feature Theory, which masks broken features of separated sound to make them unavailable to ASR. This new interface improved the recognition performance of three simultaneous speeches up to about 90%. Finally, since the ASR uses only a single acoustic model that is direction- and speaker-independent and created under clean environments, the processing of the whole system was made very light and fast.
Shun'ichi Yamamoto, Kazuhiro Nakadai, Hiroshi Tsujino, Toshio Yokoyama, Hiroshi G. Okuno
ICRA5
2004 Recognition of Emotional States in Spoken Dialogue with a Robot
Kazunori Komatani, Ryosuke Ito, Tatsuya Kawahara, Hiroshi G. Okuno
IEA/AIE4
2004 Disambiguation in determining phonemes of sound-imitation words for environmental sound recognition
abstract
Onomatopoeia, or sound-imitation words (SIWs) are important in informing sound events in human-computer communication. One problem is listener-dependency in recognizing environmental sounds by means of SIWs, that is, different listener hears the same environmental sound as a different SIW even under the same condition. Therefore, the use of usual Japanese phonemes is not adequate to express SIWs. To cope with this ambiguity problem of phoneme determination, we designed a set of new phonemes, referred to as the basic phoneme-groups, to represent environmental sounds. The basic phonemegroup consists of one or more Japanese phonemes, and thus the ambiguity problem is resolved based on it by generating one or more SIWs for a sound event. An HMM-based scheme is adopted to recognize SIWs using the phoneme-groups. Listening experiments with seven subjects showed that automatic SIW recognition based on the basic phoneme-groups outperformed ones based on the other types of phonemes. The recall and precision rate were 56.4% and 72.2%, respectively.
Kazushi Ishihara, Yuya Hattori, Tomohiro Nakatani, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
INTERSPEECH6
2004 Robot motion control using listener's back-channels and head gesture information
abstract
A novel method is described for robot gestures and utterances during a dialogue based on the listener’s understanding and interest, which are recognized from back-channels and head gestures. “Back-channels” are defined as sounds like ‘uhhuh’ uttered by a listener during a dialogue, and “head gestures” are defined as nod and tilt motions of the listener’s head. The back-channels are recognized using sound features such as power and fundamental frequency. The head gestures are recognized using the movement of the skin-color area and the optical flow data. Based on the estimated understanding and interest of the listener, the speed and size of robot motions are changed. This method was implemented in a humanoid robot called SIG2. Experiments with six participants demonstrated that the proposed method enabled the robot to increase the listener’s level of interest against the dialogue.
Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno, Tsuyoshi Tasaki, Takeshi Yamaguchi
INTERSPEECH3
2004 Assessment of general applicability of robot audition system by recognizing three simultaneous speeches
abstract
Robot audition is a critical technology in creating an intelligent robot operating in daily environments. We have developed such a robot audition system by using a new interface between sound source separation and automatic speech recognition (ASR). A mixture of speeches captured with a pair of microphones installed in the ear positions of a humanoid is separated into each speech by using active direction-pass filter (ADPF). The ADPF extracts a sound source originating from a specific direction in real-time by using interaural phase and intensity differences. The separated speech is recognized by a speech recognizer based on the missing feature theory (MFT). By using a missing feature mask, the MFT based ASR neglects distorted and missing features caused during the speech separation. A missing feature mask for each separated speech is generated in speech separation and is sent to the ASR with the separated speech. Thus, this new integration improves the performance of ASR. However, the generality of this robot audition system has not been assessed so far. In this paper, we assess its general applicability by implementing it on the three humanoids, i.e., ASIMO of Honda, SIG2, and Replie of Kyoto University. By using three simultaneous speeches as benchmarks, the robot audition system improved the performance of ASR over 50% in every humanoid, and thus its general applicability was confirmed.
Shun'ichi Yamamoto, Kazuhiro Nakadai, Hiroshi Tsujino, Hiroshi G. Okuno
IROS4
2004 Incremental Methods to Select Test Sentences for Evaluating Translation Ability
Yasuhiro Akiba, Eiichiro Sumita, Hiromi Nakaiwa, Seiichi Yamamoto, Hiroshi G. Okuno
LREC5
2004 Automatic Sound-Imitation Word Recognition from Environmental Sounds Focusing on Ambiguity Problem in Determining Phonemes
Kazushi Ishihara, Tomohiro Nakatani, Tetsuya Ogata, Hiroshi G. Okuno
PRICAI4
2004 Sound and Visual Tracking for Humanoid Robot
Hiroshi G. Okuno, Kazuhiro Nakadai, Tino Lourens, Hiroaki Kitano
Appl. Intell.1
2004 Improvement of recognition of simultaneous speech signals using AV integration and scattering theory for humanoid robots
Kazuhiro Nakadai, Daisuke Matsuura, Hiroshi G. Okuno, Hiroshi Tsujino
Speech Commun.3
2004 Effects of increasing modalities in recognizing three simultaneous speeches
Hiroshi G. Okuno, Kazuhiro Nakadai, Hiroaki Kitano
Speech Commun.1
2003 Flexible Guidance Generation Using User Model in Spoken Dialogue Systems
abstract
We address appropriate user modeling in order to generate cooperative responses to each user in spoken dialogue systems. Unlike previous studies that focus on user's knowledge or typical kinds of users, the user model we propose is more comprehensive. Specifically, we set up three dimensions of user models: skill level to the system, knowledge level on the target domain and the degree of hastiness. Moreover, the models are automatically derived by decision tree learning using real dialogue data collected by the system. We obtained reasonable classification accuracy for all dimensions. Dialogue strategies based on the user modeling are implemented in Kyoto city bus information system that has been developed at our laboratory. Experimental evaluation shows that the cooperative responses adaptive to individual users serve as good guidance for novice users without increasing the dialogue duration for skilled users.
Kazunori Komatani, Shinichi Ueno, Tatsuya Kawahara, Hiroshi G. Okuno
ACL4
2003 Chunk-Based Statistical Translation
abstract
This paper describes an alternative translation model based on a text chunk under the framework of statistical machine translation. The translation model suggested here first performs chunking. Then, each word in a chunk is translated. Finally, translated chunks are reordered. Under this scenario of translation modeling, we have experimented on a broad-coverage Japanese-English traveling corpus and achieved improved performance.
Taro Watanabe, Eiichiro Sumita, Hiroshi G. Okuno
ACL3
2003 Privacy-Enhanced SPKI Access Control on PKIX and Its Application to Web Server
abstract
Access control using PKIX (Public Key Infrastructure with X.509) may cause a privacy problem. It is caused mainly by the fact that a server can know a client's ID. To solve this problem, we proposed a restricted anonymous access control scheme using SPKI (Simple Public Key Infrastructure). It can make a server provide service to an authorized client. It still has another problem: SPKI is not so popular as PKIX. PKIX has many efficient technologies such like SSL (Secure Socket Layer), but SPKI can't directly use these technologies. In this paper our implementation utilizes the slightest extension of PKIX, namely, we use an X.509 Certificate as an Authorization Certificate and PKIX technologies, i.e. SSL. Therefore, our approach can make some proposed SPKI schemes practical and useful. In this paper the proposed scheme is applied to access control of the Web server. The system demonstrates that it succeeds in adding privacy-enhanced access control to SSL mutual authentication. We also describe and discuss the details of implementations.
Takamichi Saito, Kentaro Umesawa, Toshiyuki Kito, Hiroshi G. Okuno
AINA4
2003 Musical instrument identification based on F0-dependent multivariate normal distribution
abstract
The pitch dependency of timbres has not been fully exploited in musical instrument identification. In this paper, we present a method using an F0-dependent multivariate normal distribution of which mean is represented by a function of fundamental frequency (F0). This F0-dependent mean function represents the pitch dependency of each feature, while the F0-normalized covariance represents the nonpitch dependency. Musical instrument sounds are first analyzed by the F0-dependent multivariate normal distribution, and then identified by using the discriminant function based on the Bayes decision rule. Experimental results of identifying 6247 solo tones of 19 musical instruments by 10-fold cross validation showed that the proposed method improved the recognition rate at individual-instrument level from 75.73% to 79.73%, and the recognition rate at category level from 88.20% to 90.65%.
Tetsuro Kitahara, Masataka Goto, Hiroshi G. Okuno
ICASSP (5)3
2003 Musical instrument identification based on F0-dependent multivariate normal distribution
abstract
The pitch dependency of timbres has not been fully exploited in musical instrument identification. In this paper, we present a method using an F0-dependent multivariate normal distribution of which mean is represented by a function of fundamental frequency (FO). This F0-dependent mean function represents the pitch dependency of each feature, while the F0-normalized covariance represents the non-pitch dependency. Musical instrument sounds are first analyzed by the F0-dependent multivariate normal distribution, and then identified by using the discriminant function based on the Bayes decision rule. Experimental results of identifying 6,247 solo tones of 19 musical instruments by 10-fold cross validation showed that the proposed method improved the recognition rate at individual-instrument level from 75.73% to 79.73%, and the recognition rate at category level from 88.20% to 90.65%.
Tetsuro Kitahara, Masataka Goto, Hiroshi G. Okuno
ICME3
2003 Robot recognizes three simultaneous speech by active audition
abstract
Robots should listen to and recognize speeches with their own ears under noisy environments and simultaneous speeches to attain smooth communications with people in a real world. This paper presents three simultaneous speech recognition based on active audition which integrates audition with motion. Our robot audition system consists of three modules - a real-time human tracking system, an active direction-pass filter (ADPF) and a speech recognition system using multiple acoustic models. The real-time human tracking realizes robust and accurate sound source localization and tracking by audio-visual integration. The performance of localization shows that the resolution of the center of the robot is much higher than that of the peripheral. We call this phenomenon "auditory fovea" because it is similar to visual fovea (high resolution in the center of the human eye). Active motions such as being directed at the sound source improve localization because of making the best use if the auditory fovea. The ADPF realizes accurate and fast sound separation by using a pair of microphones. The ADPF separates sounds originating from the specified direction obtained by the real-time human tracking system. Because the performance of separation depends on the accuracy of localization, the extraction of sound from the front direction is more accurate than that of sound from the periphery. This means that the pass range of ADPF should be narrower in the front direction than in periphery. In other words, such active pass range control improves sound separation. The separated speech is recognized by the speech recognition using multiple acoustic models that integrates multiple results to output the result with the maximum likelihood. Active motions such as being directed at a sound source improve speech recognition because it realizes not only improvement of sound extraction but also easier integration of the results using face ID by face recognition. The robot audition system improved by active audition is implemented on an upper-torso humanoid. The system attains localization, separation and recognition of three simultaneous speeches and the results proves the efficiency of active audition.
Kazuhiro Nakadai, Hiroshi G. Okuno, Hiroaki Kitano
ICRA2
2003 Realizing personality in audio-visually triggered non-verbal behaviors
abstract
Controlling robot behaviors becomes more important recently as active perception for robot, in particular active audition in addition to active vision, has made remarkable progress. We are studying how to create social humanoids that perform actions empowered by real-time audio-visual tracking of multiple talkers. In this paper, we present personality as means of controlling on-verbal behaviors. It consists of two dimensions, dominance vs. submissiveness and friendliness vs. hostility, based on the interpersonal theory in psychology. The upper-torso humanoid SIG equipped with real-time audio-visual multiple-talker tracking system is used as a testbed for social interaction. As a companion robot, with friendly personality, it turns toward new sound source in order to show its attention, while with hostile personality, it turns away a new sound source. As a receptionist robot with dominant personality, it focuses its attention on the current customer, while with submissive personality, its attention to the current customer is interrupted by a new one.
Hiroshi G. Okuno, Kazuhiro Nakadai, Hiroaki Kitano
ICRA1
2003 Pitch-Dependent Musical Instrument Identification and Its Application to Musical Sound Ontology
Tetsuro Kitahara, Masataka Goto, Hiroshi G. Okuno
IEA/AIE3
2003 Design and Implementation of Personality of Humanoids in Human Humanoid Non-verbal Interaction
Hiroshi G. Okuno, Kazuhiro Nakadai, Hiroaki Kitano
IEA/AIE1
2003 Automatic transformation of environmental sounds into sound-imitation words based on Japanese syllable structure
abstract
Abstract Sound-imitation words , a sound-related subset of ono-matopoeia , are important for computer-human interaction andautomatic tagging of sound archives. The main problemof automatic recognition of sound-imitation word is that theliteral representation of such words is dependent on listen-ers and influenced by a particular cultural history. Basedon our preliminary experiments of such dependency and thesonority theory, we discovered that the process of transform-ing environmental sounds into syllable-structure expressions ismostly listener-independent whilethatoftransformingsyllable-structure expressions into sound-imitation words is mostly listener-dependent and influenced by culture. This paper fo-cuses on the former lister-independent process and presents thethree-stagearchitecture of automatictransformation of environ-mentalsounds tosound-imitationwords; segmenting soundsig-nalstosyllables, identifying syllablestructureasmora, and rec-ognizing mora as phonemes. 1. Introduction The recent development of automatic speech recognition sys-tems (ASR) has enhanced human-computer interaction and en-abled speech input over normal or cellular phones. CurrentASR’s, however, fail in recognizing non-speech sounds, in par-ticular environmental sounds such as animal voices, instrumen-tal, natural, or machine sounds. Japanese speaking people of-ten use
Kazushi Ishihara, Yasushi Tsubota, Hiroshi G. Okuno
INTERSPEECH3
2003 User modeling in spoken dialogue systems for flexible guidance generation
abstract
We address appropriate user modeling in order to generate cooperative responses to each user in spoken dialogue systems. Unlike previous studies that focus on users’ knowledge or typical kinds of users, the proposed user model is more comprehensive. Specifically, we set up three dimensions of user models: skill level to the system, knowledge level on the target domain and degree of hastiness. Moreover, the models are automatically derived by decision tree learning using real dialogue data. We obtained reasonable classification accuracy for all dimensions. Dialogue strategies based on the user modeling are implemented in Kyoto city bus information system that has been developed at our laboratory. Experimental evaluation shows that the cooperative responses adaptive to individual users serve as good guidance for novice users without increasing the dialogue duration for skilled users.
Kazunori Komatani, Shinichi Ueno, Tatsuya Kawahara, Hiroshi G. Okuno
INTERSPEECH4
2003 Three simultaneous speech recognition by integration of active audition and face recognition for humanoid
Kazuhiro Nakadai, Daisuke Matsuura, Hiroshi G. Okuno, Hiroshi Tsujino
INTERSPEECH3
2003 Applying scattering theory to robot audition system: robust sound source localization and extraction
abstract
Robot audition by its own ears (microphones) is essential for natural human-robot communication and interface. Since a microphone is embedded in the head of a robot, the head-related transfer function (HRTF) plays an important role in sound source localization and extraction. Usually, from binaural input, the interaural phase difference (IPD) and interaural intensity difference (IID) are calculated, and then the direction is determined by using IPD and IID with HRTF. The problem of HRTF-based sound source localization is that a HRTF should be measured for each robot in an anechoic chamber, because it depends on the shape of robot's head; HRTF should be interpolated to manipulate a moving talker, because it is available only for discrete azimuth and elevation. To cope with these problems of HRTF, we proposed the auditory epipolar geometry as a continuous function of IPD and IID to dispense with HRTF and have developed a real-time multiple-talker tracking system. This auditory epipolar geometry, however, does not give a good approximation to IID of all range and IPD of peripheral areas. In this paper, the scattering theory in physics is employed to take into consideration the diffraction of sounds around robot's head for better approximation of IID and IPD. The resulting system shows that it is efficient for localization and extraction of sound at higher frequency and from side directions.
Kazuhiro Nakadai, Daisuke Matsuura, Hiroshi G. Okuno, Hiroaki Kitano
IROS3
2003 Experimental comparison of MT evaluation methods: RED vs.BLEU
abstract
This paper experimentally compares two automatic evaluators, RED and BLEU, to determine how close the evaluation results of each automatic evaluator are to average evaluation results by human evaluators, following the ATR standard of MT evaluation. This paper gives several cautionary remarks intended to prevent MT developers from drawing misleading conclusions when using the automatic evaluators. In addition, this paper reports a way of using the automatic evaluators so that their results agree with those of human evaluators.
Yasuhiro Akiba, Eiichiro Sumita, Hiromi Nakaiwa, Seiichi Yamamoto, Hiroshi G. Okuno
MTSummit5
2002 Efficient Dialogue Strategy to Find Users' Intended Items from Information Query Results
Kazunori Komatani, Tatsuya Kawahara, Ryosuke Ito, Hiroshi G. Okuno
COLING4
2002 Real-Time Speaker Localization and Speech Separation by Audio-Visual Integration
abstract
Robot audition in real-world should cope with motor and other noises caused by the robot's own movements in addition to environmental noises and reverberation. This paper reports how auditory processing is improved by audio-visual integration with active movements. The key idea resides in hierarchical integration of auditory and visual streams to disambiguate auditory or visual processing. The system runs in real-time by using distributed processing on 4 PCs connected by a Gigabit Ethernet. The system implemented in a upper-torso humanoid tracks multiple talkers and extracts speech from a mixture of sounds. The performance of epipolar geometry based sound source localization and sound source separation by active and adaptive direction-pass filtering is also reported.
Kazuhiro Nakadai, Ken-ichi Hidai, Hiroshi G. Okuno, Hiroaki Kitano
ICRA3
2002 Social Interaction of Humanoid RobotBased on Audio-Visual Tracking
Hiroshi G. Okuno, Kazuhiro Nakadai, Hiroaki Kitano
IEA/AIE1
2002 Real-time sound source localization and separation for robot audition
abstract
Robot audition in the real world should cope with environment noises and reverberation and motor noises caused by the robot's own movements. This paper presents the active direction-pass filter (ADPF) to separate sounds originating from the specified direction with a pair of microphones. The ADPF is implemented by hierarchical integration of visual and auditory processing with hypothetical reasoning on interaural phase difference (IPD) and interaural intensity difference (IID) for each subband. In creating hypotheses, the reference data of IPD and IID is calculated by the auditory epipolar geometry on demand. Since the performance of the ADPF depends on the direction, the ADPF controls the direction by motor movement. The human tracking and sound source separation based on the ADPF is implemented on an upper-torso humanoid and runs in real-time with 4 PCs connected over Gigabit ethernet. The signal-to-noise ratio (SNR) of each sound separated by the ADPF from a mixture of two speeches with the same loudness is improved to about 10 dB from 0 dB.
Kazuhiro Nakadai, Hiroshi G. Okuno, Hiroaki Kitano
INTERSPEECH2
2002 Auditory fovea based speech enhancement and its application to human-robot dialog system
Kazuhiro Nakadai, Hiroshi G. Okuno, Hiroaki Kitano
INTERSPEECH2
2002 Belief network based disambiguation of object reference in spoken dialogue system for robot
abstract
We are studying joint activity in which a remote robot finds an object by communicating with the user over a voice-only channel. We focus on how the robot disambiguates the reference of the uttered word or phrase to the target object. For example, by “cup”, one may refer to a “teacup”, a “coffee cup”, or even a “glass” under some situations. This reference (hereafter, “object reference”) is user-dependent. We confirm that a user model of object references is significant by conducting a survey of 12 subjects. In addition to ambiguity of object reference, actual systems should cope with two other sources of uncertainty in speech and image recognition. We present a Belief Network based probabilistic reasoning system to determine the object reference. The resulting system demonstrates that the number of interactions needed to find a common reference is reduced as the user model is refined.
Yoko Yamakata, Tatsuya Kawahara, Hiroshi G. Okuno
INTERSPEECH3
2002 Auditory fovea based speech separation and its application to dialog system
abstract
This paper presents an active direction-pass filter (ADPF) that separates sounds originating from the specified direction by using a pair of microphones. Its application to front-end processing for speech recognition is also reported. Since the performance of sound source separation by the ADPF depends on the accuracy of sound source localization (direction), various localization modules including the interaural phase difference, interaural intensity difference for each sub-band, and other visual and auditory processing are integrated hierarchically. The resulting performance of auditory localization varies according to the relative position of the sound source. The resolution of the center of the robot is much higher than that of peripherals, indicating similar property of visual fovea. To make the best use of this property, the ADPF controls the direction of a head by motor movement. In order to recognize sound streams separated by the ADPF, a hidden Markov model based automatic speech recognition is built with multiple acoustic models trained by the output of the ADPF under different conditions. A preliminary dialog system is thus implemented on an upper-torso humanoid. The experimental results prove that it works well even when two speakers speak simultaneously.
Kazuhiro Nakadai, Hiroshi G. Okuno, Hiroaki Kitano
IROS2
2002 Realizing Audio-Visually Triggered ELIZA-Like Non-verbal Behaviors
Hiroshi G. Okuno, Kazuhiro Nakadai, Hiroaki Kitano
PRICAI1
2001 A computational model of monkey grating cells for oriented repetitive alternating patterns
Tino Lourens, Kazuhiro Nakadai, Hiroshi G. Okuno, Hiroaki Kitano
ESANN3
2001 Graph extraction from color images
Tino Lourens, Kazuhiro Nakadai, Hiroshi G. Okuno, Hiroaki Kitano
ESANN3
2001 Sound and Visual Tracking for Humanoid Robot
Hiroshi G. Okuno, Kazuhiro Nakadai, Tino Lourens, Hiroaki Kitano
IEA/AIE1
2001 Real-Time Auditory and Visual Multiple-Object Tracking for Humanoids
Kazuhiro Nakadai, Ken-ichi Hidai, Hiroshi Mizoguchi, Hiroshi G. Okuno, Hiroaki Kitano
IJCAI4
2001 Real-time multiple speaker tracking by multi-modal integration for mobile robots
abstract
In this paper, real-time multiple speaker tracking is addressed, because it is essential in robot perception and humanrobot social interaction. The difficulty lies in treating a mixture of sounds, occlusion (some talkers are hidden) and real-time processing. Our approach consists of three components; (1) the extraction of the direction of each speaker by using interaural phase difference and interaural intensity difference, (2) the resolution of each speaker's direction by multi-modal integration of audition, vision and motion with canceling inevitable motor noises in motion in case of an unseen or silent speaker, and (3) the distributed implementation to three PCs connected by TCP/IP network to attain real-time processing. As a result, we...
Kazuhiro Nakadai, Ken-ichi Hidai, Hiroshi G. Okuno, Hiroaki Kitano
INTERSPEECH3
2001 Separating three simultaneous speeches with two microphones by integrating auditory and visual processing
abstract
This paper addresses the problem of automatic recognition of three simultaneous speeches with two microphones, that is, that of sound source separation where the number of sound sources is greater than that of microphones. The approach used is the direction-pass filter, which is implemented by hypothetical reasoning on the interaural phase difference (IPD) and interaural intensity difference (IID). Auditory processing calculates IPD and IID for each subband, and generates hypotheses for precalculated IPD and IID for every direction including one obtained by visual processing. Then the system calculates the belief factor of hypothesis by Dempster-Shafer theory and determines the direction of each subband. Subbands of the specific direction are collected and then converted to a wave form by inverse FFT. With 200 benchmarks of three simultaneous utterances of Japanese words, the average 1-best and 10-best recognition rates of extracted speeches are 60% and 81%, respectively.
Hiroshi G. Okuno, Kazuhiro Nakadai, Tino Lourens, Hiroaki Kitano
INTERSPEECH1
2001 An Access Control with handling Private Information
abstract
An Internet user may want to provide only necessary information in order to access servers without disclosing his/her personal information, or with disclosing minimal personal information to the server. This requirement on privacy is neither realized by PKIX (Public Key Infrastructure with X.509 certificates) based access control nor by anonymous access. In this paper, we define a privacy-enhanc ed access control as the right of controlling the exposure of personal information and propose a privacy-enhanc ed access control mechanism by using Authorization Certificate of SPKI (Simple Public Key Infrastructure). This implementation shows that the SPKI-based WWW (World Wide Web) access control can easily replace the conventional one and that it also introduces a new service with some regulations such as ages, sex, or other features. We also discuss about the security issues of the proposed access control system.
Takamichi Saito, Kentaro Umesawa, Hiroshi G. Okuno
IPDPS3
2001 Epipolar geometry based sound localization and extraction for humanoid audition
abstract
Sound localization for a robot or an embedded system is usually solved by using inter-aural phase difference (IPD) and inter-aural intensity difference (IID). These values are calculated by using head-related transfer function (HRTF). However, the HRTF depends on the shape of the head and also on changes of the environments. Therefore, sound localization without HRTF is needed for real-world applications. In this paper, we present a new sound localization method based on auditory epipolar geometry with motion control. The auditory epipolar geometry is an extension of an epipolar geometry in stereo vision to audition, and auditory and visual epipolar geometries can share the sound source direction. The key idea is to exploit additional inputs obtained by the motor control in order to compensate damages in the IPD and IID caused by reverberation of the room and the body of the robot. The proposed system can localize and extract simultaneously two sound sources in a real-world room.
Kazuhiro Nakadai, Hiroshi G. Okuno, Hiroaki Kitano
IROS2
2001 Human-robot interaction through real-time auditory and visual multiple-talker tracking
abstract
Nakadai et al. (2001) have developed a real-time auditory and visual multiple-talker tracking technique. In this paper, this technique is applied to human-robot interaction including a receptionist robot and a companion robot at a party. The system includes face identification, speech recognition, focus-of-attention control, and sensorimotor task in tracking multiple talkers. The system is implemented on a upper-torso humanoid and the talker tracking is attained by distributed processing on three nodes connected by 100Base-TX network. The delay of tracking is 200 msec. Focus-of-attention is controlled by associating auditory and visual streams by using the sound source direction and talker position as a clue. Once an association is established, the humanoid keeps its face to the direction of the associated talker.
Hiroshi G. Okuno, Kazuhiro Nakadai, Ken-ichi Hidai, Hiroshi Mizoguchi, Hiroaki Kitano
IROS1
2000 A framework for integrating sensory information in a humanoid robot
abstract
We propose a framework towards the integration of information sensors based on the idea that the stimulus perceived through different sensors are spatial-time correlated for a short time period. Applications in robotics need to be able to process information from multiple sensors, for instance, in the case of a visible talking person. How can we relate this kind of information in a simple way, without making use of high level representation? This is the question that we want to address. A new framework based on a correlation measure of low level data information is proposed. This low level correlation measure can be used as an integration data engine to support high level task description. In the paper a coherent approach from sensor level to task level for developing a robot which can handle a large number of sensors and actuators is developed. An example how this approach can be used for a visual-sound integration task is also presented.
Iris Fermin, Hiroshi G. Okuno, Hiroshi Ishiguro, Hiroaki Kitano
IROS2
2000 Design and architecture of SIG the humanoid: an experimental platform for integrated perception in RoboCup humanoid challenge
abstract
In this paper, we report an initial design of humanoid head platform for RoboCup humanoid challenge. While many researches in RoboCup humanoid challenge naturally focus on walking and running behaviors, we focus on perception and high-level behavior issues using an upper-torso humanoid. We have designed a head/neck part of the humanoid with aesthetically designed appearance, and various processing for early perception and associated reflex. The goal of this paper is to illustrate issues in designing humanoid head platform for high-level cognition research, and provide some of initial insights obtained during the early stage of implementations. With regards to perception, the emphasis will be placed on how interaction between different perception channel and motor control can complement each other to improve processing.
Hiroaki Kitano, Hiroshi G. Okuno, Kazuhiro Nakadai, Theo Sabisch, Tatsuya Matsui
IROS2
2000 Active audition system and humanoid exterior design
abstract
We present a humanoid active audition system with improved noise cancellation using humanoid cover acoustics and the cover design for an industrial exterior design. For active audition, it is an important problem to distinguish between the internal sounds like motor noises and sounds which originate from the outer world. It is important that inevitable motor noise is cancelled while the humanoid is in motion. Therefore, humanoids require the exterior to cancel such internal noises because it separates humanoid inner world from the outer world. We report an active audition system focused on the sound source tracking by integrating audition, vision and motor movements using sound separation of the humanoid cover. The experiments show that the humanoid can track and localize sound sources more accurately, the noise cannot always be cancelled out optimally and further improvements are required.
Kazuhiro Nakadai, Tatsuya Matsui, Hiroshi G. Okuno, Hiroaki Kitano
IROS3
2000 Humanoid Active Audition System Improved by the Cover Acoustics
Kazuhiro Nakadai, Hiroshi G. Okuno, Hiroaki Kitano
PRICAI2
2000 And the Fans Are Going Wild! SIG plus MIKE
Ian Frank, Kumiko Tanaka-Ishii, Hiroshi G. Okuno, Junichi Akita, Yukiko Nakagawa, Kazuaki Maeda, Kazuhiro Nakadai, Hiroaki Kitano
RoboCup3
2000 Bridging Gap between the Simulation and Robotics with a Global Vision System
Yukiko Nakagawa, Hiroshi G. Okuno, Hiroaki Kitano
RoboCup2
1999 Harmonic sound stream segregation using localization and its application to speech stream segregation
Tomohiro Nakatani, Hiroshi G. Okuno
Speech Commun.2
1999 Listening to two simultaneous speeches
Hiroshi G. Okuno, Tomohiro Nakatani, Takeshi Kawabata
Speech Commun.1
1998 On the Properties of Combination Set Operations
Hiroshi G. Okuno, Shin-ichi Minato, Hideki Isozaki
Inf. Process. Lett.1
1997 Understanding Three Simultaneous Speeches
Hiroshi G. Okuno, Tomohiro Nakatani, Takeshi Kawabata
IJCAI (1)1
1996 Localization by harmonic structure and its application to harmonic sound stream segregation
abstract
Sound stream segregation is essential for understanding auditory events in the real-world. In this paper, we present a new method for sound stream segregation using harmonic structure and localization, or direction, in the horizontal plane. The direction of the sound source is determined by using the harmonic structure extracted from binaural inputs. The fundamental frequency of each sound is then refined by using the direction of its source. This paper discusses how the effectiveness of the harmonic-based stream segregation system (HBSS) is improved by incorporating the new method and presents the binaural HBSS (Bi-HBSS). In particular, experimental results show that the Bi-HBSS reduces the spectrum distortions and the fundamental frequency errors, compared with the HBSS and with a direction-based stream segregation system.
Tomohiro Nakatani, Masataka Goto, Hiroshi G. Okuno
ICASSP3
1996 Design and Implementation of Multiple-Context Truth Maintenance System with Binary Decision Diagram
Hiroshi G. Okuno, Osamu Shimokuni, Hidehiko Tanaka
IEA/AIE1
1996 A new speech enhancement: speech stream segregation
abstract
Speech stream segregation is presented as a new speech enhancement for automatic speech recognition.Two issues are addressed: speech stream segregation from a mixture of sounds, and interfacing speech stream segregation with automatic speech recognition.Speech stream segregation is modeled as a process of extracting harmonic fragments, grouping these extracted harmonic fragments, and substituting non-harmonic residue for non-harmonic parts of groups.The main problem in interfacing speech stream segregation with HMM-based speech recognition is how to improve the degradation of recognition performance due to spectral distortion of segregated sounds, which is caused mainly by transfer function of a binaural input.Our solution is to re-train the parameters of HMM with training data binauralized for four directions.Experiments with 500 mixtures of two women' s utterances of a word showed that the cumulative accuracy of word recognition up to the 10th candidate of each woman' s utterance is, on average, 75%.
Hiroshi G. Okuno, Tomohiro Nakatani, Takeshi Kawabata
ICSLP1
1995 A computational model of sound stream segregation with multi-agent paradigm
abstract
This paper presents a new computation model for sound stream segregation based on a multi-agent paradigm. Sound streams are thought to play a key role in auditory scene analysis, which provides a general framework for auditory research including voiced speech and music. Each agent is dynamically allocated to a sound stream, and it segregates the stream by focusing on consistent attributes. Agents interact with each other to resolve stream interference. We design agents to segregate harmonic streams and a noise stream. The presented system can segregate all the streams from a mixture of a male and a female voiced speech and a background non-harmonic noise.
Tomohiro Nakatani, Takeshi Kawabata, Hiroshi G. Okuno
ICASSP3
1995 Residue-Driven Architecture for Computational Auditory Scene Analysis
Tomohiro Nakatani, Hiroshi G. Okuno, Takeshi Kawabata
IJCAI2
1994 Auditory Stream Segregation in Auditory Scene Analysis with a Multi-Agent System
Tomohiro Nakatani, Hiroshi G. Okuno, Takeshi Kawabata
AAAI2
1994 Unified architecture for auditory scene analysis and spoken language processing
Tomohiro Nakatani, Takeshi Kawabata, Hiroshi G. Okuno
ICSLP3
1992 Experience of parallel AI programming with parallel Lisp
Hiroshi G. Okuno
Future Gener. Comput. Syst.1