EDBT 2026 Demo / reviewers in the wild / expert
Hiroshi Tsujino
dblp:69/1036
· DBLP profile ↗
55ranked-venue papers
0as first author
0since 2021 · last 2019
0000-0001-8042-2796ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 49Systems, architecture and hardware · 19Graphics, computer vision, multimedia, augmented reality and games · 12Human-computer interaction and ubiquitous computing · 5Applied, interdisciplinary, general and emerging computing · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Human-computer interaction and pervasive computing
4 papers |
Human-robot interaction · 94% Wearable and physiological sensing · 6% | |
| Artificial intelligence
5 papers |
Speech recognition and synthesis · 80% Multi-agent systems · 20% | |
| Computer graphics and multimedia
2 papers |
Audio and music processing · 100% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Speech recognition and synthesis › speech separation › computational auditory scene analysis
robot audition |
0.3 | 4 | 2010 | A hybrid framework for ego noise cancellation of a robot · ICRA 2010 A robot referee for rock-paper-scissors sound games · ICRA 2008 Improvement of Robot Audition by Interfacing Sound Source Separation and Automatic Speech Recognition with Missing Feature Theory · ICRA 2004 |
Natural language and speech › Speech recognition and synthesis
sound source separation |
0.1 | 2 | 2008 | A robot referee for rock-paper-scissors sound games · ICRA 2008 Improvement of Robot Audition by Interfacing Sound Source Separation and Automatic Speech Recognition with Missing Feature Theory · ICRA 2004 |
Human-robot interaction
physical human-robot interaction |
0.1 | 1 | 2011 | Rhythmic reference of a human while a rope turning task · HRI 2011 |
Knowledge, reasoning and agents › Multi-agent systems › human-agent interaction
human-agent collaboration |
0.1 | 1 | 2019 | Personal Partner Agents for Cooperative Intelligence · HRI 2019 |
Audio and music processing › source separation
blind source separation |
0.1 | 1 | 2010 | Blind Source Separation With Parameter-Free Adaptive Step-Size Method for Robot Audition · IEEE Trans. Speech Audio Process. 2010 |
Audio and music processing
source separation |
0.1 | 1 | 2010 | Blind Source Separation With Parameter-Free Adaptive Step-Size Method for Robot Audition · IEEE Trans. Speech Audio Process. 2010 |
Audio and music processing › adaptive filtering
step-size control |
0.1 | 1 | 2010 | Blind Source Separation With Parameter-Free Adaptive Step-Size Method for Robot Audition · IEEE Trans. Speech Audio Process. 2010 |
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.0 | 2 | 2010 | A hybrid framework for ego noise cancellation of a robot · ICRA 2010 Improvement of Robot Audition by Interfacing Sound Source Separation and Automatic Speech Recognition with Missing Feature Theory · ICRA 2004 |
Wearable and physiological sensing
multimodal sensing |
0.0 | 1 | 2011 | Rhythmic reference of a human while a rope turning task · HRI 2011 |
Audio and music processing › computational auditory scene analysis
robot audition |
0.0 | 1 | 2010 | Blind Source Separation With Parameter-Free Adaptive Step-Size Method for Robot Audition · IEEE Trans. Speech Audio Process. 2010 |
Human-robot interaction › robot communication
dialogue system |
0.0 | 1 | 2008 | A robot referee for rock-paper-scissors sound games · ICRA 2008 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › robust speech recognition
missing feature theory |
0.0 | 1 | 2004 | Improvement of Robot Audition by Interfacing Sound Source Separation and Automatic Speech Recognition with Missing Feature Theory · ICRA 2004 |
Methods — techniques the papers use, named apart from their topics
coordinative intelligence · 0.8collective intelligence · 0.8collaborative intelligence · 0.8adaptive intelligence · 0.8microphone array processing · 0.2missing feature theory · 0.2microphone array · 0.2geometric source separation · 0.2masking experiment · 0.1template subtraction · 0.1source separation · 0.1independent component analysis · 0.1ultrasonic intermodulation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Personal Partner Agents for Cooperative IntelligenceabstractWe advocate cooperative intelligence (CI) that achieves its goals in cooperating with other agents, particularly human beings, with limited resources but in complex and dynamic environments. CI is important because it delivers better performances in achieving a broad range of tasks; furthermore, cooperativeness is key to human intelligence, and the processes of cooperation can contribute to help people gain several life values. This paper discusses elements in CI and our research approach to CI. We identify the four aspects of CI: adaptive intelligence, collective intelligence, coordinative intelligence, and collaborative intelligence. We also take an approach that focuses on the implementation of coordinative intelligence in the form of personal partner agents (PPAs) and consider the design of our robotic research platform to physically realize PPAs. Kotaro Funakoshi, Hideaki Shimazaki, Takatsune Kumada, Hiroshi Tsujino |
HRI | 4 |
| 2011 | Rhythmic reference of a human while a rope turning taskabstractThis paper addresses the rhythmic reference in physical human-robot interaction. Human refers to a rhythm from multiple sensing modalities when turning a rope with another human synchronously. This study verifies a hypothesis that some humans mix several rhythms of the modalities into a rhythm (rhythmic reference). Six participants, four males and two females, 21-23 years old, took part in eight experiments which examined the hypothesis. In each experiment, we masked the perception of each participant using eight combination of three kinds of masks, an eye-mask, headphones, and a force mask. Each participant interacted with an operator that turned a rope with a constant frequency. As a result of the experiments, a participant increased the controlling error as the number of masks was increased regardless the types of masked modalities. The result strongly supported our hypothesis. Kenta Yonekura, Chyon Hae Kim, Kazuhiro Nakadai, Hiroshi Tsujino, Shigeki Sugano |
HRI | 4 |
| 2011 | Online motion selection for semi-optimal stabilization using reverse-time treeabstractThis paper presents a general method for creating an approximately optimal online stabilization system. An optimal stabilization system is an ideal online system that can calculate each optimal motion leading to a stable mechanical goal state depending on the current state. We propose a system that selects each semi-optimal motion according to the current state from a reverse-time tree. To create the reverse-time tree, we applied rapid semi-optimal motion planning method (RASMO) to a reverse-time search from a stable state. We also developed an online motion selection technique. To validate the proposed method, we simulated the stabilization of a double inverted pendulum. When we used an optimization criteria, time optimal, the system quickly stabilized the pendulum's posture and velocity. When we used higher resolution RASMO, the time approached the optimal time. The general framework proposed here is applicable to a variety of machines. Chyon Hae Kim, Hiroshi Tsujino, Shigeki Sugano |
IROS | 2 |
| 2011 | Ego noise cancellation of a robot using missing feature masks
Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Hiroshi Tsujino, Jun-ichi Imura |
Appl. Intell. | 4 |
| 2011 | A multi-expert model for dialogue and behavior control of conversational robots and agents
Mikio Nakano, Yuji Hasegawa, Kotaro Funakoshi, Johane Takeuchi, Toyotaka Torii, Kazuhiro Nakadai, Naoyuki Kanda, Kazunori Komatani, Hiroshi G. Okuno, Hiroshi Tsujino |
Knowl. Based Syst. | 10 |
| 2011 | Real-time simulation of a spiking neural network model of the basal ganglia circuitry using general purpose computing on graphics processing units
Jun Igarashi, Osamu Shouno, Tomoki Fukai, Hiroshi Tsujino |
Neural Networks | 4 |
| 2010 | One-Shot Supervised Reinforcement Learning for Multi-targeted Tasks: RL-SAS
Johane Takeuchi, Hiroshi Tsujino |
ICANN (2) | 2 |
| 2010 | A hybrid framework for ego noise cancellation of a robotabstractNoise generated due to the motion of a robot is not desired, because it deteriorates the quality and intelligibility of the sounds recorded by robot-embedded microphones. It must be reduced or cancelled to achieve automatic speech recognition with a high performance. In this work, we divide ego-motion noise problem into three subdomains of arm, leg and head motion noise, depending on their complexity and intensity levels. We investigate methods that make use of single-channel and multi-channel processing in order to suppress ego noise separately. For this purpose, a framework consisting of a microphone-array-based geometric source separation, a consequent post filtering process and a parallel module for template subtraction is used. Furthermore, a control mechanism is proposed, which is based on signal-to-noise ratio and instantaneously detected motions, to switch to the most suitable method to deal with the current type of noise. We evaluate the proposed techniques on a humanoid robot using automatic speech recognition (ASR). The preliminary results of isolated word recognition show the effectiveness of our methods by increasing the word correct rates up to 50% compared to the single channel recognition in arm and leg motion noises and up to 25% in very strong head motion noises. Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Yuji Hasegawa, Hiroshi Tsujino, Jun-ichi Imura |
ICRA | 5 |
| 2010 | Robust Ego Noise Suppression of a Robot
Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Hiroshi Tsujino, Jun-ichi Imura |
IEA/AIE (1) | 4 |
| 2010 | A robust speech recognition system against the ego noise of a robot
Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Hiroshi Tsujino, Jun-ichi Imura |
INTERSPEECH | 4 |
| 2010 | Multi-talker speech recognition under ego-motion noise using Missing Feature TheoryabstractThis paper presents a system that gives a mobile robot the ability to recognize target speaker's speech, even if the robot performs an action and there are multiple speakers talking in the room. Associated problems to this system are twofold: (1) While the robot is moving, the joints inevitably generate ego-motion noise due to its motors. (2) Recognizing target speech against other interfering speech signals is a difficult task. Since typical solutions to (1) and (2), motor noise suppression and sound source separation, both introduce distortion to the processed signals, the performance of automatic speech recognition (ASR) deteriorates. Instead of removing the ego-motion noise with conventional noise suppression methods, in this work, we investigate methods to eliminate the unreliable parts of the audio features that are contaminated by the ego-motion noise. For this purpose, we model masks that filter unreliable speech features based on the ratio of speech and motor noise energies. We analyze the performance of the proposed technique under various test conditions by comparing it to the performance of existing Missing Feature Theory-based ASR implementations. Finally, we propose an integration framework for two different masks that are designed to eliminate ego noise and to filter the leakage energy of interfering sound sources. We demonstrate that the proposed methods achieve a high ASR accuracy. Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Hiroshi Tsujino, Jun-ichi Imura |
IROS | 4 |
| 2010 | Blind Source Separation With Parameter-Free Adaptive Step-Size Method for Robot AuditionabstractThis paper proposes an adaptive step-size method for blind source separation (BSS) suitable for robot audition systems. The design of the step-size parameter is a critical consideration when we apply BSS to real-world applications such as robot audition systems, because the surrounding environment dynamically changes in the real world. It is common to use a fixed step-size parameter that was obtained empirically. However, because of environmental changes and noise, the performance of BSS with a fixed step-size parameter deteriorates and the separation matrix sometimes diverges. Several adaptive step-size methods for BSS have been proposed. However, there are difficulties when applying them to robot audition systems for example, low-computational cost requirements, being free from manual parameter adjustment and so on. We propose an adaptive step-size method suitable for robot audition systems. The proposed method has the following merits: 1) low computational cost; 2) no parameters to be adjusted manually; and 3) no additional preprocessing requirements. We applied our method to six different BSS algorithms for an eight-channel microphone array embedded in Honda's ASIMO robot. The method improved the performance of all six algorithms in experiments on separation and recognition of simultaneous speech. Moreover, the method increased the amount of calculation by less than 10% compared with the original calculation used in most BSS algorithms. Hirofumi Nakajima, Kazuhiro Nakadai, Yuji Hasegawa, Hiroshi Tsujino |
IEEE Trans. Speech Audio Process. | 4 |
| 2009 | Sound source separation of moving speakers for robot auditionabstractThis paper addresses sound source separation and speech recognition for moving sound sources. Real-world applications such as robots should cope with both moving and stationary sound sources. However, most studies assume only stationary sound sources. We introduce two key techniques to cope with moving sources, that is, Adaptive Step-size control (AS) and Optima Controlled Recursive Average (OCRA) to improve blind source separation. We implemented a real-time robot audition system with these techniques for our humanoid robot ASIMO with an 8ch microphone array by using HARK which is our open-source software for robot audition. The performance of the system will be shown through sound source separation for moving sources and automatic speech recognition of separated speeches. Kazuhiro Nakadai, Hirofumi Nakajima, Yuji Hasegawa, Hiroshi Tsujino |
ICASSP | 4 |
| 2009 | Ego noise suppression of a robot using template subtractionabstractWhile a robot is moving, the joints inevitably generate noise due to its motors, i.e. ego-motion noise. This problem is very crucial, especially in humanoid robots, because it tends to have a lot of joints and the motors are located closer to the microphones than the sound sources. In this work, we investigate methods for the prediction and suppression of the ego-motion noise. In the first part, we analyze the performance of different noise subtraction strategies, assuming that the noise prediction problem has been solved. In the second part, we present some results for a noise prediction scheme based on the current robot joint status. Performance is evaluated for a number of criteria, including Automatic Speech Recognition (ASR). We demonstrate that our method improves recognition performance during ego-motion considerably. Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Yuji Hasegawa, Hiroshi Tsujino, Jun-ichi Imura |
IROS | 5 |
| 2009 | Thereminist robot: Development of a robot theremin player with feedforward and feedback arm control based on a Theremin's pitch modelabstractWe propose a Thereminist robot system that plays the Theremin based on a Theremin's pitch model. The theremin, which is a 1920s electronic musical instrument, is played by moving a player's hand position in the air without touching it. It is difficult to play the Theremin because the relationship between the hand position and Theremin's pitch (pitch characteristics) is non-linear and varies according to the electromagnetic field (hereafter called environment). These characteristics cause two problems: (1) adapting to the environment change is required and (2) a nai¿ve design tends to depend on robot's particular hardware. We implement the coarse-to-fine control system on the Thereminist robot using newly proposed two pitch models: parametric and nonparametric ones. The Thereminist robot works as below: first, the robot calibrates the pitch model by parameter fitting with the Levenberg-Marquardt method. Second, the robot moves its hand in a coarse manner by feedforward control based on the pitch model. Finally, the robot adjusts its position by feedback control (proportional-integral control). In these steps, the robot can play a required pitch quickly, because the robot moves its hand using the pitch model without listening to the Theremin's sound Thus, the time to play the exact pitch is shorter than when only feedback control is used. Three experiments were conducted to evaluate the robustness against the number of samples, environment change, and types of robots. The results revealed that our pitch model describes using only 12 samples of pitches for estimation of the parameters, and adapts if the environment changes. In addition, our system works on two different robots: HRP-2 and ASIMO. Takeshi Mizumoto, Hiroshi Tsujino, Toru Takahashi 0001, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 2 |
| 2009 | Intelligent sound source localization for dynamic environmentsabstractAs robotic technology plays an increasing role in human lives, ¿robot audition¿, human-robot communication, is of great interest, and robot audition needs to be robust and adaptable for dynamic environments. This paper addresses sound source localization working in dynamic environments for robots. Previously, noise robustness and dynamic localized sound selection have been enormous issues for practical use. To correct the issues, a new localization system ¿Selective Attention System¿ is proposed. The system has four new functions: localization with Generalized EigenValue Decomposition of correlation matrices for noise robustness(¿Localization with GEVD¿), sound source cancellation and focus (¿Target Source Selection¿), human-like dynamic Focus of Attention (¿Dynamic FoA¿), and correlation matrix estimation for robotic head rotation (¿Correlation Matrix Estimation¿). All are achieved by the dynamic design of correlation matrices. The system is implemented into a humanoid robot, and the experimental validation is successfully verified even when the robot microphones move dynamically. Keisuke Nakamura, Kazuhiro Nakadai, Futoshi Asano, Yuji Hasegawa, Hiroshi Tsujino |
IROS | 5 |
| 2009 | A Model for Learning Topographically Organized Parts-Based Representations of Objects in Visual Cortex: Topographic Nonnegative Matrix FactorizationabstractObject representation in the inferior temporal cortex (IT), an area of visual cortex critical for object recognition in the primate, exhibits two prominent properties: (1) objects are represented by the combined activity of columnar clusters of neurons, with each cluster representing component features or parts of objects, and (2) closely related features are continuously represented along the tangential direction of individual columnar clusters. Here we propose a learning model that reflects these properties of parts-based representation and topographic organization in a unified framework. This model is based on a nonnegative matrix factorization (NMF) basis decomposition method. NMF alone provides a parts-based representation where nonnegative inputs are approximated by additive combinations of nonnegative basis functions. Our proposed model of topographic NMF (TNMF) incorporates neighborhood connections between NMF basis functions arranged on a topographic map and attains the topographic property without losing the parts-based property of the NMF. The TNMF represents an input by multiple activity peaks to describe diverse information, whereas conventional topographic models, such as the self-organizing map (SOM), represent an input by a single activity peak in a topographic map. We demonstrate the parts-based and topographic properties of the TNMF by constructing a hierarchical model for object recognition where the TNMF is at the top tier for learning high-level object features. The TNMF showed better generalization performance over NMF for a data set of continuous view change of an image and more robustly preserving the continuity of the view change in its object representation. Comparison of the outputs of our model with actual neural responses recorded in the IT indicates that the TNMF reconstructs the neuronal responses better than the SOM, giving plausibility to the parts-based learning of the model. Kenji Hosoda, Masataka Watanabe, Heiko Wersing, Edgar Körner, Hiroshi Tsujino, Hiroshi Tamura, Ichiro Fujita |
Neural Comput. | 5 |
| 2008 | Modular Neural Networks for Model-Free Behavioral Learning
Johane Takeuchi, Osamu Shouno, Hiroshi Tsujino |
ICANN (1) | 3 |
| 2008 | Adaptive step-size parameter control for real-world blind source separationabstractThis paper describes a method to adaptively control a step-size parameter which is used for updating a separation matrix to extract a target sound source accurately in blind source separation (BSS). The design of the step-size parameter is essential when we apply BSS to real-world applications such as robot audition systems, because the surrounding environment dynamically changes in the real world. It is common to use a fixed step-size parameter that is obtained empirically. However, due to environmental changes and noises, the performance of BSS with the fixed step-size parameter deteriorates and the separation matrix sometimes diverges. We propose a general method that allows adaptive step-size control. The proposed method is an extension of Newton’s method utilizing a complex gradient theory and is applicable to any BSS algorithm. Actually, we applied it to six types of BSS algorithms for an 8 ch microphone array embedded in Honda ASIMO. Experimental results show that the proposed method improves the performance of these six BSS algorithms through experiments of separation and recognition for two simultaneous speeches. Hirofumi Nakajima, Kazuhiro Nakadai, Yuji Hasegawa, Hiroshi Tsujino |
ICASSP | 4 |
| 2008 | Smoothing human-robot speech interactions by using a blinking-light as subtle expressionabstractSpeech overlaps, undesired collisions of utterances between systems and users, harm smooth communication and degrade the usability of systems. We propose a method to enable smooth speech interactions between a user and a robot, which enables subtle expressions by the robot in the form of a blinking LED attached to its chest. In concrete terms, we show that, by blinking an LED from the end of the user's speech until the robot's speech, the number of undesirable repetitions, which are responsible for speech overlaps, decreases, while that of desirable repetitions increases. In experiments, participants played a last-and-first game with the robot. The experimental results suggest that the blinking-light can prevent speech overlaps between a user and a robot, speed up dialogues, and improve user's impressions. Kotaro Funakoshi, Kazuki Kobayashi, Mikio Nakano, Seiji Yamada, Yasuhiko Kitamura, Hiroshi Tsujino |
ICMI | 6 |
| 2008 | Self-Referential Event Lists for Self-Organizing Modular Reinforcement Learning
Johane Takeuchi, Osamu Shouno, Hiroshi Tsujino |
ICONIP (2) | 3 |
| 2008 | Adaptive habituation detection to build human computer interactive systems using a real-time cross-modal computationabstractWe propose a new habituation detection system using a cross-modal computation. The cross-modal sensory data comprised of eye-movement and skin potential level (SPL) for our habituation detection system has the substantial temporal/spatial nonstationarity. Therefore, it was difficult for conventional classification methods to detect the boundary of the habituation state from the sensory data. Hence, we introduced an Allen-Cahn type partial differential equation (PDE) method to deal with the uncertainty, and developed a new real-time habituation detection system. The result demonstrates that our proposed method performs better classification of the nonstationary data than conventional methods even with a small amount of data. Motohri Kon, Takamasa Koshizen, Kazuyuki Aihara, Hiroshi Tsujino |
ICPR | 4 |
| 2008 | A robot referee for rock-paper-scissors sound gamesabstractThis paper describes a robot referee for “rockpaper-scissors (RPS)” sound games; the robot decides the winner from a combination of rock, paper and scissors uttered by two or three people simultaneously without using any visual information. In this referee task, the robot has to cope with speech with low signal-to-noise ratio (SNR) due to a mixture of speeches, robot motor noises, and ambient noises. Our robot referee system, thus, consists of two subsystems - a real-time robot audition subsystem and a dialog subsystem focusing on RPS sound games. The robot audition subsystem can recognize simultaneous speeches by exploiting two key ideas; preprocessing consisting of sound source localization and separation with a microphone array, and system integration based on missing feature theory (MFT). Preprocessing improves the SNR of a target sound signal using geometric source separation with a multi-channel post-filter. MFT uses only reliable acoustic features in speech recognition and masks out unreliable parts caused by interfering sounds and preprocessing. MFT thus provides smooth integration between preprocessing and automatic speech recognition. The dialog subsystem is implemented as a system-initiative dialog system for multiple players based on deterministic finite automata. It first waits for a trigger command to start an RPS sound game, controls the dialog with players in the game, and finally decides the winner of the game. The referee system is constructed for Honda ASIMO with an 8-ch microphone array. In the case with two players, we attained a 70% task completion rate for the games on average. Kazuhiro Nakadai, Shun'ichi Yamamoto, Hiroshi G. Okuno, Hirofumi Nakajima, Yuji Hasegawa, Hiroshi Tsujino |
ICRA | 6 |
| 2008 | Rapid Prototyping of Robust Language Understanding Modules for Spoken Dialogue Systems
Yuichiro Fukubayashi, Kazunori Komatani, Mikio Nakano, Kotaro Funakoshi, Hiroshi Tsujino, Tetsuya Ogata, Hiroshi G. Okuno |
IJCNLP | 5 |
| 2008 | A robot uses its own microphone to synchronize its steps to musical beats while scatting and singingabstractMusical beat tracking is one of the effective technologies for human-robot interaction such as musical sessions. Since such interaction should be performed in various environments in a natural way, musical beat tracking for a robot should cope with noise sources such as environmental noise, its own motor noises, and self voices, by using its own microphone. This paper addresses a musical beat tracking robot which can step, scat and sing according to musical beats by using its own microphone. To realize such a robot, we propose a robust beat tracking method by introducing two key techniques, that is, spectro-temporal pattern matching and echo cancellation. The former realizes robust tempo estimation with a shorter window length, thus, it can quickly adapt to tempo changes. The latter is effective to cancel self noises such as stepping, scatting, and singing. We implemented the proposed beat tracking method for Honda ASIMO. Experimental results showed ten times faster adaptation to tempo changes and high robustness in beat tracking for stepping, scatting and singing noises. We also demonstrated the robot times its steps while scatting or singing to musical beats. Kazumasa Murata, Kazuhiro Nakadai, Kazuyoshi Yoshii, Ryu Takeda, Toyotaka Torii, Hiroshi G. Okuno, Yuji Hasegawa, Hiroshi Tsujino |
IROS | 8 |
| 2008 | High performance sound source separation adaptable to environmental changes for robot auditionabstractThis paper describes a novel sound source separation method for a robot that needs to cope with dynamically changing noises in the real world. The sound source separation method, Geometric Source Separation (GSS), is promising because it has high separation performance and requires low computational cost. One of the most important factors in GSS performance is a step-size parameter to update a separation matrix which is generally used for extracting a target sound source. A fixed value that was obtained empirically is commonly used as the step-size parameter. However, in the real world, the surrounding environment changes dynamically. Thus, conventional GSS with a fixed step-size parameter sometimes results in poor separation results, or divergence of the separation matrix. Another important factor is the weight parameter, which adjusts the balance between geometric errors and separation errors and also affects performance. If this parameter is set to a small value, GSS becomes similar to a Blind Source Separation method, by which the output signal may contain errors based on indefinite source amplitudes and orders. In contrast, if this parameter is set to a large value, GSS becomes similar to a delay-and-sum Beamforming method, which does not have high separation performance. GSS gives good performance when the parameters are tuned to an optimum value, which changes according to the environment. We propose two effective methods that can be used for general BSS’s. One is an adaptive step-size parameter control method. By using this method, the step-size and the weight parameters are automatically set to optimum values and are able to adapt to environmental changes. The other is an optima controlled recursive average method for correlation matrix estimation. This method can improve the estimation of a separation matrix, and achieve high separation performance. We evaluated the proposed GSS algorithm with an 8ch microphone array embedded in Honda ASIMO. Experimental results showed that the proposed method improved sound source separation even in dynamically changing environments. Hirofumi Nakajima, Kazuhiro Nakadai, Yuji Hasegawa, Hiroshi Tsujino |
IROS | 4 |
| 2008 | Smoothing human-robot speech interaction with blinking-light expressionsabstractWe propose a method to enable smooth speech interactions between a user and a robot. Our method is based on subtle expression whereby a robot blinks a small LED attached to its chest. We performed experiments in which participants played a last-and-first games and counted the number of repetitions made by the participants and analyzed their impression of the game and the robot. The experimental results suggested that the blinking-light could prevent utterance collisions between a user and a robot and could create familiar and attentive impressions about the game on users. Kazuki Kobayashi, Kotaro Funakoshi, Seiji Yamada, Mikio Nakano, Yasuhiko Kitamura, Hiroshi Tsujino |
RO-MAN | 6 |
| 2007 | Design and implementation of a robot audition system for automatic speech recognition of simultaneous speechabstractThis paper addresses robot audition that can cope with speech that has a low signal-to-noise ratio (SNR) in real time by using robot-embedded microphones. To cope with such a noise, we exploited two key ideas; Preprocessing consisting of sound source localization and separation with a microphone array, and system integration based on missing feature theory (MFT). Preprocessing improves the SNR of a target sound signal using geometric source separation with multichannel post-filter. MFT uses only reliable acoustic features in speech recognition and masks unreliable parts caused by errors in preprocessing. MFT thus provides smooth integration between preprocessing and automatic speech recognition. A real-time robot audition system based on these two key ideas is constructed for Honda ASIMO and Humanoid SIG2 with 8-ch microphone arrays. The paper also reports the improvement of ASR performance by using two and three simultaneous speech signals. Shun'ichi Yamamoto, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ASRU | 4 |
| 2007 | The Design of Phoneme Grouping for Coarse Phoneme Recognition
Kazuhiro Nakadai, Ryota Sumiya, Mikio Nakano, Koichi Ichige, Yasuo Hirose, Hiroshi Tsujino |
IEA/AIE | 6 |
| 2007 | Modular Neural Networks for Reinforcement Learning with Temporal Intrinsic RewardsabstractInspired by intrinsic motivation that is thought to play a crucial role in animal development and learning, several artificial learning systems with built in intrinsic rewards were recently studied. Here we suggest an intrinsically rewarded learning system for autonomous task achievements that copes with several kinds of transitions. The system consists of neural networks equipped with a modular reinforcement learning algorithm. The modular system that decomposes the observed state space stabilizes the intrinsic rewards calculated from prediction errors. On-line learning via the proposed system takes place under various kinds of transitions, including deterministic, probabilistic and partially observable, without any specific adjustments of parameters for each transition. The combined system with both the modular network and the intrinsic reward generator led to performance that converged to the optimal sequences of actions in all tested transitions, in which external rewards were delivered only at the completion of tasks. Johane Takeuchi, Osamu Shouno, Hiroshi Tsujino |
IJCNN | 3 |
| 2007 | Robust acquisition and recognition of spoken location names by domestic robotsabstractThis paper presents a method that enables a conversational domestic robot to learn location names through speech interaction. Each acquired name is associated with a point on the map coordinate system of the robot. Both for acquisition and recognition of location names, a bag- of-words-based categorization technique is used. Namely, the robot acquires a location name as a frequency pattern of words, and recognizes a spoken location name by computing similarity between the patterns. This makes the robot robust not only against speech recognition errors but also against out- of-vocabulary names. We designed a dialogue and behavior management subsystem that learns location names by using our proposed method and navigates to indicated locations, and implemented the subsystem on an omnidirectional cart robot. The result of a preliminary evaluation of the implemented robot with human subjects suggested this approach is promising. Kotaro Funakoshi, Mikio Nakano, Toyotaka Torii, Yuji Hasegawa, Hiroshi Tsujino, Noriyuki Kimura, Naoto Iwahashi |
IROS | 5 |
| 2007 | A biped robot that keeps steps in time with musical beats while listening to music with its own earsabstractWe aim at enabling a biped robot to interact with humans through real-world music in daily-life environments, e.g., to autonomously keep its steps (stamps) in time with musical beats. To achieve this, the robot should be able to robustly predict the beat times in real time while listening to musical performance with its own ears (head-embedded microphones). However, this has not previously been addressed in most studies on music-synchronized robots due to the difficulty in predicting the beat times in real-world music. To solve this problem, we implemented a beat-tracking method developed in the field of music information processing. The predicted beat times are then used by a feedback-control method that adjusts the robot's step intervals to synchronize its steps in time with the beats. The experimental results show that the robot can adjust its steps in time with the beat times as the tempo changes. The resulting robot needed about 25 [s] to recognize the tempo change after it and then synchronize its steps. Kazuyoshi Yoshii, Kazuhiro Nakadai, Toyotaka Torii, Yuji Hasegawa, Hiroshi Tsujino, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 5 |
| 2007 | A markup language for describing interactive humanoid robot presentationsabstractThis paper presents a multi-modal presentation markup language for humanoid robots, MPML-HR ver. 3.0, which is able to describe presentation contents including speech-based interactions with audiences. Previous versions of MPML-HR do not feature any interaction functionality which dynamically changes the presentation according to the utterances by audiences, although such interaction makes the presentation more effective and understandable. Since MPML-HR ver. 3.0 inherits simple descriptions of previous versions of MPML-HR, the content designer can describe interactive presentations without configuring conventional complicated multi-modal interactive systems. Yoshitaka Nishimura, Shinichiro Minotsu, Hiroshi Dohi, Mitsuru Ishizuka, Mikio Nakano, Kotaro Funakoshi, Johane Takeuchi, Yuji Hasegawa, Hiroshi Tsujino |
IUI | 9 |
| 2006 | Robust Tracking of Multiple Sound Sources by Spatial Integration of Room And Robot Microphone ArraysabstractSound source tracking is an important function for a robot operating in a daily environment, because the robot should recognize where a sound event such as speech, music and other environmental sounds originates from. This paper addresses sound source tracking by integrating a room and a robot microphone array. The room microphone array consists of 64 microphones attached to the walls. It provides 2D (x-y) sound source localization based on a weighted delay-and-sum beamforming method. The robot microphone array consists of eight microphones installed on a robot head, and localizes multiple sound sources in azimuth. The localization results are integrated to track sound sources by using a particle filter for multiple sound sources. The experimental results show that particle filter based integration reduces localization errors and provides accurate and robust 2D sound source tracking. Kazuhiro Nakadai, Hirofumi Nakajima, Masamitsu Murase, Satoshi Kaijiri, Kentaro Yamada, Takahiro Nakamura, Yuji Hasegawa, Hiroshi G. Okuno, Hiroshi Tsujino |
ICASSP (4) | 9 |
| 2006 | Genetic Algorithm-Based Improvement of Robot Hearing Capabilities in Separating and Recognizing Simultaneous Speech Signals
Shun'ichi Yamamoto, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Ryu Takeda, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE | 4 |
| 2006 | Connectionist Reinforcement Learning with Cursory Intrinsic Motivations and Linear Dependencies to Multiple RepresentationsabstractA significant feature of brain intelligence is flexibility. This is generally lacking in current machine intelligence We think that learning that effectively uses the combination of multiple information representations is the key to constructing flexible machine intelligence. This hypothesis is demonstrated by means of a simple connectionist model of intrinsically motivated reinforcement learning. A linear approximation of reward functions that depends on multiple representations is engaged in our model. We show preliminary results for a model network that enables a flexible learning response to several different situations. Multiple representations in our model accelerate the learning not only in complex situations that need many kinds of information, but also in simple situations. Johane Takeuchi, Osamu Shouno, Hiroshi Tsujino |
IJCNN | 3 |
| 2006 | Real-Time Tracking of Multiple Sound Sources by Integration of In-Room and Robot-Embedded Microphone ArraysabstractReal-time and robust sound source tracking is an important function for a robot operating in a daily environment, because the robot should recognize where a sound event such as speech, music and other environmental sounds originate from. This paper addresses real-time sound source tracking by real-time integration of an in-room microphone array (IRMA) and a robot-embedded microphone array (REMA). The IRMA system consists of 64 ch microphones attached to the walls. It localizes multiple sound sources based on weighted delay-and-sum beam-forming on a 2D plane. The REMA system localizes multiple sound sources in azimuth using eight microphones attached to a robot's head on a rotational table. The localization results are integrated to track multiple sound sources by using a particle filter in real-time. The experimental results show that particle filter based integration improved accuracy and robustness in multiple sound source tracking even when the robot's head was in rotation Kazuhiro Nakadai, Hirofumi Nakajima, Masamitsu Murase, Hiroshi G. Okuno, Yuji Hasegawa, Hiroshi Tsujino |
IROS | 6 |
| 2006 | Comparison of a Humanoid Robot and an On-Screen Agent as Presenters to AudiencesabstractBoth on-screen agents and humanoid robots are increasingly used as human-computer interfaces. This study evaluates an on-screen agent and a humanoid robot in the task of one-sided presentations. We compared the participant's subjective impressions of nearly identical presentation contents performed by each presenter. The results derived by the semantic differential (SD) method and the direct questioning show that each presenter has different functional advantages. We infer that on-screen agent and robot can complement each other in presentations Johane Takeuchi, Kushida Kazutaka, Yoshitaka Nishimura, Hiroshi Dohi, Mitsuru Ishizuka, Mikio Nakano, Hiroshi Tsujino |
IROS | 7 |
| 2006 | Real-Time Robot Audition System That Recognizes Simultaneous Speech in The Real WorldabstractThis paper presents a robot audition system that recognizes simultaneous speech in the real world by using robot-embedded microphones. We have previously reported missing feature theory (MFT) based integration of sound source separation (SSS) and automatic speech recognition (ASR) for building robust robot audition. We demonstrated that a MFT-based prototype system drastically improved the performance of speech recognition even when three speakers talked to a robot simultaneously. However, the prototype system had three problems; being offline, hand-tuning of system parameters, and failure in voice activity detection (VAD). To attain online processing, we introduced FlowDesigner-based architecture to integrate sound source localization (SSL), SSS and ASR. This architecture brings fast processing and easy implementation because it provides a simple framework of shared-object-based integration. To optimize the parameters, we developed genetic algorithm (GA) based parameter optimization, because it is difficult to build an analytical optimization model for mutually dependent system parameters. To improve VAD, we integrated new VAD based on a power spectrum and location of a sound source into the system, since conventional VAD relying only on power often fails due to low signal-to-noise ratio of simultaneous speech. We, then, constructed a robot audition system for Honda ASIMO. As a result, we showed that the system worked online and fast, and had a better performance in robustness and accuracy through experiments on recognition of simultaneous speech in a noisy and echoic environment Shun'ichi Yamamoto, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 4 |
| 2006 | Recognition of Simultaneous Speech by Estimating Reliability of Separated Signals for Robot Audition
Shun'ichi Yamamoto, Ryu Takeda, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
PRICAI | 5 |
| 2005 | Towards New Human-Humanoid Communication: Listening During Speaking by Using Ultrasonic Directional SpeakerabstractThis paper presents a new human-humanoid communication system by using a directional speaker. The directional speaker produces directional sound beams by using intermodulation of ultrasonic sound beams and non linearity in air. This technology solves problems and brings new ways of human-humanoid communication as follows: 1) A humanoid can recognize human speeches during speaking because a microphone installed in the humanoid does not capture self voices played by the directional speaker. 2) A humanoid can speak to a specific person as if it whispers. The directional speaker is installed at the position of the mouth of the humanoid. Preliminary experiments show the efficiency of the directional speaker to realize the above two functions. Kazuhiro Nakadai, Hiroshi Tsujino |
ICRA | 2 |
| 2005 | Sound source tracking with directivity pattern estimation using a 64 ch microphone arrayabstractIn human-robot communication, a robot should distinguish between voices uttered by a human and those played by a loudspeaker such as on a TV or a radio. This paper addresses detection of actual human voices by using a microphone array as an extension of auditory function of the robot to support environmental understanding by the robot. We introduce a 64 ch microphone array system in a room and propose a new method based on weighted delay-and-sum beamforming to estimate a directivity pattern of a sound source. The microphone array system localizes a sound source and estimates its directivity pattern. The directivity pattern estimation has two advantages as follows: One is that the system can detect whether the sound source is an actual human voice or not by comparing the estimated directivity pattern with prerecorded directivity patterns. The other is that the heading of the sound source is estimated by detecting the angle with the highest power in the directivity pattern. As a result, we proved the effectiveness of our microphone array through sound source tracking with orientation and detection of actual human voices based on directivity pattern estimation. Kazuhiro Nakadai, Hirofumi Nakajima, Kentaro Yamada, Yuji Hasegawa, Takahiro Nakamura, Hiroshi Tsujino |
IROS | 6 |
| 2005 | A two-layer model for behavior and dialogue planning in conversational service robotsabstractThis paper presents a model for the behavior and dialogue planning module of conversational service robots. Most of the previously built conversational robots cannot perform dialogue management necessary for accurately recognizing human intentions and providing information to humans. This model integrates robot behavior planning models with spoken dialogue management that is robust enough to engage in mixed-initiative dialogues in specific domains. It has two layers; the upper layer is responsible for global task planning using hierarchical planning and the lower layer engages in local planning by utilizing modules called experts, which are specialized for performing certain kind of tasks by performing physical actions and engaging in dialogues. This model enables switching and canceling tasks based on recognized human intentions. A preliminary implementation of the model, which has been integrated with Honda ASIMO, has shown its effectiveness. Mikio Nakano, Yuji Hasegawa, Kazuhiro Nakadai, Takahiro Nakamura, Johane Takeuchi, Toyotaka Torii, Hiroshi Tsujino, Naoyuki Kanda, Hiroshi G. Okuno |
IROS | 7 |
| 2005 | Learning to estimate user interest utilizing the variational Bayes estimatorabstractMany studies of man-machine interaction using eye trackers have been tackled over recent decades. In this paper, we present a new learning system to estimate user interest with gaze sensory information. In short, a statistical learning scheme, especially the variational Bayes (VB), is incorporated for building probabilistic model parameters, dealing with the uncertainty of estimated user interest. Several computational results show how the VB can cope with user interest estimation, by selectively modeling their uncertainty. Taiji Suzuki, Takamasa Koshizen, Kazuyuki Aihara, Hiroshi Tsujino |
ISDA | 4 |
| 2004 | Improvement of Robot Audition by Interfacing Sound Source Separation and Automatic Speech Recognition with Missing Feature TheoryabstractWe have been developed robot audition system using the active direction-pass filter (ADPF) with the Scattering Theory, and demonstrated that the humanoid SIG could separate and recognize three simultaneous speeches originating from different directions. This is the first result that a robot can listen to several things simultaneously. However, its general applicability to other robots is not yet confirmed. Since automatic speech recognition (ASR) requires direction- and speaker-dependent acoustic models, it is difficult to adapt various kinds of environments. In addition ASR with lots of acoustic models causes slow processing. In this paper, these three problems are resolved. First, we confirmed the generality of the ADPF by applying it to two humanoids, SIG2 and Replie, under different environments. Next, we present the new interface between ADPF and ASR based on the Missing Feature Theory, which masks broken features of separated sound to make them unavailable to ASR. This new interface improved the recognition performance of three simultaneous speeches up to about 90%. Finally, since the ASR uses only a single acoustic model that is direction- and speaker-independent and created under clean environments, the processing of the whole system was made very light and fast. Shun'ichi Yamamoto, Kazuhiro Nakadai, Hiroshi Tsujino, Toshio Yokoyama, Hiroshi G. Okuno |
ICRA | 3 |
| 2004 | Multimodal expression for humanoid robots by integration of human speech mimicking and facial color
Tokitomo Ariyoshi, Kazuhiro Nakadai, Hiroshi Tsujino |
INTERSPEECH | 3 |
| 2004 | Assessment of general applicability of robot audition system by recognizing three simultaneous speechesabstractRobot audition is a critical technology in creating an intelligent robot operating in daily environments. We have developed such a robot audition system by using a new interface between sound source separation and automatic speech recognition (ASR). A mixture of speeches captured with a pair of microphones installed in the ear positions of a humanoid is separated into each speech by using active direction-pass filter (ADPF). The ADPF extracts a sound source originating from a specific direction in real-time by using interaural phase and intensity differences. The separated speech is recognized by a speech recognizer based on the missing feature theory (MFT). By using a missing feature mask, the MFT based ASR neglects distorted and missing features caused during the speech separation. A missing feature mask for each separated speech is generated in speech separation and is sent to the ASR with the separated speech. Thus, this new integration improves the performance of ASR. However, the generality of this robot audition system has not been assessed so far. In this paper, we assess its general applicability by implementing it on the three humanoids, i.e., ASIMO of Honda, SIG2, and Replie of Kyoto University. By using three simultaneous speeches as benchmarks, the robot audition system improved the performance of ASR over 50% in every humanoid, and thus its general applicability was confirmed. Shun'ichi Yamamoto, Kazuhiro Nakadai, Hiroshi Tsujino, Hiroshi G. Okuno |
IROS | 3 |
| 2004 | Expectation maximization of prefrontal-superior temporal network by indicator component-based approach
Takamasa Koshizen, Bernd Heisele, Hiroshi Tsujino |
Neurocomputing | 3 |
| 2004 | Improvement of recognition of simultaneous speech signals using AV integration and scattering theory for humanoid robots
Kazuhiro Nakadai, Daisuke Matsuura, Hiroshi G. Okuno, Hiroshi Tsujino |
Speech Commun. | 4 |
| 2003 | Three simultaneous speech recognition by integration of active audition and face recognition for humanoid
Kazuhiro Nakadai, Daisuke Matsuura, Hiroshi G. Okuno, Hiroshi Tsujino |
INTERSPEECH | 4 |
| 2003 | Active Detection of Anomalous Region as Primitive Processing for Visual Object Segregation
Shinichi Nagai, Koji Akatsuka, Tetsuya Ido, Hiroshi Kondo, Atsushi Miura, Hiroshi Tsujino |
KES | 6 |
| 2003 | Semantic rewiring mechanism of neural cross-supramodal integration based on spatial and temporal properties of attention
Takamasa Koshizen, So Yamada, Hiroshi Tsujino |
Neurocomputing | 3 |
| 2002 | A computational model of attentive visual system induced by cortical neural network
Takamasa Koshizen, Koji Akatsuka, Hiroshi Tsujino |
Neurocomputing | 3 |
| 1998 | Evaluation of the PCA Encoding Ability in Speech Recognition
Koji Akatsuka, Hiroshi Tsujino |
ICONIP | 2 |
| 1997 | A Cortical-type Modular Neural Network for Hypothetical Reasoning
Edgar Körner, Hiroshi Tsujino, Tomohiko Masutani |
Neural Networks | 2 |