EDBT 2026 Demo / reviewers in the wild / expert
Yuji Hasegawa
dblp:01/1913
· DBLP profile ↗
23ranked-venue papers
0as first author
2since 2021 · last 2022
0000-0001-5175-0408ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 2 since 2021Systems, architecture and hardware · 17 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Motion planning and robot control · 52% Robot navigation and mapping · 19% Reinforcement learning · 19% | |
| Computer graphics and multimedia
1 paper |
Audio and music processing · 100% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Robot navigation and mapping › social navigation
crowd navigation |
0.6 | 1 | 2022 | Learning Crowd-Aware Robot Navigation from Challenging Environments via Distributed Deep Reinforcement Learning · ICRA 2022 |
Machine learning › Reinforcement learning › deep reinforcement learning
deep reinforcement learning for navigation |
0.6 | 1 | 2022 | Learning Crowd-Aware Robot Navigation from Challenging Environments via Distributed Deep Reinforcement Learning · ICRA 2022 |
Robotics › Motion planning and robot control
robot learning |
0.6 | 1 | 2022 | Learning Crowd-Aware Robot Navigation from Challenging Environments via Distributed Deep Reinforcement Learning · ICRA 2022 |
Robotics › Motion planning and robot control
collision avoidance |
0.5 | 1 | 2021 | Dynamic Window Approach with Human Imitating Collision Avoidance · ICRA 2021 |
Robotics › Motion planning and robot control › motion planning › reactive motion generation
dynamic window approach |
0.5 | 1 | 2021 | Dynamic Window Approach with Human Imitating Collision Avoidance · ICRA 2021 |
Natural language and speech › Speech recognition and synthesis › speech separation › computational auditory scene analysis
robot audition |
0.2 | 2 | 2010 | A hybrid framework for ego noise cancellation of a robot · ICRA 2010 A robot referee for rock-paper-scissors sound games · ICRA 2008 |
Audio and music processing › source separation
blind source separation |
0.1 | 1 | 2010 | Blind Source Separation With Parameter-Free Adaptive Step-Size Method for Robot Audition · IEEE Trans. Speech Audio Process. 2010 |
Audio and music processing
source separation |
0.1 | 1 | 2010 | Blind Source Separation With Parameter-Free Adaptive Step-Size Method for Robot Audition · IEEE Trans. Speech Audio Process. 2010 |
Audio and music processing › adaptive filtering
step-size control |
0.1 | 1 | 2010 | Blind Source Separation With Parameter-Free Adaptive Step-Size Method for Robot Audition · IEEE Trans. Speech Audio Process. 2010 |
Natural language and speech › Speech recognition and synthesis
sound source separation |
0.1 | 1 | 2008 | A robot referee for rock-paper-scissors sound games · ICRA 2008 |
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.0 | 1 | 2010 | A hybrid framework for ego noise cancellation of a robot · ICRA 2010 |
Audio and music processing › computational auditory scene analysis
robot audition |
0.0 | 1 | 2010 | Blind Source Separation With Parameter-Free Adaptive Step-Size Method for Robot Audition · IEEE Trans. Speech Audio Process. 2010 |
Human-robot interaction › robot communication
dialogue system |
0.0 | 1 | 2008 | A robot referee for rock-paper-scissors sound games · ICRA 2008 |
Methods — techniques the papers use, named apart from their topics
distributed reinforcement learning · 0.6deep reinforcement learning · 0.6ape-x · 0.6imitation learning · 0.5deep learning · 0.5microphone array processing · 0.2microphone array · 0.2geometric source separation · 0.2template subtraction · 0.1source separation · 0.1independent component analysis · 0.1missing feature theory · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Learning Crowd-Aware Robot Navigation from Challenging Environments via Distributed Deep Reinforcement LearningabstractThis paper presents a deep reinforcement learning (DRL) sframework for safe and efficient navigation in crowded environments. Here, the robot learns cooperative behavior using a new reward function that penalizes robot actions interfering with the pedestrian's movement. Also, we propose a simulated pedestrian policy reflecting data from actual pedestrian movements. Furthermore, we introduce a collision detection that considers the pedestrian's personal space to generate affinity robot behavior. To efficiently explore this simulation environment, we propose distributed learning using Ape-X [1]. We deployed the robot in a real environment and verified its crowd-aware navigation performance compared with an actual human in terms of path length, travel time, and the number of abrupt avoidances. Sango Matsuzaki, Yuji Hasegawa |
ICRA | 2 |
| 2021 | Dynamic Window Approach with Human Imitating Collision AvoidanceabstractThe autonomous navigation in the crowded environment is a challenging task due to the sensor occlusion and the complex nature of the abstract social interactions. And yet, humans are capable of navigating in such complex environment. In this paper, we propose an effective navigation method that combines the learning-based and model-based methods in a way that a cost function that includes human imitation factor learned via deep learning is integrated into the dynamic window approach (DWA) [1]. The experiments conducted on simulations show that by training the robot to imitate the human trajectory, our navigation method is safer and more efficient than the state-of-the-art methods. Additionally, we successfully deployed a physical robot in an actual environment, and we validate that our navigation quality shares similar tendency with human in the path length, travel time, and the collision avoidance. Sango Matsuzaki, Shinta Aonuma, Yuji Hasegawa |
ICRA | 3 |
| 2011 | A multi-expert model for dialogue and behavior control of conversational robots and agents
Mikio Nakano, Yuji Hasegawa, Kotaro Funakoshi, Johane Takeuchi, Toyotaka Torii, Kazuhiro Nakadai, Naoyuki Kanda, Kazunori Komatani, Hiroshi G. Okuno, Hiroshi Tsujino |
Knowl. Based Syst. | 2 |
| 2010 | A hybrid framework for ego noise cancellation of a robotabstractNoise generated due to the motion of a robot is not desired, because it deteriorates the quality and intelligibility of the sounds recorded by robot-embedded microphones. It must be reduced or cancelled to achieve automatic speech recognition with a high performance. In this work, we divide ego-motion noise problem into three subdomains of arm, leg and head motion noise, depending on their complexity and intensity levels. We investigate methods that make use of single-channel and multi-channel processing in order to suppress ego noise separately. For this purpose, a framework consisting of a microphone-array-based geometric source separation, a consequent post filtering process and a parallel module for template subtraction is used. Furthermore, a control mechanism is proposed, which is based on signal-to-noise ratio and instantaneously detected motions, to switch to the most suitable method to deal with the current type of noise. We evaluate the proposed techniques on a humanoid robot using automatic speech recognition (ASR). The preliminary results of isolated word recognition show the effectiveness of our methods by increasing the word correct rates up to 50% compared to the single channel recognition in arm and leg motion noises and up to 25% in very strong head motion noises. Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Yuji Hasegawa, Hiroshi Tsujino, Jun-ichi Imura |
ICRA | 4 |
| 2010 | Sound source separation and automatic speech recognition for moving sourcesabstractThis paper addresses sound source separation and speech recognition for moving sound sources. Real-world applications such as robots should cope with both moving and stationary sound sources. However, most studies assume only stationary sound sources. We introduce three key techniques to cope with moving sources, that is, Adaptive Step-size control (AS), Optima Controlled Recursive Average (OCRA), and Separation Parameter Switching (SPS). We implemented a real-time robot audition system with these techniques for our humanoid robot with an 8ch microphone array by using HARK which is our open-source software for robot audition. Preliminary results show that the performance of recognition of moving sound sources improved drastically, and also the performance of the system is shown through two speech dialog scenarios which requires sound source separation and automatic speech recognition for moving sources. Kazuhiro Nakadai, Hirofumi Nakajima, Gökhan Ince, Yuji Hasegawa |
IROS | 4 |
| 2010 | An easily-configurable robot audition system using Histogram-based Recursive Level EstimationabstractThis paper presents an easily-configurable robot audition system using the Histogram-based Recursive Level Estimation (HRLE) method. In order to achieve natural human-robot interaction, a robot should recognize human speeches even if there are some noises and reverberations. Since the precision of automatic speech recognizers (ASR) have been degraded by such interference, many systems applying speech enhancement processes have been reported. However, performance of most reported systems suffer from acoustical environmental changes. For example, an enhancement process optimized for steady-state noise, such as fan noise, yields low performance when the process is used for non-steady-state noises, such as background music. The primary reason is mismatches of parameters because the appropriate parameters change according to the acoustical environments. To solve this problem, we propose a robot audition system that optimizes parameters adaptively and automatically. Our system applies and non-linear enhancement sub-processes. For the linear sub-process, we used Geometric Source Separation with the Adaptive Step-size method (GSS-AS). This adjusts the parameters adaptively and does not have any manual parameters. For the non-linear sub-process, we applied a spectral subtraction-based enhancement method with the HRLE method that is newly introduced in this paper. Since HRLE controls the threshold level parameter implicitly based on the statistical characteristics of noise and speech levels, our system has high robustness against acoustical environmental changes. For robot audition systems, all processes should be performed in real-time. We also propose implementation techniques to make HRLE run in real-time and show the effectiveness. We evaluate performance of our system and compare it to conventional systems based on the Minima Controlled Recursive Average (MCRA) method and Minimum Mean Square Error (MMSE) method. The experimental results show that our system achieves better performance than the conventional systems. Hirofumi Nakajima, Gökhan Ince, Kazuhiro Nakadai, Yuji Hasegawa |
IROS | 4 |
| 2010 | Blind Source Separation With Parameter-Free Adaptive Step-Size Method for Robot AuditionabstractThis paper proposes an adaptive step-size method for blind source separation (BSS) suitable for robot audition systems. The design of the step-size parameter is a critical consideration when we apply BSS to real-world applications such as robot audition systems, because the surrounding environment dynamically changes in the real world. It is common to use a fixed step-size parameter that was obtained empirically. However, because of environmental changes and noise, the performance of BSS with a fixed step-size parameter deteriorates and the separation matrix sometimes diverges. Several adaptive step-size methods for BSS have been proposed. However, there are difficulties when applying them to robot audition systems for example, low-computational cost requirements, being free from manual parameter adjustment and so on. We propose an adaptive step-size method suitable for robot audition systems. The proposed method has the following merits: 1) low computational cost; 2) no parameters to be adjusted manually; and 3) no additional preprocessing requirements. We applied our method to six different BSS algorithms for an eight-channel microphone array embedded in Honda's ASIMO robot. The method improved the performance of all six algorithms in experiments on separation and recognition of simultaneous speech. Moreover, the method increased the amount of calculation by less than 10% compared with the original calculation used in most BSS algorithms. Hirofumi Nakajima, Kazuhiro Nakadai, Yuji Hasegawa, Hiroshi Tsujino |
IEEE Trans. Speech Audio Process. | 3 |
| 2009 | Sound source separation of moving speakers for robot auditionabstractThis paper addresses sound source separation and speech recognition for moving sound sources. Real-world applications such as robots should cope with both moving and stationary sound sources. However, most studies assume only stationary sound sources. We introduce two key techniques to cope with moving sources, that is, Adaptive Step-size control (AS) and Optima Controlled Recursive Average (OCRA) to improve blind source separation. We implemented a real-time robot audition system with these techniques for our humanoid robot ASIMO with an 8ch microphone array by using HARK which is our open-source software for robot audition. The performance of the system will be shown through sound source separation for moving sources and automatic speech recognition of separated speeches. Kazuhiro Nakadai, Hirofumi Nakajima, Yuji Hasegawa, Hiroshi Tsujino |
ICASSP | 3 |
| 2009 | Ego noise suppression of a robot using template subtractionabstractWhile a robot is moving, the joints inevitably generate noise due to its motors, i.e. ego-motion noise. This problem is very crucial, especially in humanoid robots, because it tends to have a lot of joints and the motors are located closer to the microphones than the sound sources. In this work, we investigate methods for the prediction and suppression of the ego-motion noise. In the first part, we analyze the performance of different noise subtraction strategies, assuming that the noise prediction problem has been solved. In the second part, we present some results for a noise prediction scheme based on the current robot joint status. Performance is evaluated for a number of criteria, including Automatic Speech Recognition (ASR). We demonstrate that our method improves recognition performance during ego-motion considerably. Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Yuji Hasegawa, Hiroshi Tsujino, Jun-ichi Imura |
IROS | 4 |
| 2009 | Real-time sound source orientation estimation using a 96 channel microphone arrayabstractThis paper proposes real-time sound source orientation estimation based on orientation-extended amplitude beamforming (OE-ABF). To recognize a sound source orientation (such as face orientation) is an important function for a robot who can achieve natural human-robot interaction because the function is required to distinguish the human target from a robot or another person. We developed a sound source orientation system using orientation-extended beamforming (OE-BF) and showed the system worked properly at least under a specific controlled environment. However, in practical use, this system does not work properly because the system doesn't take into account the differences between the supposed model in OE-BF and in practical situations. For example, the system model supposes that there is neither noise nor reverberation, however, this is not a realistic assumption. To solve this assumption mismatch problem, we propose sound source orientation estimation based on OE-ABF, and constructed a real-time sound source orientation estimation system with the proposed method using a 96ch microphone array. Evaluation results of our proposed system show that the average error of estimated angles is lower than 5°, while the error of our previously reported system was greater than 20°. With this system, the robot is able to distinguish that the utterance target of a person standing 1m in front is itself or another person standing 0.2m to the left of the robot. This is valuable for human-robot interaction. Hirofumi Nakajima, Keiko Kikuchi, Touru Daigo, Yutaka Kaneda, Kazuhiro Nakadai, Yuji Hasegawa |
IROS | 6 |
| 2009 | Intelligent sound source localization for dynamic environmentsabstractAs robotic technology plays an increasing role in human lives, ¿robot audition¿, human-robot communication, is of great interest, and robot audition needs to be robust and adaptable for dynamic environments. This paper addresses sound source localization working in dynamic environments for robots. Previously, noise robustness and dynamic localized sound selection have been enormous issues for practical use. To correct the issues, a new localization system ¿Selective Attention System¿ is proposed. The system has four new functions: localization with Generalized EigenValue Decomposition of correlation matrices for noise robustness(¿Localization with GEVD¿), sound source cancellation and focus (¿Target Source Selection¿), human-like dynamic Focus of Attention (¿Dynamic FoA¿), and correlation matrix estimation for robotic head rotation (¿Correlation Matrix Estimation¿). All are achieved by the dynamic design of correlation matrices. The system is implemented into a humanoid robot, and the experimental validation is successfully verified even when the robot microphones move dynamically. Keisuke Nakamura, Kazuhiro Nakadai, Futoshi Asano, Yuji Hasegawa, Hiroshi Tsujino |
IROS | 4 |
| 2008 | Adaptive step-size parameter control for real-world blind source separationabstractThis paper describes a method to adaptively control a step-size parameter which is used for updating a separation matrix to extract a target sound source accurately in blind source separation (BSS). The design of the step-size parameter is essential when we apply BSS to real-world applications such as robot audition systems, because the surrounding environment dynamically changes in the real world. It is common to use a fixed step-size parameter that is obtained empirically. However, due to environmental changes and noises, the performance of BSS with the fixed step-size parameter deteriorates and the separation matrix sometimes diverges. We propose a general method that allows adaptive step-size control. The proposed method is an extension of Newton’s method utilizing a complex gradient theory and is applicable to any BSS algorithm. Actually, we applied it to six types of BSS algorithms for an 8 ch microphone array embedded in Honda ASIMO. Experimental results show that the proposed method improves the performance of these six BSS algorithms through experiments of separation and recognition for two simultaneous speeches. Hirofumi Nakajima, Kazuhiro Nakadai, Yuji Hasegawa, Hiroshi Tsujino |
ICASSP | 3 |
| 2008 | A robot referee for rock-paper-scissors sound gamesabstractThis paper describes a robot referee for “rockpaper-scissors (RPS)” sound games; the robot decides the winner from a combination of rock, paper and scissors uttered by two or three people simultaneously without using any visual information. In this referee task, the robot has to cope with speech with low signal-to-noise ratio (SNR) due to a mixture of speeches, robot motor noises, and ambient noises. Our robot referee system, thus, consists of two subsystems - a real-time robot audition subsystem and a dialog subsystem focusing on RPS sound games. The robot audition subsystem can recognize simultaneous speeches by exploiting two key ideas; preprocessing consisting of sound source localization and separation with a microphone array, and system integration based on missing feature theory (MFT). Preprocessing improves the SNR of a target sound signal using geometric source separation with a multi-channel post-filter. MFT uses only reliable acoustic features in speech recognition and masks out unreliable parts caused by interfering sounds and preprocessing. MFT thus provides smooth integration between preprocessing and automatic speech recognition. The dialog subsystem is implemented as a system-initiative dialog system for multiple players based on deterministic finite automata. It first waits for a trigger command to start an RPS sound game, controls the dialog with players in the game, and finally decides the winner of the game. The referee system is constructed for Honda ASIMO with an 8-ch microphone array. In the case with two players, we attained a 70% task completion rate for the games on average. Kazuhiro Nakadai, Shun'ichi Yamamoto, Hiroshi G. Okuno, Hirofumi Nakajima, Yuji Hasegawa, Hiroshi Tsujino |
ICRA | 5 |
| 2008 | A robot uses its own microphone to synchronize its steps to musical beats while scatting and singingabstractMusical beat tracking is one of the effective technologies for human-robot interaction such as musical sessions. Since such interaction should be performed in various environments in a natural way, musical beat tracking for a robot should cope with noise sources such as environmental noise, its own motor noises, and self voices, by using its own microphone. This paper addresses a musical beat tracking robot which can step, scat and sing according to musical beats by using its own microphone. To realize such a robot, we propose a robust beat tracking method by introducing two key techniques, that is, spectro-temporal pattern matching and echo cancellation. The former realizes robust tempo estimation with a shorter window length, thus, it can quickly adapt to tempo changes. The latter is effective to cancel self noises such as stepping, scatting, and singing. We implemented the proposed beat tracking method for Honda ASIMO. Experimental results showed ten times faster adaptation to tempo changes and high robustness in beat tracking for stepping, scatting and singing noises. We also demonstrated the robot times its steps while scatting or singing to musical beats. Kazumasa Murata, Kazuhiro Nakadai, Kazuyoshi Yoshii, Ryu Takeda, Toyotaka Torii, Hiroshi G. Okuno, Yuji Hasegawa, Hiroshi Tsujino |
IROS | 7 |
| 2008 | High performance sound source separation adaptable to environmental changes for robot auditionabstractThis paper describes a novel sound source separation method for a robot that needs to cope with dynamically changing noises in the real world. The sound source separation method, Geometric Source Separation (GSS), is promising because it has high separation performance and requires low computational cost. One of the most important factors in GSS performance is a step-size parameter to update a separation matrix which is generally used for extracting a target sound source. A fixed value that was obtained empirically is commonly used as the step-size parameter. However, in the real world, the surrounding environment changes dynamically. Thus, conventional GSS with a fixed step-size parameter sometimes results in poor separation results, or divergence of the separation matrix. Another important factor is the weight parameter, which adjusts the balance between geometric errors and separation errors and also affects performance. If this parameter is set to a small value, GSS becomes similar to a Blind Source Separation method, by which the output signal may contain errors based on indefinite source amplitudes and orders. In contrast, if this parameter is set to a large value, GSS becomes similar to a delay-and-sum Beamforming method, which does not have high separation performance. GSS gives good performance when the parameters are tuned to an optimum value, which changes according to the environment. We propose two effective methods that can be used for general BSS’s. One is an adaptive step-size parameter control method. By using this method, the step-size and the weight parameters are automatically set to optimum values and are able to adapt to environmental changes. The other is an optima controlled recursive average method for correlation matrix estimation. This method can improve the estimation of a separation matrix, and achieve high separation performance. We evaluated the proposed GSS algorithm with an 8ch microphone array embedded in Honda ASIMO. Experimental results showed that the proposed method improved sound source separation even in dynamically changing environments. Hirofumi Nakajima, Kazuhiro Nakadai, Yuji Hasegawa, Hiroshi Tsujino |
IROS | 3 |
| 2007 | Robust acquisition and recognition of spoken location names by domestic robotsabstractThis paper presents a method that enables a conversational domestic robot to learn location names through speech interaction. Each acquired name is associated with a point on the map coordinate system of the robot. Both for acquisition and recognition of location names, a bag- of-words-based categorization technique is used. Namely, the robot acquires a location name as a frequency pattern of words, and recognizes a spoken location name by computing similarity between the patterns. This makes the robot robust not only against speech recognition errors but also against out- of-vocabulary names. We designed a dialogue and behavior management subsystem that learns location names by using our proposed method and navigates to indicated locations, and implemented the subsystem on an omnidirectional cart robot. The result of a preliminary evaluation of the implemented robot with human subjects suggested this approach is promising. Kotaro Funakoshi, Mikio Nakano, Toyotaka Torii, Yuji Hasegawa, Hiroshi Tsujino, Noriyuki Kimura, Naoto Iwahashi |
IROS | 4 |
| 2007 | Fast decision making of autonomous robot under dynamic environment by sampling real-time Q-MDP value methodabstractIn this paper, the sampling real-time QMDP value method is proposed and applied to a goalkeeper task of a soccer robot. Soccer is a challenging task for an autonomous robot that has poor computing resources and sensors, and a good subject in the study of decision making in dynamic environments. A robot frequently decides its behavior without enough observation of the environment. In the proposed method, the risk of a score is solved beforehand toward every set of position of the goalkeeper and motion of the ball. The shortage of observation is calculated and represented by particle filters. The proposed method uses the risk function and the particle filters so as to choose appropriate behavior of the goalkeeper. The experiment with an actual robot suggests that the method can decide actions reflexively toward the motion of the ball. Yoshiaki Jitsukawa, Ryuichi Ueda, Tamio Arai, Kazutaka Takeshita, Yuji Hasegawa, Shota Kase, Takashi Okuzumi, Kazunori Umeda, Hisashi Osumi |
IROS | 5 |
| 2007 | A biped robot that keeps steps in time with musical beats while listening to music with its own earsabstractWe aim at enabling a biped robot to interact with humans through real-world music in daily-life environments, e.g., to autonomously keep its steps (stamps) in time with musical beats. To achieve this, the robot should be able to robustly predict the beat times in real time while listening to musical performance with its own ears (head-embedded microphones). However, this has not previously been addressed in most studies on music-synchronized robots due to the difficulty in predicting the beat times in real-world music. To solve this problem, we implemented a beat-tracking method developed in the field of music information processing. The predicted beat times are then used by a feedback-control method that adjusts the robot's step intervals to synchronize its steps in time with the beats. The experimental results show that the robot can adjust its steps in time with the beat times as the tempo changes. The resulting robot needed about 25 [s] to recognize the tempo change after it and then synchronize its steps. Kazuyoshi Yoshii, Kazuhiro Nakadai, Toyotaka Torii, Yuji Hasegawa, Hiroshi Tsujino, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 4 |
| 2007 | A markup language for describing interactive humanoid robot presentationsabstractThis paper presents a multi-modal presentation markup language for humanoid robots, MPML-HR ver. 3.0, which is able to describe presentation contents including speech-based interactions with audiences. Previous versions of MPML-HR do not feature any interaction functionality which dynamically changes the presentation according to the utterances by audiences, although such interaction makes the presentation more effective and understandable. Since MPML-HR ver. 3.0 inherits simple descriptions of previous versions of MPML-HR, the content designer can describe interactive presentations without configuring conventional complicated multi-modal interactive systems. Yoshitaka Nishimura, Shinichiro Minotsu, Hiroshi Dohi, Mitsuru Ishizuka, Mikio Nakano, Kotaro Funakoshi, Johane Takeuchi, Yuji Hasegawa, Hiroshi Tsujino |
IUI | 8 |
| 2006 | Robust Tracking of Multiple Sound Sources by Spatial Integration of Room And Robot Microphone ArraysabstractSound source tracking is an important function for a robot operating in a daily environment, because the robot should recognize where a sound event such as speech, music and other environmental sounds originates from. This paper addresses sound source tracking by integrating a room and a robot microphone array. The room microphone array consists of 64 microphones attached to the walls. It provides 2D (x-y) sound source localization based on a weighted delay-and-sum beamforming method. The robot microphone array consists of eight microphones installed on a robot head, and localizes multiple sound sources in azimuth. The localization results are integrated to track sound sources by using a particle filter for multiple sound sources. The experimental results show that particle filter based integration reduces localization errors and provides accurate and robust 2D sound source tracking. Kazuhiro Nakadai, Hirofumi Nakajima, Masamitsu Murase, Satoshi Kaijiri, Kentaro Yamada, Takahiro Nakamura, Yuji Hasegawa, Hiroshi G. Okuno, Hiroshi Tsujino |
ICASSP (4) | 7 |
| 2006 | Real-Time Tracking of Multiple Sound Sources by Integration of In-Room and Robot-Embedded Microphone ArraysabstractReal-time and robust sound source tracking is an important function for a robot operating in a daily environment, because the robot should recognize where a sound event such as speech, music and other environmental sounds originate from. This paper addresses real-time sound source tracking by real-time integration of an in-room microphone array (IRMA) and a robot-embedded microphone array (REMA). The IRMA system consists of 64 ch microphones attached to the walls. It localizes multiple sound sources based on weighted delay-and-sum beam-forming on a 2D plane. The REMA system localizes multiple sound sources in azimuth using eight microphones attached to a robot's head on a rotational table. The localization results are integrated to track multiple sound sources by using a particle filter in real-time. The experimental results show that particle filter based integration improved accuracy and robustness in multiple sound source tracking even when the robot's head was in rotation Kazuhiro Nakadai, Hirofumi Nakajima, Masamitsu Murase, Hiroshi G. Okuno, Yuji Hasegawa, Hiroshi Tsujino |
IROS | 5 |
| 2005 | Sound source tracking with directivity pattern estimation using a 64 ch microphone arrayabstractIn human-robot communication, a robot should distinguish between voices uttered by a human and those played by a loudspeaker such as on a TV or a radio. This paper addresses detection of actual human voices by using a microphone array as an extension of auditory function of the robot to support environmental understanding by the robot. We introduce a 64 ch microphone array system in a room and propose a new method based on weighted delay-and-sum beamforming to estimate a directivity pattern of a sound source. The microphone array system localizes a sound source and estimates its directivity pattern. The directivity pattern estimation has two advantages as follows: One is that the system can detect whether the sound source is an actual human voice or not by comparing the estimated directivity pattern with prerecorded directivity patterns. The other is that the heading of the sound source is estimated by detecting the angle with the highest power in the directivity pattern. As a result, we proved the effectiveness of our microphone array through sound source tracking with orientation and detection of actual human voices based on directivity pattern estimation. Kazuhiro Nakadai, Hirofumi Nakajima, Kentaro Yamada, Yuji Hasegawa, Takahiro Nakamura, Hiroshi Tsujino |
IROS | 4 |
| 2005 | A two-layer model for behavior and dialogue planning in conversational service robotsabstractThis paper presents a model for the behavior and dialogue planning module of conversational service robots. Most of the previously built conversational robots cannot perform dialogue management necessary for accurately recognizing human intentions and providing information to humans. This model integrates robot behavior planning models with spoken dialogue management that is robust enough to engage in mixed-initiative dialogues in specific domains. It has two layers; the upper layer is responsible for global task planning using hierarchical planning and the lower layer engages in local planning by utilizing modules called experts, which are specialized for performing certain kind of tasks by performing physical actions and engaging in dialogues. This model enables switching and canceling tasks based on recognized human intentions. A preliminary implementation of the model, which has been integrated with Honda ASIMO, has shown its effectiveness. Mikio Nakano, Yuji Hasegawa, Kazuhiro Nakadai, Takahiro Nakamura, Johane Takeuchi, Toyotaka Torii, Hiroshi Tsujino, Naoyuki Kanda, Hiroshi G. Okuno |
IROS | 2 |