Keisuke Nakamura

dblp:85/2131 · DBLP profile ↗
← Back
65ranked-venue papers
13as first author
12since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 54 · 12 first-author · 12 since 2021Systems, architecture and hardware · 34 · 8 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 19 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Computer networks · 1
YearPublicationVenuePosition
2024 Exceptions with Priorities by Ethical Values for Logic Programming in Natural Language
abstract
We developed a system that can perform predicate logic operations similar to PROLOG based on natural language. We then made it possible to express and reason about "strong negation" (taboo expressions) that is morally or scientifically forbidden. This is different from weak negation, i.e., the impossibility of proving in the closed world of PROLOG, and was very useful for describing ethical taboo knowledge that constitutes big data, but it did not support exceptions to taboos or taboos of those exceptions. In this paper, we propose a knowledge representation method that can simply describe the priority switching and logical integration of the "strong negation" (level 1), affirmation as the "stronger" exception (level 2), and negation as the much "stronger" exception(level 3) of that exception (level 2), etc., with priority. We implemented the improved system that can be interpreted and inferred by a computer, and proved its practical use through simple experiments.
Keisuke Nakamura, Narumi Naya, Tatsuyoshi Ando
IEEE Big Data1
2023 GAN-Based Interactive Reinforcement Learning from Demonstration and Human Evaluative Feedback
abstract
Generative adversarial imitation learning (GAIL) — a general model-free imitation learning method, allows robots to directly learn policies from expert trajectories in large environments. However, GAIL shares the limitation of other imitation learning methods that they can seldom surpass the performance of demonstrations. In this paper, to address the limit of GAIL, we propose GAN-based interactive reinforcement learning (GAIRL) from demonstrations and human evaluative feedback, by combining the advantages of GAIL and interactive reinforcement learning. We test GAIRL in six physics-based control tasks, ranging from simple low-dimensional control tasks — Cart Pole, Mountain Car and Lunar Lander, to difficult high-dimensional tasks — Inverted Double Pendulum, Hopper and HalfCheetah. Our results suggest that, the GAIRL agent can generally surpass the performance of demonstrations in both low-dimensional and high-dimensional tasks and get an optimal or close to optimal policy.
Jiangshan Hao, Rongshun Juan, Randy Gomez, Keisuke Nakamura, Guangliang Li
ICRA5
2023 Sim-to-Real Policy and Reward Transfer with Adaptive Forward Dynamics Model
abstract
Deep reinforcement learning has shown promise in learning robust skills for robot control, but typically requires a large amount of samples to achieve good performance. Sim-to-real transfer learning has been developed to solve this problem, but the policy trained in simulation usually has unsatisfactory performance in the real world because simulators inevitably model the dynamics of reality imperfectly. To enable sample-efficient learning in the real world, we proposed progressive policy transfer with adaptive dynamics model (PPTADM). PPTADM assumes the dynamics of simulation and real world do not match but the state space is the same, transfers policy from simulation via progressive neural network (PNN) and further improves the policy with a learned forward dynamics model in reality. In addition, for real-world tasks in which reward functions are difficult or even impossible to define and verify the effectiveness, PPTADM can learn in real world solely from a transferred reward function that is estimated from simulation even though their dynamics do not match. Our results in five simulated tasks and on a real robot arm show that with PPTADM, the robot's learning efficiency and performance in the real world can be significantly improved.
Rongshun Juan, Hao Ju 0003, Randy Gomez, Keisuke Nakamura, Guangliang Li
ICRA5
2022 A TABOO-NOT in OPEN World Assumption for A Natural Language based Logic Programming
abstract
Traditional logic programs are written in some formal languages easy to “unify” each other or symbol based ones such as in Prolog, and their “NOT” mechanisms are “negation by failure of proof” according to “CLOSED world assumption” and are not assumed as “true negation”s.On the other hand, we suppose that BigData for moral, ethics and values are written originally in various different natural languages over the world.In order for both human and AI to easily understand and automatically operate such BigData, we hope that (1) human could logic-program the contents of such Big-Data, by semi-dead-copying the original-natural-language-texts as "logic-programs" without hard translation, formalization and/or normalization,(2) “NOT”s in the language text that mean some TABOOs in any culture shall be operated correctly as such (not as “negataion by failure of proof”s but as the TABOOs) even in case of concretizing objects from different cultures, as in OPEN world assumption, because any TABOO from any one culture shall be rigidly respected, even during meetings between people from different cultures including the one culture with the TABOO.For the above purpose, this document explains a method of natural-language-based-logic-programming where variables for knowledge abstraction and concretization are embedded into the natural language texts, and also explains the above strong "TABOO-NOT" implementation therefor.
Keisuke Nakamura, Tatsuyoshi Ando
IEEE Big Data1
2022 An Automatically Inferable Format for Creating Big Data of Morals, Ethics and ORDER OF VALUES Written Almost in Natural Language
abstract
If there is a format that is easy for both humans and AI (computers) to understand and infer, and also if the humans let the AI learn various morals, ethics and values in the above format, these data may be operated automatically by the AI in a way that the humans concerned are able to check for mistakes / lacks of thoughts in the automatic judgments by the AI, avoiding arbitrary thoughts / judgments by both the providers(humans) of such data and the AI.On the other hand, morals, etc. are diverse depending on the country, history, culture, social position, and/or individual. So in order to accurately express such data for reducing interpretation mistakes, lacks of required conditions of thought-rules and so on, we need to allow, as such data, expressions almost in (or from completely including all) the natural languages used in each cultural area.In this paper, we (1) propose a format based on natural language that we called as "situation-specific-value-order-formula" as a starting point for an expression format that satisfy the above needs, (2) explain its components: situations / viewpoints and items ( positive / 0 / negative ) ordered by the value all based on natural language, (3) explain related things including the specific l ogical i nference f eature o f t he n atural language based programming where variables are embedded in the natural language.We conclude that our proposed format is feasible and practical to some extent for the above Purpose and need.
Keisuke Nakamura, Tatsuyoshi Ando
IEEE Big Data1
2022 Developing The Bottom-up Attentional System of A Social Robot
abstract
This paper describes the development of a 3- stage signalling framework to trigger a social robot's bottom- up reactive behavior inspired by a biological model. In the first stage, low-level firing of stimuli due to external sources is constructed through perception grounding. This is followed by a saliency classifier which fires-up high level salient signals that require attention and are used to trigger the robot's reactive behavior. The whole framework evolves primarily on the knowledge ontology that defines the characteristics of the social robot and the querying mechanism that correlates the perceived stimuli with the ontology to trigger the reactive behavior. We evaluated the performance of our system with timing metrics and we achieved good results for our application.
Randy Gomez, Álvaro Páez, Yu Fang 0007, Serge Thill, Luis Merino, Eric Nichols, Keisuke Nakamura, Heike Brock
ICRA7
2022 Affective Behavior Learning for Social Robot Haru with Implicit Evaluative Feedback
abstract
We propose a human-in-the-loop reinforcement learning mechanism to help robots learn emotional behavior. Unlike the previous methods of providing explicit feedback via pressing keyboard buttons or mouse clicks, we provide a more natural way for ordinary people to train social robots how to perform social tasks according to their preferences - facial expressions. The whole experiment is carried out on the desktop robot Haru, which is mainly used for the research of emotion and empathy participation. Our experimental results show that through learning from implicit feedback of facial features, Haru can quickly understand and dynamically adapt to individual preferences, and obtain a similar performance to learning from explicit feedback. In addition, we observe that the recognition error of human feedback will cause a “temporary regress” of the robot's learning performance, which is more obvious at the beginning of the training process. This phenomenon is shown to be correlated with the accuracy of recognizing negative implicit feedback.
Hui Wang 0141, Jinying Lin, Yurii Vasylkiv, Heike Brock, Keisuke Nakamura, Randy Gomez, Bo He 0002, Guangliang Li
IROS6
2022 Shaping Haru's Affective Behavior with Valence and Arousal Based Implicit Facial Feedback
abstract
Social robots that are able to express emotions can potentially improve human’s well-being. Whether and how they can learn from interactions between them and human being in a natural way will be key to their success and acceptance by ordinary people. In this paper, we proposed to shape social robot Haru affective behaviors with predicted continuous rewards based on received implicit facial feedback via human-centered reinforcement learning. The implicit facial feedback was estimated with the valence and arousal of received implicit facial feedback using Russell’s circumplex model, which can provide a more accurate estimation of the subtle psychological changes of human user, resulting in more effective robot behavior learning. The whole experiment is conducted on the desktop robot Haru, which is primarily used to study emotional interactions with human in different scenarios. Our experimental results show that with our proposed method, Haru can obtain a similar performance to learning from explicit feedback, eliminating the need for human users to get familiar with training interface in advance and resulting in an unobtrusive learning process.
Hui Wang 0141, Randy Gomez, Keisuke Nakamura, Bo He 0002, Guangliang Li
RO-MAN4
2021 Automating Behavior Selection for Affective Telepresence Robot
abstract
The tabletop robot Haru, used for affective telepresence research, enables a teleoperator to communicate affects from a distance. The robot’s expressiveness offers myriad ways of communicating affects through the execution of emotive routines. The teleoperator reacts to input modalities such as the user’s facial expression, gestures and speech-based intent as perceived by the robot’s perception system. However, due to the sheer number of routines to select from, the task of choosing the appropriate or the most preferred routine is becoming cumbersome. In this paper, we propose a human-in-the-loop reinforcement learning mechanism in which an agent learns the teleoperator’s selection preference as a function of the input modalities and aids the routine selection process by narrowing it to n-best optimal choices. Our experimental results show that with only a few number of interactions from the teleoperator, the system can learn to recommend optimal routine behaviors for all perceived modalities, which greatly reduces the workload of the teleoperator.
Yurii Vasylkiv, Guangliang Li, Eleanor Sandry, Heike Brock, Keisuke Nakamura, Pourang Irani, Randy Gomez
ICRA6
2021 Shaping Progressive Net of Reinforcement Learning for Policy Transfer with Human Evaluative Feedback
abstract
Deep reinforcement learning has achieved significant success in many fields, but will confront sampling efficiency and safety problems when applying to robot control in the real world. Sim-to-real transfer learning was proposed to make use of samples in the simulation and overcome the gap between simulation and real world. In this paper, we focus on improving Progressive Neural Network — an effective sim-to-real learning method, by proposing Interactive Progressive Network Learning (IPNL). IPNL integrates progressive network and interactive reinforcement learning (interactive RL) which learns from evaluative feedback provided by an observing human trainer. We test our method using five RL tasks with discrete or continuous actions in OpenAI Gym and a sinusoids curve following task with AUV simulator on the Gazebo platform. Our results suggest that while Progressive Network has good performance when transferring from tasks with low-dimensional state space to those with high-dimensional one but has little effect for transferring from high-dimensional tasks to low-dimensional ones, IPNL allows an agent to learn a more stable policy with better performance faster for both cases. More importantly, our further analysis indicate that there is a synergy between Progressive Network and interactive RL for improving the agent’s learning. Our results in the path following of AUV shed light on the potential of applying our method in the real world tasks.
Rongshun Juan, Randy Gomez, Keisuke Nakamura, Qixin Sha, Bo He 0002, Guangliang Li
IROS4
2021 Exploring Affective Storytelling with an Embodied Agent
abstract
In this paper, we explore the storytelling potential of a robot. We exploit the use of creative contents that maximize the embodied communication affordance of the empathic robot Haru. We identify the elements in storytelling such as narration, agency, engagement and education and synthesized these into the robot. Through effective design we investigated the possible answers that could leverage the limitations and the challenges in developing storytelling applications through a robotic medium. Our preliminary findings show that the use of an embodied agent such as a robot in storytelling only has meaning when its communicative affordance (i.e. embodiment, expressiveness, and other modalities) is tapped, adding new dimension to the experience. Otherwise, traditional storytelling delivery (e.g. tablet) without the use of embodiment will suffice. Hence, robots need to be performers rather than just mere props in storytelling.
Randy Gomez, Deborah Szapiro, Kerl Galindo, Luis Merino, Heike Brock, Keisuke Nakamura, Yu Fang 0007, Eric Nichols
RO-MAN6
2021 Shaping Affective Robot Haru's Reactive Response
abstract
We describe a method of teaching a robot its empathic behavioural response from its interaction with people. We used the input modalities such as relative spatial information, facial expressions, body gestures and speech information as perception input that triggers the robot’s empathic response. First, we bootstrap the training through a pre-learning mechanism in which training is conducted by users who know the robotic system. This phase provides simulation-based training using a simple graphical user interface to simulate the input, rewards and correction feedback. In the second phase, we developed an online learning scheme for naive users to personalize their robot further, building on top of the bootstrapped model. Here, we developed a natural user interface that enables natural human-robot interaction via the suite of sensors that allows the users to provide evaluative feedback during the interaction with the robot. We evaluated the system and our results show that bootstrapping is an efficient tool to hasten the robot’s learning while online learning provided some form of personalization in the real environment with naive users.
Yurii Vasylkiv, Guangliang Li, Heike Brock, Keisuke Nakamura, Pourang Irani, Randy Gomez
RO-MAN5
2020 A Holistic Approach in Designing Tabletop Robot's Expressivity
abstract
Defining a robot's expressivity is a difficult task that requires thoughtful consideration of the potential of various robot modalities and a model of communication that humans understand. Humanoid and zoomorphic-designed robots can easily take cues from human and animals, respectively when designing their expressivity. However, a robot design that is neither human nor animal-like does not have a clear model to follow in terms of designing expressivity. Animation presents a potential model in these circumstances as animated characters in movies take various forms, sizes, shapes and styles, and are successful in defining expressivity that is widely accepted across different languages and cultures. In this paper, we discuss the development and design of the expressivity of Haru, a table top robot that is neither human nor animal-like and the application of animation expertise to the holistic treatment of the different modalities. The method maximizes animation techniques and expertise normally applied to movies to generate expressivity that is then transferred to the robot hardware. Experimental results show that the robot's expressivity generated using our method is easily understood and are preferred to the conventional approach of generating expressions.
Randy Gomez, Deborah Szapiro, Luis Merino, Keisuke Nakamura
ICRA4
2020 Robust Real-Time Hand Gestural Recognition for Non-Verbal Communication with Tabletop Robot Haru
abstract
In this paper, we present our work in close-distance non-verbal communication with tabletop robot Haru through hand gestural interaction. We implemented a novel hand gestural understanding system by training a machine-learning architecture for real-time hand gesture recognition with the Leap Motion. The proposed system is activated based on the velocity of a user's palm and index finger movement, and subsequently labels the detected movement segments under an early classification scheme. Our system is able to combine multiple gesture labels for recognition of consecutive gestures without clear movement boundaries. System evaluation is conducted on data simulating real human-robot interaction conditions, taking into account relevant performance variables such as movement style, timing and posture. Our results show robustness in hand gesture classification performance under variant conditions. We furthermore examine system behavior under sequential data input, paving the way towards seamless and natural real-time close-distance hand-gestural communication in the future.
Heike Brock, Selma Sabanovic, Keisuke Nakamura, Randy Gomez
RO-MAN3
2020 Human Social Feedback for Efficient Interactive Reinforcement Agent Learning
abstract
As a branch of reinforcement learning, interactive reinforcement learning mainly studies the interaction process between humans and agents, allowing agents to learn from the intentions of human users and adapt to their preferences. In most of the current studies, human users need to intentionally provide explicit feedback via pressing keyboard buttons or mouse clicks. However, in our paper, we proposed an interactive reinforcement learning method that facilitates an agent to learn from human social signals - facial feedback via a ordinary camera and gestural feedback via a leap motion sensor. Our method provides a natural way for ordinary people to train agents how to perform a task according to their preferences. We tested our method in two reinforcement learning benchmarking domains - LoopMaze and Tetris, and compared to the state of the art - the TAMER framework. Our experimental results show that when learning from facial feedback the recognition of which is very low, the TAMER agent can get a similar performance to that of learning from keypress feedback with slightly more feedback. When learning from gestural feedback with a more accurate recognition, the TAMER agent can obtain a similar performance to that of learning from keypress feedback with much less feedback received. Moreover, our results indicate that the recognition error of facial feedback has a large effect on the agent performance in the beginning training process than in the later training stage. Finally, our results indicate that with enough recognition accuracy, human social signals can effectively improve the learning efficiency of agents with less human feedback.
Jinying Lin, Qilei Zhang, Randy Gomez, Keisuke Nakamura, Bo He 0002, Guangliang Li
RO-MAN4
2019 Expressivity for Sustained Human-Robot Interaction
abstract
Expressivity - the use of multiple, non-verbal, modalities to convey or augment the communication of internal states and intentions - is a core component of human social interactions. Studying expressivity in contexts of artificial agents has led to explicit considerations of how robots can leverage these abilities in sustained social interactions. Research on this covers aspects such as animation, robot design, mechanics, as well as cognitive science and developmental psychology. This workshop provides a forum for scientists from diverse disciplines to come together and advance the state of the art in developing expressive robots. Participants will discuss points of methodological opportunities and limitations, to develop a shared vision for next steps in expressive social robots.
Vicky Charisi, Selma Sabanovic, Serge Thill, Emilia Gómez, Keisuke Nakamura, Randy Gomez
HRI5
2019 Generation of expressive motions for a tabletop robot interpolating from hand-made animations
abstract
Motion is an important modality for human-robot interaction. Besides a fundamental component to carry out tasks, through motion a robot can express intentions and expressions as well. In this paper, we focus on a tabletop robot in which motion, among other modalities, is used to convey expressions. The robot incorporates a set of pre-programmed motion animations that show different expressions with various intensities. These have been created by designers with expertise in animation. The objective in the paper is to analyze if these examples can be used as demonstrations, and combined by the robot to generate additional richer expressions. Challenges are the representation space used, and the scarce number of examples. The paper compares three different learning from demonstration approaches for the task at hand. A user study is presented to evaluate the resultant new expressive motions automatically generated by combining previous demonstrations.
Gonzalo Mier, Fernando Caballero, Keisuke Nakamura, Luis Merino, Randy Gomez
RO-MAN3
2019 Human-Centered Reinforcement Learning: A Survey
abstract
Human-centered reinforcement learning (RL), in which an agent learns how to perform a task from evaluative feedback delivered by a human observer, has become more and more popular in recent years. The advantage of being able to learn from human feedback for a RL agent has led to increasing applicability to real-life problems. This paper describes the state-of-the-art human centered RL algorithms and aims to become a starting point for researchers who are initiating their endeavors in human-centered RL. Moreover, the objective of this paper is to present a comprehensive survey of the recent breakthroughs in this field and provide references to the most interesting and successful works. After starting with an introduction of the concepts of RL from environmental reward, this paper discusses the origins of human-centered RL and its difference from traditional RL. Then we describe different interpretations of human evaluative feedback, which have produced many human-centered RL algorithms in the past decade. In addition, we describe research on agents learning from both human evaluative feedback and environmental rewards as well as on improving the efficiency of human-centered RL. Finally, we conclude with an overview of application areas and a discussion of future work and open questions.
Guangliang Li, Randy Gomez, Keisuke Nakamura, Bo He 0002
IEEE Trans. Hum. Mach. Syst.3
2018 Comparison of Region of Interest Segmentation Methods for Video-Based Heart Rate Measurements
abstract
Conventional contact photoplethysmography (PPG) sensors are not suitable in situations of skin damage or when unconstrained movement is required. As a consequence, remote photoplethysmography (rPPG) has recently emerged because it provides remote physiological measurements without expensive hardware and improves comfort for long term monitoring. RPPG estimation methods use the spatially averaged RGB values of pixels in a Region Of Interest (ROI) to generate a temporal RGB signal. The selection of ROI is a critical first step to obtain reliable pulse signals and must contain as many skin pixels as possible with a low percentage of non-skin pixels. In this paper, we experimentally compare seven ROI segmentation methods in the perspective of heart rate (HR) measurements with dedicated metrics. The algorithms are compared using our in-house database UBFC-RPPG, comprising of 53 videos specifically geared towards rPPG analysis.
Peixi Li, Yannick Benezeth, Keisuke Nakamura, Randy Gomez, Chao Li 0005, Fan Yang 0019
BIBE3
2018 PageFlip: Leveraging Page-Flipping Gestures for Efficient Command and Value Selection on Smartwatches
abstract
Selecting an item of interest on smartwatches can be tedious and time-consuming as it involves a series of swipe and tap actions. We present PageFlip, a novel method that combines into a single action multiple touch operations such as command invocation and value selection for efficient interaction on smartwatches. PageFlip operates with a page flip gesture that starts by dragging the UI from a corner of the device. We first design PageFlip by examining its key design factors such as corners, drag directions and drag distances. We next compare PageFlip to a functionally equivalent radial menu and a standard swipe and tap method. Results reveal that PageFlip improves efficiency for both discrete and continuous selection tasks. Finally, we demonstrate novel smartwatch interaction opportunities and a set of applications that can benefit from PageFlip.
Teng Han, Jiannan Li, Khalad Hasan, Keisuke Nakamura, Randy Gomez, Ravin Balakrishnan, Pourang Irani
CHI4
2018 Haru: Hardware Design of an Experimental Tabletop Robot Assistant
abstract
This paper discusses the design and development of an experimental tabletop robot called "Haru" based on design thinking methodology. Right from the very beginning of the design process, we have brought an interdisciplinary team that includes animators, performers and sketch artists to help create the first iteration of a distinctive anthropomorphic robot design based on a concept that leverages form factor with functionality. Its unassuming physical affordance is intended to keep human expectation grounded while its actual interactive potential stokes human interest. The meticulous combination of both subtle and pronounced mechanical movements together with its stunning visual displays, highlight its affective affordance. As a result, we have developed the first iteration of our tabletop robot rich in affective potential for use in different research fields involving long-term human-robot interaction.
Randy Gomez, Deborah Szapiro, Kerl Galindo, Keisuke Nakamura
HRI4
2018 Interactive Reinforcement Learning from Demonstration and Human Evaluative Feedback
abstract
Programing robots to perform tasks is difficult in the real world because of its richness and uncertainty. For robots and agents to be more useful, they must be able to learn quickly from ordinary people via natural interactions. In this paper, we investigate how an agent can learn from demonstration and positive and negative evaluative feedback provided by a human teacher. Specifically, we proposed a model-based method-IRL-TAMER-by combining learning from demonstration via inverse reinforcement learning (IRL) and learning from human reward via the TAMER framework. We tested our method in the Grid World domain and compared with the TAMER framework using different discount factors on human reward. Our results suggest that although an agent learning via IRL can learn a useful value function indicating which states are good based on the demonstration, it cannot obtain an effective policy navigating to the goal state with one demonstration. However, learning from demonstration can reduce the number of human reward needed to obtain an optimal policy, especially the number of negative feedback. That is to say, learning from demonstration can be a jump-start for agent's learning from human reward and reduce the number of mistakes-incorrect actions. Furthermore, our results show that learning from demonstration can only be useful for agent's learning from human reward when the discount factor is small, i.e., learning from myopic human reward.
Guangliang Li, Bo He 0002, Randy Gomez, Keisuke Nakamura
RO-MAN4
2018 A robust multispectral palmprint matching algorithm and its evaluation for FPGA applications
Chao Li 0005, Yannick Benezeth, Keisuke Nakamura, Randy Gomez, Fan Yang 0019
J. Syst. Archit.3
2017 Improving separation of overlapped speech for meeting conversations using uncalibrated microphone array
abstract
In this paper, we propose a novel approach of sound source separation for meeting conversations even when using an uncalibrated microphone array. Our method can blindly estimate three parameters for separation, namely Steering Vectors (SVs), speaker indices, and activity periods of each speaker. First, we estimate the number of speakers and SVs by clustering Time Delay Of Arrival (TDOA) of the observed signal and selecting major clusters to compute TDOA-based SVs. Then, speaker indices and activity periods are estimated by thresholding spatial spectrum using estimated SVs, whose threshold is blindly obtained. Finally, we separate overlapped speeches/noise based on dynamic design of noise correlation matrices of the minimum variance distortionless response (MVDR) beamformer using blindly estimated parameters. The proposed algorithm was evaluated in both separation objective measure and recognition correct rate and showed improvements in both single and simultaneous speech scenarios in a reverberant meeting room. Moreover, the blindly estimated parameters improved separation and recognition compared to geometrically obtained parameters.
Keisuke Nakamura, Randy Gomez
ASRU1
2017 Exploring data augmentation methods in reverberant human-robot voice communication
abstract
Collecting training data is not an easy task especially in situation involving robots that require tremendous physical effort. The ability to augment data through synthetic means is a convenient tool to solve this problem. Therefore it is important to evaluate the extent of the usefulness of augmented data. In this paper, we will explore data augmentation schemes in reverberant environment and investigate a method to effectively select data. We experiment in a real reverberant environment condition and investigate both the traditional automatic speech recognition (ASR) system based on gaussian mixture model-hidden markov model (GMM-HMM) and the most current system based on Deep Neural Networks (i.e, HMM-DNN). Our results show that the combination of data augmentation and data selection, further improves system performance. In our experiments, we used real test data in a reverberant hands-free human-robot communication scenario.
Randy Gomez, Keisuke Nakamura
RO-MAN2
2017 SoundCraft: Enabling Spatial Interactions on Smartwatches using Hand Generated Acoustics
abstract
We present SoundCraft, a smartwatch prototype embedded with a microphone array, that localizes angularly, in azimuth and elevation, acoustic signatures: non-vocal acoustics that are produced using our hands. Acoustic signatures are common in our daily lives, such as when snapping or rubbing our fingers, tapping on objects or even when using an auxiliary object to generate the sound. We demonstrate that we can capture and leverage the spatial location of such naturally occurring acoustics using our prototype. We describe our algorithm, which we adopt from the MUltiple SIgnal Classification (MUSIC) technique [31], that enables robust localization and classification of the acoustics when the microphones are required to be placed at close proximity. SoundCraft enables a rich set of spatial interaction techniques, including quick access to smartwatch content, rapid command invocation, in-situ sketching, and also multi-user around device interaction. Via a series of user studies, we validate SoundCraft's localization and classification capabilities in non-noisy and noisy environments.
Teng Han, Khalad Hasan, Keisuke Nakamura, Randy Gomez, Pourang Irani
UIST3
2016 Online simultaneous localization and mapping of multiple sound sources and asynchronous microphone arrays
abstract
This paper presents an online method of simultaneous localization and mapping (SLAM) for estimating the positions of multiple moving sound sources and stationary robots and synchronizing microphone arrays attached to those robots. Since each robot with a microphone array can solely estimate the directions of sound sources, the two-dimensional source positions can be estimated from the source directions estimated by multiple robots using a triangulation method. In addition, sound mixtures can be separated accurately by regarding distributed microphone arrays as one big array. To perform these tasks, some methods have been proposed for localizing and synchronizing microphone arrays. These methods, however, can be used only if a single sound source exists because the time differences of arrival (TDOAs) between microphones are assumed to be directly observed. To overcome this limitation, we propose a unified state-space model that encodes the source and robot positions and the time offsets between microphone arrays in a latent space. Given the TDOAs and directions of arrival (DOAs) estimated by separating observed mixture sounds into source sounds, the latent variables are estimated jointly in an online manner using a FastSLAM2.0 algorithm that can deal with an unknown time-varying number of moving sound sources.
Kouhei Sekiguchi, Yoshiaki Bando, Keisuke Nakamura, Kazuhiro Nakadai, Katsutoshi Itoyama, Kazuyoshi Yoshii
IROS3
2016 Robust sound source mapping using three-layered selective audio rays for mobile robots
abstract
This paper investigates sound source mapping in a real environment using a mobile robot. Our approach is based on audio ray tracing which integrates occupancy grids and sound source localization using a laser range finder and a microphone array. Previous audio ray tracing approaches rely on all observed rays and grids. As such observation errors caused by sound reflection, sound occlusion, wall occlusion, sounds at misdetected grids, etc. can significantly degrade the ability to locate sound sources in a map. A three-layered selective audio ray tracing mechanism is proposed in this work. The first layer conducts frame-based unreliable ray rejection (sensory rejection) considering sound reflection and wall occlusion. The second layer introduces triangulation and audio tracing to detect falsely detected sound sources, rejecting audio rays associated to these misdetected sounds sources (short-term rejection). A third layer is tasked with rejecting rays using the whole history (long-term rejection) to disambiguate sound occlusion. Experimental results under various situations are presented, which proves the effectiveness of our method.
Daobilige Su, Keisuke Nakamura, Kazuhiro Nakadai, Jaime Valls Miró
IROS2
2016 Construction of Japanese Audio-Visual Emotion Database and Its Application in Emotion Recognition
Nurul Lubis, Randy Gomez, Sakriani Sakti, Keisuke Nakamura, Koichiro Yoshino, Satoshi Nakamura 0001, Kazuhiro Nakadai
LREC4
2016 Leveraging phantom signals for improved voice-based human-robot interaction
abstract
Voice-based system used in human-robot interaction is susceptible to challenging environment conditions. In an enclosed environment, the speech signal is often reflected which causes smearing as it is observed in the microphone. This phenomenon creates mismatch with the acoustic model, degrading the recognition performance and the robot's ability to understand and execute commands. Moreover, phantoms increase false-alarm in robot's attention system. To address these issues, environment-matched training and model adaptation may be used. The former requires enormous amount of training data to exhaustively cover different matched conditions whereas the latter needs several adaptation data to be collected at runtime. It is important to stress that data collection and the wait time are luxuries in a robot setup. In this paper, we extend our previous work that mitigates these problem by combining environment-adaptive training, speech enhancement with phantom awareness and fast model update, respectively. As a result, we achieve a robust voice-based system that enhances the observed speech, rejects phantoms and automatically updates the model at runtime to minimize the mismatch. Results show that the proposed method significantly outperforms our previous work.
Randy Gomez, Yurii Vasylkiv, Keisuke Nakamura, Takeshi Mizumoto, Kazuhiro Nakadai
RO-MAN3
2015 Temporal smearing compensation in reverberant environment for speech-based human-robot interaction
abstract
Speech-based human-robot interaction is often plagued with issues such as reverberation and changes in speaker position that impacts overall performance. In this paper, we show a method in compensating the joint effects of reverberation and the change in speaker position. The acoustic perturbation caused by these two takes its toll on the Automatic Speech Recognition (ASR) and then the Spoken Language Understanding (SLU). Consequently, these will lead to a failure in the human-robot interaction experience. The proposed method is specifically designed to address the challenging environment condition in which robots are deployed. First, we analyze the impact of reverberation in the form of temporal smearing per change in speaker position. Then, we extract the smearing coefficients that capture the joint dynamics between the speech signal at current position and the room acoustics as observed by the robot. These coefficients are utilized to update the room transfer function (RTF) and the suppression parameters are stored offline. Moreover, all of these processes are optimized in the context of the ASR system for robot application. In the online mode, the reverberant data at an arbitrary position is processed using the parameters pre-computed offline. This effectively compensates the joint effects of reverberation at the arbitrary speaker position. Experimental results using real data gathered in a human-robot communication setting show that the proposed method outperforms existing methods.
Randy Gomez, Keisuke Nakamura, Takeshi Mizumoto, Kazuhiro Nakadai
ICRA2
2015 On-the-spot calibration of microphone array Transfer Functions for robot audition
abstract
This paper investigates the calibration of a microphone array based robot audition system, namely calibration of microphone array Transfer Functions (TFs). There are mainly two methods to obtain TFs: geometrical calculation and measurement. The geometrical calculation has difficulty in simulating robot- and room-acoustics such as diffraction and reflection properties of a robot's body and a room, and the measurement is accurate but time-consuming and requires expertise on acoustics. Thus, we propose fast and simple on-the-spot calibration of TFs including robot- and room-acoustics. The proposed approach first estimates microphone location and clock-difference by Simultaneous Localization And Mapping (SLAM) using hand clap acoustic signals while a human is walking around a robot. Second, TFs with robot- and room-acoustics are estimated by hand clap acoustic signals and interpolated so that the TFs can be roundly arranged at regular intervals in an online manner. In the evaluation, we calibrated TFs only by 20 hand claps (took only 20 seconds), and the TFs showed considerable improvements in sound source localization and separation compared to geometrically calculated TFs and achieved comparable performance towards the measured TFs which are calibrated by approximately 60 minute recordings.
Keisuke Nakamura, Surya Ambrose, Kazuhiro Nakadai
ICRA1
2015 Dereverberation for active human-robot communication robust to speaker's face orientation
Randy Gomez, Levko Ivanchuk, Keisuke Nakamura, Takeshi Mizumoto, Kazuhiro Nakadai
INTERSPEECH3
2015 Utilizing visual cues in robot audition for sound source discrimination in speech-based human-robot communication
abstract
It is easy for human beings to discern whether an observed acoustic signal is a direct speech, reflected speech or noise through simple listening. Relying purely on acoustic cues is enough for human beings to discriminate between the different kinds of sound sources which is not straightforward for machines. A robot equipped with the current robot audition mechanism in most cases, will fail to differentiate a direct speech from the other sound sources because acoustic information alone is insufficient for effective discrimination. Robot audition is an important topic in speech-based human-robot communication. It enables the robot to associate the incoming speech signal to the user for an effective human-robot communication. In challenging environments, this task becomes difficult due to reflections of the direct speech signal and background noise sources. To counter this problem, a robot needs to have a minimum amount of prior information to discriminate the valid speech signal (direct speech) from the contaminants (i.e., speech reflections and background noise sources). Failure to do so would lead to false speech-to-speaker association in robot audition and will gravely impact human-robot communication experience. In this paper we propose to using visual cues to augment the traditional robot audition which relies solely on acoustic information. The proposed method significantly improves accuracy of speech-to-speaker association and machine understanding performance in real environment situation. Experimental results show that our expanded system is robust in discriminating direct speech from speech reflections and background noise sources.
Randy Gomez, Levko Ivanchuk, Keisuke Nakamura, Takeshi Mizumoto, Kazuhiro Nakadai
IROS3
2015 Robot-Audition-based Human-Machine Interface for a Car
abstract
This paper describes a Robot-Audition based Car Human Machine Interface (RA-CHMI). A RA-CHMI, like a car navigation system, has difficulty dealing with voice commands, since there are many noise sources in a car, including road noise, air-conditioner, music, and passengers. Microphone array processing developed in robot audition, may overcome this problem. Robot audition techniques, including sound source localization, Voice Activity Detection (VAD), sound source separation, and barge-in-able processing, were introduced by considering the characteristics of RA-CHMI. Automatic Speech Recognition (ASR), based on a Deep Neural Network (DNN), improved recognition performance and robustness in a noisy environment. In addition, as an integrated framework, HARK-Dialog was developed to build a multi-party and multi-modal dialog system, enabling the seamless use of cloud and local services with pluggable modular architecture. The constructed multi-party and multimodal RA-CHMI system did not require a push-to-talk button, nor did it require reducing the audio volume or air-conditioner when issuing speech commands. It could also control a four-DOF robot agent to make the system's responses more understandable. The proposed RA-CHMI was validated by evaluating essential techniques in the system, such as VAD and DNN-ASR, using real speech data recorded during driving. The entire design of the RA-CHMI system, including the system response time and the proper use of cloud/local services, are also discussed.
Kazuhiro Nakadai, Takeshi Mizumoto, Keisuke Nakamura
IROS3
2015 Robot audition based Acoustic Event Identification using a Bayesian model considering spectral and temporal uncertainties
abstract
To analyze auditory scenes of robots' surrounding environments, not only speeches but also non-speech sounds are important, which are spatially distributed and have different spectral and temporal characteristics. Thus, this paper investigates Acoustic Event Identification (AEI) which includes problems of localization, detection, and identification of sound sources. To achieve AEI by a robot in a real environment, we first propose to use a robot audition framework including sound source localization and separation to localize, detect, and separate acoustic events. For the identification, we propose two Bayesian models, iterative Latent Dirichlet Allocation (it-LDA) and Nested Pitman-Yor process with Uncertainty Compensation (NPY-UC). it-LDA and NPY-UC extract noise-robust sound-units and sound-words, respectively, and they consider probabilistic spectral and temporal uncertainties to robustify AEI against harsh environments such as noise and reverberation, etc. We have implemented these proposed methods using a robot-embedded microphone array. The preliminary results showed 5–18 pts improvement compared to a conventional GMM method in noisy environments thanks to the Bayesian framework in consideration of uncertainties.
Keisuke Nakamura, Kazuhiro Nakadai
IROS1
2015 Interactive sound source localization using robot audition for tablet devices
abstract
This paper investigates localization of sound sources in a real environment using a tablet device. For the localization, we use build-in sensors on a tablet device and additionally mount a cover with a microphone array. Because of the flat shape and limited sensor performance, the localization has mainly the following three issues; 1) the flat microphone array allows only azimuth estimation but elevation estimation; 2) the measurement noise of built-in sensors degrades the direction of arrival estimation performance; 3) the direction of arrival is not sufficient information to localize sound sources in three dimensional space to explore. To solve these issues, we propose interactive sound source localization using robot audition. For 1), we propose elevation estimation by integrating sound azimuth estimation and device motion information under an active audition framework in robot audition. For 2), we introduced a constrained optimization to robustify the direction of arrival estimation. For 3), we propose an interactive interface based on augmented reality for tablet devices, which navigates the device camera view to sound sources to explore. The proposed system was evaluated both subjectively and objectively and showed validity in real cases of sound source localization.
Keisuke Nakamura, Lana Sinapayen, Kazuhiro Nakadai
IROS1
2015 A case study of an automatic volume control interface for a telepresence system
abstract
The study of the telepresence robot as a tool for telecommunication from a remote location is attracting a considerable amount of attention. However, the problem arises that a telepresence robot system does not allow the volume of the user's utterance to be adjusted precisely, because it does not consider varying conditions in the sound environment, such as noise. In addition, when talking with several people in remote location, the user would like to be able to change the speaker volume freely according to the situation. In a previous study, a telepresence robot was proposed that has a function that automatically regulates the volume of the user's utterance. However, the manner in which the user exploits this function in a practical situation needs to be investigated. We propose a telepresence conversation robot system called “TeleCoBot.” TeleCoBot includes an operator's user interface, through which the volume of the user's utterance can be automatically regulated according to the distance between the robot and the conversation partner and the noise level in the robot's environment. We conducted a case study, in which the participants played a game using TeleCoBot's interface. The results of the study reveal the manner in which the participants used TeleCoBot and the additional factors that the system requires.
Masaaki Takahashi, Masa Ogata, Michita Imai, Keisuke Nakamura, Kazuhiro Nakadai
RO-MAN4
2014 Volume adaptation and visualization by modeling the volume level in noisy environments for telepresence system
abstract
The Lombard effect is the involuntary tendency of speakers to increase their vocal effort when speaking in a loud noise to enhance the audibility of their voice. There is a problem in telecommunication due to the Lombard effect. A speaker talks at a louder volume than necessary for the conversation partner at a remote location. This paper proposes a volume model that is required in order to automatically adjust the volume of an operator's voice at a remote communication via a telepresence robot, and develops an optimal volume control system LombaBot equipped on a telepresence robot with the model. The volume model measures the level of noise around the robot and the distance between a conversation partner and the robot to adjust the volume of the operator's voice. It has two types of volume adjustments. Those are called comfortable volume and secret talk volume. LombaBot enables people at a remote site to listen comfortably to the voice of a robot operator. Moreover, the operator is able to talk in low voices when s/he wants to talk in secret with nearby people. We confirmed that LombaBot adjusted the volume of an operator's voice properly in the noisy remote location.
Akira Hayamizu, Michita Imai, Keisuke Nakamura, Kazuhiro Nakadai
HAI3
2014 Speech-based human-robot interaction robust to acoustic reflections in real environment
abstract
Acoustic reflection inside an enclosed environment is detrimental to human-robot interaction. Reflection may manifest as phantom sources emanating from unknown directions. In effect, a single speaker may falsely manifest as multiple speakers to the robot audition system, impeding the robot's ability to correctly associate the speech command to the actual speaker. Moreover, speech reflection smears the original speech signal due to reverberation. This degrades speech recognition and understanding performance. Conventional robot audition schemes that rely purely on acoustics and spatial information are very sensitive to acoustic reflection which ultimately leads to the failure in human-robot interaction. We propose a method for human-robot interaction robust to the effect of acoustic reflection. First, visual information is utilized and head tracking scheme is employed to reinforce the acoustic information with the visual presence of a prospect user. Second, we employ a model-based sound event identification scheme and scrutinize whether the acoustic information is likely to be speech or non-speech. Using all the information we have gathered, we create a simple rule construct to effectively discriminate the original source (actual speaker) from phantom sources (reflection). Consequently, the corresponding source identified as phantom (reflection) is used to estimate the unwanted smearing for effective suppression via speech enhancement. Experiments are conducted in human-robot interaction setting in which the proposed method outperforms the conventional method.
Randy Gomez, Koji Inoue, Keisuke Nakamura, Takeshi Mizumoto, Kazuhiro Nakadai
IROS3
2014 Improvement in outdoor sound source detection using a quadrotor-embedded microphone array
abstract
This paper addresses sound source detection in an outdoor environment using a quadrotor with a microphone array. Since the previously reported method has a high computational cost, we proposed a sound source detection algorithm called MUltiple SIgnal Classification based on incremental Generalized Singular Value Decomposition (iGSVD-MUSIC), which detects sound source location and temporal activity with low computational cost. In addition, to relax an over-esitimation problem of noise correlation matrix which is used in iGSVD-MUSIC, we proposed Correlation Matrix Scaling (CMS), which realizes soft whitening of noise. The protptype system based on the proposed methods were evaluated with two types of microphone arrays in an outdoor environment. Experimental results showed that the combination of iGSVD-MUSIC and CMS improves sound source detection performance drastically and achieves real-time processing.
Takuma Ohata, Keisuke Nakamura, Takeshi Mizumoto, Taiki Tezuka, Kazuhiro Nakadai
IROS2
2014 Making a robot dance to diverse musical genre in noisy environments
abstract
In this paper we address the problem of musical genre recognition for a dancing robot with embedded microphones capable of distinguishing the genre of a musical piece while moving in a real-world scenario. For this purpose, we assess and compare two state-of-the-art musical genre recognition systems, based on Support Vector Machines and Markov Models, in the context of different real-world acoustic environments. In addition, we compare different preprocessing robot audition variants (single channel and separated signal from multiple channels) and test different acoustic models, learned a priori, to tackle multiple noise conditions of increasing complexity in the presence of noises of different natures (e.g., robot motion, speech). The results with six different musical genres suggest improved results, in the order of 43.6pp for the most complex conditions, when recurring to Sound Source Separation and acoustic models trained in similar conditions to the testing scenarios. A robot dance demonstration session confirms the applicability of the proposed integration for genre-adaptive dancing robots in real-world noisy environments.
João Lobato Oliveira, Keisuke Nakamura, Thibault Langlois, Fabien Gouyon, Kazuhiro Nakadai, Angelica Lim, Luís Paulo Reis, Hiroshi G. Okuno
IROS2
2014 Auditory-aware navigation for mobile robots based on reflection-robust sound source localization and visual SLAM
abstract
Autonomous robot navigation using Simultaneous localization and mapping (SLAM) is essential for scene understanding by robots. Most existing systems use visual information, and even though such visual-based technologies are robust and useful for many situations, they have difficulty dealing with certain scenarios such as occluded goal or when the goal is out of frame. Introducing audio information to the navigation system solves these issues effectively. Several audio-based methods have been developed in the past to deal with these issues. However, these existing audio-based methods have been developed with the assumption that the space around the robot is open i.e. no sound reflection occurs. Hence, the invisible goals where the sound reflection can be localized have not been fully considered. This paper proposes a reflection-robust sound source localization method using visual SLAM. This method can deal with sound sources whose direct paths are not available, and using localization, we can set a goal only for the actual sound source. Also, to correct the drift present in the local estimates of Visual Odometry (VO), SLAM was integrated with the system, thus increasing the accuracy and robustness of mapping and navigation of our proposed method. The performance of the proposed system is compared to conventional methods, and it proves to be very efficient and robust especially in extremely reverberant situations. It was found that the integration of VO and SLAM improved the average error of a map by approximately 50 pts, and the accuracy of SSL for direct path of sounds was improved by approximately 8 pts. With the online implementation of these methods we successfully achieved audio visual navigation for the actual sound sources.
Gautam Narang, Keisuke Nakamura, Kazuhiro Nakadai
SMC2
2013 Robustness to speaker position in distant-talking automatic speech recognition
abstract
In this paper, we show a method that significantly improved our previous work in single-channel dereverberation. The proposed method is more robust to changes in speaker position in distant talking ASR. First, we update the room transfer function (RTF) and weighting parameters for dereverberation to the target speaker position. This scheme corrects speech power variation as a function of position in the waveform level. Consequently, its impact to the acoustic model is verified. Then, we implement a fast acoustic model update reflective of the speech power level of the target speaker position. Furthermore, the scheme in updating the model is simple and precludes time-consuming model re-estimation. As a result, the proposed method can be executed online. The synergy of these corrective measures significantly minimizes the mismatch between training and testing conditions. We test our method using real reverberant data with different locations inside the room. Experimental results show that the proposed method outperforms the conventional methods in terms of ASR performance. Moreover, our fast acoustic model update scheme is at par in terms of recognition performance against time-consuming model re-estimation.
Randy Gomez, Keisuke Nakamura, Kazuhiro Nakadai
ICASSP2
2013 Hands-free human-robot communication robust to speaker's radial position
abstract
In this paper we present a method in room transfer function (RTF) estimation, employed specifically for dereverberation in hands-free human-robot communication.We introduce a radial distance compensation scheme which significantly improved the RTF estimate robust to the speech power variation due to changes in speaker's radial position. The proposed method is implemented in two levels; first, waveform-level compensation is executed to reflect the change in power caused by the change of radial position to the RTF. We generated possible RTF estimates within a close neighbourhood based on curve fitting. Then, we select among these estimates the optimal RTF based on acoustic model likelihood criterion, the same criterion employed in automatic speech recognition (ASR) systems. The latter is referred to as acoustic model-level compensation, which links the generated RTF to the ASR. We note that in ASR application, both waveform and acoustic models play an important role in achieving optimal performance. Thus, the synergistic effect of the two processes guarantee ASR performance improvement when used in conjunction with our ASR-based dereverberation scheme. Experimental evaluation show robustness in recognition performance when used in hands-free human-robot communication environment.
Randy Gomez, Keisuke Nakamura, Kazuhiro Nakadai, Ui-Hyun Kim, Hiroshi G. Okuno, Tatsuya Kawahara
ICRA2
2013 Dereverberation robust to speaker's azimuthal orientation in multi-channel human-robot communication
abstract
The acoustical dynamics of reverberation in an enclosed environment poses a problem to human-robot communication. Any change in the azimuthal orientation of the speaker contributes to unpredictable acoustical activity resulting in a degradation in the performance of the automatic speech recognition (ASR) system. Thus, dereverberation techniques need to address this issue prior to ASR. Dereverberation in multi-channel applications primarily evolves in the adoption of a suitable reverberant model that results to a computationally feasible solution and at the same time yields an accurate estimate of the harmful reflections (i.e., late reflection) for effective suppression. In this paper we address this problem by introducing a hybrid method based on multi-channel processing on a singlechannel reverberant model platform. The proposed method is capable of accurate signal estimation, a property inherent to a multi-channel system, and at the same time bears the computational efficiency derived from single-channel reverberant model approach. The proposed method is summarized as follows; First, multi-channel sound-source processing is employed to obtain the full reverberant and the late reflection signal estimates. Then, equalization is employed to update the late reflection estimate reflective of the change in azimuth prior to dereverberation. The equalization parameters for azimuthal change are obtained through an offline optimization procedure. Experimental evaluation in an actual human-robot communication environment shows that the proposed method outperforms existing methods in terms of robustness in the ASR performance.
Randy Gomez, Keisuke Nakamura, Kazuhiro Nakadai
IROS2
2013 Real-time super-resolution three-dimensional sound source localization for robots
abstract
This paper investigates Sound Source Localization (SSL) for a robot in a real world. Previously, we focused on one-dimensional SSL for azimuth and assumed that target sources are distributed close to a horizontal plane. Without this assumption, the SSL performance is drastically degraded. Thus, three-dimensional SSL is essential to improve the localization for sound sources distributed in a three-dimensional space. Compared to one-dimensional SSL, three-dimensional SSL mainly has the following problems: 1) a massive number of Transfer Function (TF) measurements for microphone array calibration are required for three dimensions to maintain the spatial resolution of SSL sufficiently-high, 2) the computational cost for searching for sound sources drastically increases in high-dimensional spaces. For the first issue, we extend the previously-proposed one-dimensional TF interpolation method, integrating time-domain-based and frequency-domain-based interpolation, to three dimensions. The interpolation achieves three-dimensional super-resolution SSL and reduction of the number of TF measurements while maintaining the spatial resolution of SSL. For the second issue, we propose optimal hierarchical SSL, which reduces computational cost for searching for sound sources by introducing a hierarchical search algorithm instead of using greedy search in localization. We previously proposed the concept of the algorithm. This paper additionally discusses theoretical optimality in hierarchization to minimize the total computational cost of SSL. The method determines the number of hierarchies and the resolution of each hierarchy depending on desired spatial resolution. These techniques are integrated into an SSL system using a robot. The experimental result showed: 1) the proposed interpolation method achieved super-resolution SSL working better than that with pre-measured TFs, 2) the optimal hierarchical SSL drastically reduced computational cost by approximately 97%.
Keisuke Nakamura, Randy Gomez, Kazuhiro Nakadai
IROS1
2012 Multi-party human-robot interaction with distant-talking speech recognition
abstract
Speech is one of the most natural medium for human communication, which makes it vital to human-robot interaction. In real environments where robots are deployed, distant-talking speech recognition is difficult to realize due to the effects of reverberation. This leads to the degradation of speech recognition and understanding, and hinders a seamless human-robot interaction. To minimize this problem, traditional speech enhancement techniques optimized for human perception are adopted to achieve robustness in human-robot interaction. However, human and machine perceive speech differently: an improvement in speech recognition performance may not automatically translate to an improvement in human-robot interaction experience (as perceived by the users). In this paper, we propose a method in optimizing speech enhancement techniques specifically to improve automatic speech recognition (ASR) with emphasis on the human-robot interaction experience. Experimental results using real reverberant data in a multi-party conversation, show that the proposed method improved human-robot interaction experience in severe reverberant conditions compared to the traditional techniques.
Randy Gomez, Tatsuya Kawahara, Keisuke Nakamura, Kazuhiro Nakadai
HRI3
2012 Online audio beat tracking for a dancing robot in the presence of ego-motion noise in a real environment
abstract
This paper presents the design and implementation of a real-time real-world beat tracking system which runs on a dancing robot. The main problem of such a robot is that, while it is moving, ego noise is generated due to its motors, and this directly degrades the quality of the audio signal features used for beat tracking. Therefore, we propose to incorporate ego noise reduction as a pre-processing stage prior to our tempo induction and beat tracking system. The beat tracking algorithm is based on an online strategy of competing agents sequentially processing a continuous musical input, while considering parallel hypotheses regarding tempo and beats. This system is applied to a humanoid robot processing the audio from its embedded microphones on-the-fly, while performing simplistic dancing motions. A detailed and multi-criteria based evaluation of the system across different music genres and varying stationary/non-stationary noise conditions is presented. It shows improved performance and noise robustness, outperforming our conventional beat tracker (i.e., without ego noise suppression) by 15.2 points in tempo estimation and 15.0 points in beat-times prediction.
João Lobato Oliveira, Gökhan Ince, Keisuke Nakamura, Kazuhiro Nakadai
ICRA3
2012 Online learning for template-based multi-channel ego noise estimation
abstract
This paper presents a system that gives a robot the ability to diminish its own disturbing noise (i.e., ego noise) by utilizing template-based ego noise estimation, an algorithm previously developed by the authors. In pursuit of an autonomous, online and adaptive template learning system in this work, we specifically focus on eliminating the requirement of an offline training session performed in advance to build the essential templates, which represent the ego noise. The idea of discriminating ego noise from all other sound sources in the environment enables the robot to learn the templates online without requiring any prior information. Based on the directionality/diffuseness of the sound sources, the robot can easily decide whether the template should be discarded because it is corrupted by external noises, or it should be inserted into the database because the template consists of pure ego noise only. Furthermore, we aim to update the template database optimally by introducing an additional time-variant forgetting factor parameter, which provides a balance between adaptivity and stability of the learning process automatically. Moreover, we enhanced the single-channel noise estimation system to be compatible with the multi-channel robot audition framework so that ego noise can be eliminated from all signals stemming from multiple sound sources respectively. We demonstrate that the proposed system allows the robot to have the ability of online template learning as well as a high performance of noise estimation and suppression for multiple sound sources.
Gökhan Ince, Kazuhiro Nakadai, Keisuke Nakamura
IROS3
2012 Real-time super-resolution Sound Source Localization for robots
abstract
Sound Source Localization (SSL) is an essential function for robot audition and yields the location and number of sound sources, which are utilized for post-processes such as sound source separation. SSL for a robot in a real environment mainly requires noise-robustness, high resolution and real-time processing. A technique using microphone array processing, that is, Multiple Signal Classification based on Standard EigenValue Decomposition (SEVD-MUSIC) is commonly used for localization. We improved its robustness against noise with high power by incorporating Generalized EigenValue Decomposition (GEVD). However, GEVD-based MUSIC (GEVD-MUSIC) has mainly two issues: 1) the resolution of pre-measured Transfer Functions (TFs) determines the resolution of SSL, 2) its computational cost is expensive for real-time processing. For the first issue, we propose a TF interpolation method integrating time-domain-based and frequency-domain-based interpolation. The interpolation achieves super-resolution SSL, whose resolution is higher than that of the pre-measured TFs. For the second issue, we propose two methods, MUSIC based on Generalized Singular Value Decomposition (GSVD-MUSIC), and Hierarchical SSL (H-SSL). GSVD-MUSIC drastically reduces the computational cost while maintaining noise-robustness in localization. H-SSL also reduces the computational cost by introducing a hierarchical search algorithm instead of using greedy search in localization. These techniques are integrated into an SSL system using a robot embedded microphone array. The experimental result showed: the proposed interpolation achieved approximately 1 degree resolution although we have only TFs at 30 degree intervals, GSVD-MUSIC attained 46.4% and 40.6% of the computational cost compared to SEVD-MUSIC and GEVD-MUSIC, respectively, H-SSL reduces 59.2% computational cost in localization of a single sound source.
Keisuke Nakamura, Kazuhiro Nakadai, Gökhan Ince
IROS1
2012 Outdoor auditory scene analysis using a moving microphone array embedded in a quadrocopter
abstract
This paper addresses auditory scene analysis, especially, sound source localization using an aerial vehicle with a microphone array in an outdoor environment. Since such a vehicle is able to search sound sources quickly and widely, it is useful to detect outdoor sound sources, for instance, to find distressed people in a disaster situation. In such an environment, noise is quite loud and dynamically-changing, and conventional microphone array techniques studied in the field of indoor robot audition are of less use. We, thus, proposed MUltiple SIgnal Classification based on incremental Generalized EigenValue Decomposition (iGEVD-MUSIC). It can deal with high power noise by introducing a noise correlation matrix and GEVD even when the signal-to-noise ratio is less than 0 dB. In addition, the noise correlation matrix is incrementally estimated to adapt to dynamic changes in noise. We developed a prototype system for auditory scene analysis based on the proposed method using the Parrot AR.Drone with a microphone array and a Kinect device. Experimental results using the prototype system showed that dynamically-changing noise is properly suppressed with the proposed method and multiple human voice sources are able to be localized even when the AR.Drone is moving in an outdoor environment.
Keita Okutani, Takami Yoshida, Keisuke Nakamura, Kazuhiro Nakadai
IROS3
2012 Live assessment of beat tracking for robot audition
abstract
In this paper we propose the integration of an online audio beat tracking system into the general framework of robot audition, to enable its application in musically-interactive robotic scenarios. To this purpose, we introduced a staterecovery mechanism into our beat tracking algorithm, for handling continuous musical stimuli, and applied different multi-channel preprocessing algorithms (e.g., beamforming, ego noise suppression) to enhance noisy auditory signals lively captured in a real environment. We assessed and compared the robustness of our audio beat tracker through a set of experimental setups, under different live acoustic conditions of incremental complexity. These included the presence of continuous musical stimuli, built of a set of concatenated musical pieces; the presence of noises of different natures (e.g., robot motion, speech); and the simultaneous processing of different audio sources on-the-fly, for music and speech. We successfully tackled all these challenging acoustic conditions and improved the beat tracking accuracy and reaction time to music transitions while simultaneously achieving robust automatic speech recognition.
João Lobato Oliveira, Gökhan Ince, Keisuke Nakamura, Kazuhiro Nakadai, Hiroshi G. Okuno, Luís Paulo Reis, Fabien Gouyon
IROS3
2012 An active audition framework for auditory-driven HRI: Application to interactive robot dancing
abstract
In this paper we propose a general active audition framework for auditory-driven Human-Robot Interaction (HRI). The proposed framework simultaneously processes speech and music on-the-fly, integrates perceptual models for robot audition, and supports verbal and non-verbal interactive communication by means of (pro)active behaviors. To ensure a reliable interaction, on top of the framework a behavior decision mechanism based on active audition policies the robot's actions according to the reliability of the acoustic signals for auditory processing. To validate the framework's application to general auditory-driven HRI, we propose the implementation of an interactive robot dancing system. This system integrates three preprocessing robot audition modules: sound source localization, sound source separation, and ego noise suppression; two modules for auditory perception: live audio beat tracking and automatic speech recognition; and multi-modal behaviors for verbal and non-verbal interaction: music-driven dancing and speech-driven dialoguing. To fully assess the system, we set up experimental and interactive real-world scenarios with highly dynamic acoustic conditions, and defined a set of evaluation criteria. The experimental tests revealed accurate and robust beat tracking and speech recognition, and convincing dance beat-synchrony. The interactive sessions confirmed the fundamental role of the behavior decision mechanism for actively maintaining a robust and natural human-robot interaction.
João Lobato Oliveira, Gökhan Ince, Keisuke Nakamura, Kazuhiro Nakadai, Hiroshi G. Okuno, Luís Paulo Reis, Fabien Gouyon
RO-MAN3
2011 Correlation matrix interpolation in Sound Source Localization for a robot
abstract
In microphone array processing, a Correlation Matrix (CM) between multiple channel input signals is widely utilized for various purposes such as Sound Source Localization (SSL), etc. The CM corresponds to spatial information between a microphone and a sound source, and it dynamically changes as they move in a dynamic environment. Since pre-measured CMs have difficulties in dealing with dynamically changing environments due to their discreteness, this paper addresses a CM interpolation utilizing a few discrete known CMs. The proposed method is based on an integration of eigen-value-scaling and linear interpolation, which requires a few measurements and low computational costs. Apart from conventional methods, this paper deals with an interpolation based on a correlation-matrix not based on a transfer-function, which achieves the integration approach. The evaluation shows better estimation performance compared to conventional transfer-function-based methods. We also applied the interpolation to SSL and confirmed its validity.
Keisuke Nakamura, Kazuhiro Nakadai, Hirofumi Nakajima, Gökhan Ince
ICASSP1
2011 Assessment of general applicability of ego noise estimation
abstract
Noise generated due to the motion of a robot deteriorates the quality of the desired sounds recorded by robot-embedded microphones. On top of that, a moving robot is also vulnerable to its loud fan noise that changes its orientation relative to the moving limbs where the microphones are mounted on. To tackle the non-stationary ego-motion noise and the direction changes of fan noise, we propose an estimation method based on instantaneous prediction of ego noise using parameterized templates. We verify the ego noise suppression capability of the proposed estimation method on a humanoid robot by evaluating it on two important applications in the framework of robot audition: (1) automatic speech recognition and (2) sound source localization. We demonstrate that our method improves recognition and localization performance during both head and arm motions considerably.
Gökhan Ince, Keisuke Nakamura, Futoshi Asano, Hirofumi Nakajima, Kazuhiro Nakadai
ICRA2
2011 Assessment of single-channel ego noise estimation methods
abstract
While a robot is moving, ego noise is generated due to the fans and motors of the robot. Furthermore, a robot is not only subject to the ego noise, but also to the ambient noise of the environment, both having different short-term signal characteristics. Because ego-motion noise generated by the motors is non-stationary, and the BackGround Noise (BGN) is stationary, one single noise estimation method is unable to track the changes in both noise spectra rapidly and accurately. Therefore, we propose to use the combination of two different noise estimation methods adequate for each one of co-existing noise types in a unified framework: 1) a stationary noise estimation method called Histogram-based Recursive Level Estimation (HRLE) and 2) a non-stationary noise estimation method called Template-based Estimation (TE). In this paper, we evaluate the performance of several single-channel based noise estimation techniques in terms of their prediction accuracy and quality of the speech signals enhanced by spectral subtraction methods. The experimental results show that our system, compared to the conventional single-stage noise estimation methods, achieves better performance in attaining signal quality and improving word correct rates.
Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Jun-ichi Imura, Keisuke Nakamura, Hirofumi Nakajima
IROS5
2011 Incremental learning for ego noise estimation of a robot
abstract
Using pre-recorded templates to estimate and suppress the ego noise of a robot is advantageous because this method is able to cope with the non-stationarity of this particular type of noise. However, standard template-based estimation requires human intervention in the offline training sessions, storage of large amounts of data and does not adapt to the dynamical changes in the environmental conditions. In this paper we investigate the feasibility of an incremental template learning system to tackle these drawbacks. Incremental learning enables the system to acquire new templates on the fly and update the older ones appropriately. Whilst allowing the system to continually increase its knowledge and enhancing its estimation performance, this learning scheme also reduces the size of the database. We evaluate the performance of the proposed noise estimation method in terms of its estimation accuracy, quality of speech signals enhanced by spectral subtraction method, and size of database. The experimental results show that our system compared to conventional single-channel noise estimation methods achieves better performance in attaining signal quality and improving word correct rates.
Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Jun-ichi Imura, Keisuke Nakamura, Hirofumi Nakajima
IROS5
2011 SLAM-based online calibration of asynchronous microphone array for robot audition
abstract
This paper addresses the online calibration of an asynchronous microphone array for robots. Conventional microphone array technologies require a lot of measurements of transfer functions to calibrate microphone locations, and a multi-channel A/D converter for inter-microphone synchronization. We solve these two problems using a framework combining Simultaneous Localization and Mapping (SLAM) and beamforming in an online manner. To do this, we assume that estimations of microphone locations, a sound source location, and microphone clock difference correspond to mapping, self-localization, observation errors in SLAM, respectively. In our framework, the SLAM process calibrates locations and clock differences of microphones every time a microphone array observes a sound like a human's clapping, and a beamforming process works as a cost function to decide the convergence of calibration by localizing the sound with the estimated locations and clock differences. After calibration, beamforming is used for sound source localization. We implemented a prototype system using Extended Kalman Filter (EKF) based SLAM and Delay-and-Sum Beamforming (DS-BF). The experimental results showed that microphone locations and clock differences were estimated properly with 10–15 sound events (handclaps), and the error of sound source localization with the estimated information was less than the grid size of beamforming, that is, the lowest error was theoretically attained.
Hiroaki Miura, Takami Yoshida, Keisuke Nakamura, Kazuhiro Nakadai
IROS3
2011 Intelligent sound source localization and its application to multimodal human tracking
abstract
We have assessed robust tracking of humans based on intelligent Sound Source Localization (SSL) for a robot in a real environment. SSL is fundamental for robot audition, but has three issues in a real environment: robustness against noise with high power, lack of a general framework for selective listening to sound sources, and tracking of inactive and/or noisy sound sources. To address the first issue, we extended Multiple SIgnal Classification by incorporating Generalized EigenValue Decomposition (GEVD-MUSIC) so that it can deal with high power noise and can select target sound sources. To address the second issue, we proposed Sound Source Identification (SSI) based on hierarchical gaussian mixture models and integrated it with GEVD-MUSIC to realize a selective listening function. To address the third issue, we integrated audio-visual human tracking using particle filtering. Integration of these three techniques into an intelligent human tracking system showed: 1) GEVD-MUSIC improved the noise-robustness of SSL by a signal-to-noise ratio of 5-6 dB; 2) SSI performed more than 70% in F-measure even in a noisy environment; and 3) audio-visual integration improved the average tracking error by approximately 50%.
Keisuke Nakamura, Kazuhiro Nakadai, Futoshi Asano, Gökhan Ince
IROS1
2010 Control of bipedal running by the angular-momentum-based synchronization structure
abstract
This paper investigates the running control problem for the 7-link, 6-actuator planar bipedal robot including ankle joints. The control strategy is based on the “synchronization structure” with the angular momentum of the pivot point. The synchronization not only follows the simplified joint angle dynamics of human running but also generates the uniform running speed from wide range of initial speed, which claims that the controller is robust for the error of initial speed. Moreover, it successfully verifies the acceleration of running speed from 0.1[m/s], which is almost zero speed, to a uniform running speed “without” switching controllers. Finally, the controller is applied for the running on an uneven terrain, and it successfully achieves running with the maximum slope angle 6[deg], which is highly steep terrain in the real situation.
Keisuke Nakamura, Shigeki Nakaura, Mitsuji Sampei
ICRA1
2010 Gradient descent bit flipping algorithms for decoding LDPC codes
abstract
A novel class of bit-flipping (BF) algorithm for decoding low-density parity-check (LDPC) codes is presented. The proposed algorithms, which are referred to as gradient descent bit flipping (GDBF) algorithms, can be regarded as simplified gradient descent algorithms. The proposed algorithms exhibit better decoding performance than known BF algorithms, such as the weighted BF algorithm or the modified weighted BF algorithm for several LDPC codes.
Tadashi Wadayama, Keisuke Nakamura, Masayuki Yagita, Yuuki Funahashi, Shogo Usami, Ichi Takumi
IEEE Trans. Commun.2
2009 Intelligent sound source localization for dynamic environments
abstract
As robotic technology plays an increasing role in human lives, ¿robot audition¿, human-robot communication, is of great interest, and robot audition needs to be robust and adaptable for dynamic environments. This paper addresses sound source localization working in dynamic environments for robots. Previously, noise robustness and dynamic localized sound selection have been enormous issues for practical use. To correct the issues, a new localization system ¿Selective Attention System¿ is proposed. The system has four new functions: localization with Generalized EigenValue Decomposition of correlation matrices for noise robustness(¿Localization with GEVD¿), sound source cancellation and focus (¿Target Source Selection¿), human-like dynamic Focus of Attention (¿Dynamic FoA¿), and correlation matrix estimation for robotic head rotation (¿Correlation Matrix Estimation¿). All are achieved by the dynamic design of correlation matrices. The system is implemented into a humanoid robot, and the experimental validation is successfully verified even when the robot microphones move dynamically.
Keisuke Nakamura, Kazuhiro Nakadai, Futoshi Asano, Yuji Hasegawa, Hiroshi Tsujino
IROS1
2005 Utilization of a Reusable File Format for Motional 3D Objects on an Educational System
Kenji Saito, Keisuke Nakamura, Hajime Saito, Takashi Maeda
ICCE2
2004 Noise robust real world spoken dialogue system using GMM based rejection of unintended inputs
abstract
ICSLP2004: the 8th International Conference on Spoken Language Processing, October 4-8, 2004, Jeju Island, Korea.
Akinobu Lee, Keisuke Nakamura, Ryuichi Nisimura, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH2