EDBT 2026 Demo / reviewers in the wild / expert
Tetsuya Ogata
dblp:38/1658
· DBLP profile ↗
192ranked-venue papers
15as first author
23since 2021 · last 2025
0000-0001-7015-0379ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 169 · 13 first-author · 22 since 2021Systems, architecture and hardware · 95 · 11 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 50 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 11 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Focused Blind Switching Manipulation Based on Constrained and Regional Touch States of Multi-Fingered Hand Using Deep LearningabstractTo achieve a desired grasping posture (including object position and orientation), multi-finger motions need to be conducted according to the the current touch state. Specifically, when subtle changes happen during correcting the object state, not only proprioception but also tactile information from the entire hand can be beneficial. However, switching motions with high-DOFs of multiple fingers and abundant tactile information is still challenging. In this study, we propose a loss function with constraints of touch states and an attention mechanism for focusing on important modalities depending on the touch states. The policy model is AE-LSTM which consists of Autoencoder (AE) which compresses abundant tactile information and Long Short-Term Memory (LSTM) which switches the motion depending on the touch states. Motion for cap-opening was chosen as a target task which consists of sub tasks of sliding an object and opening its cap. As a result, the proposed method achieved the best success rates with a variety of objects for real time cap-opening manipulation. Furthermore, we could confirm that the proposed model acquired the features of each subtask and attention on specific modalities. Satoshi Funabashi, Atsumu Hiramoto, Naoya Chiba, Alexander Schmitz, Shardul Kulkarni, Tetsuya Ogata |
ICRA | 6 |
| 2025 | UF-RNN: Real-Time Adaptive Motion Generation Using Uncertainty-Driven Foresight PredictionabstractTraining robots to operate effectively in environments with uncertain states—such as ambiguous object properties or unpredictable interactions—remains a longstanding challenge in robotics. Imitation learning methods typically rely on successful examples and often neglect failure scenarios where uncertainty is most pronounced. To address this limitation, we propose the Uncertainty-driven Foresight Recurrent Neural Network (UF-RNN), a model that combines standard time-series prediction with an active "Foresight" module. This module performs internal simulations of multiple future trajectories and refines the hidden state to minimize predicted variance, enabling the model to selectively explore actions under high uncertainty. We evaluate UF-RNN on a door-opening task in both simulation and a real-robot setting, demonstrating that, despite the absence of explicit failure demonstrations, the model exhibits robust adaptation by leveraging self-induced chaotic dynamics in its latent space. When guided by the Foresight module, these chaotic properties stimulate exploratory behaviors precisely when the environment is ambiguous, yielding improved success rates compared to conventional stochastic RNN baselines. These findings suggest that integrating uncertainty-driven foresight into imitation learning pipelines can significantly enhance a robot’s ability to handle unpredictable real-world conditions. Hyogo Hiruma, Tetsuya Ogata |
IROS | 3 |
| 2025 | Deep Predictive Learning with Proprioceptive and Visual Attention for Humanoid Robot Repositioning AssistanceabstractCaregiving is a vital role for domestic robots, especially the repositioning care has immense societal value, critically improving the health and quality of life of individuals with limited mobility. However, repositioning task is a challenging area of research, as it requires robots to adapt their motions while interacting flexibly with patients. The task involves several key challenges: (1) applying appropriate force to specific target areas; (2) performing multiple actions seamlessly, each requiring different force application policies; and (3) motion adaptation under uncertain positional conditions. To address these, we propose a deep neural network (DNN)-based architecture utilizing proprioceptive and visual attention mechanisms, along with impedance control to regulate the robot's movements. Using the dual-arm humanoid robot Dry-AIREC, the proposed model successfully generated motions to insert the robot's hand between the bed and a mannequin's back without applying excessive force, and it supported the transition from a supine to a lifted-up position. The project page is here: https://sites.google.com/view/caregiving-robot-airec/repositioning Tamon Miyake, Namiko Saito, Tetsuya Ogata, Shigeki Sugano |
IROS | 3 |
| 2025 | Interactive Object Detection by Mitigating Uncertainty of Robot Task Plans using Large Language ModelabstractRecently, many attempts have been made to integrate the foundation model with robotics. In most of those attempts, the model recognition results were treated as unique; however, the recognition results required for real robot tasks vary with the task goal. The recognition results of the foundation model are determined from the detection query; therefore, in the case of an ambiguous query, the query must be modified to match the purpose of the robot task. In this study, we propose an object recognition method that considers the task goal through application of an interactive task planning method using a large language model. The proposed method clarifies the purpose of the robot task by asking the user a question. Hence, uncertainty in the task plan due to ambiguous operation instructions is mitigated. During the task plan update arising from the dialog process, the object-detection results obtained from the query in the planning results are also updated to match the task goal. In our experiments, the proposed method’s effectiveness is verified quantitatively and qualitatively via object-detection tasks conducted on a custom-built verification dataset. Kanata Suzuki, Akane Ushizaka, Kazuki Hori, Tetsuya Ogata |
IROS | 4 |
| 2025 | Autonomous dialogue generation based on phase boundary detection within continuous motion for domestic robotabstractDialogue generation plays a key role in responding to user and providing transparency in motion execution in human-robot interaction. As motion planning is generally performed in terms of discrete motions, previous studies have focused on dialogue generation at the boundaries between motions. Recently, continuous motion generation was proposed to enable adapting actions to unique characteristics of the objects for domestic robots. Since a continuous motion generally involves physical and nonphysical phases, providing dialogues when the phase changes is crucial for decreasing users’ anxiety and guaranteeing safety. However, continuous motions lack clear phase boundaries, posing challenges for dialogue generation between phases. For this problem, we segmented continuous motions into discrete phases, and constructed a system to enable the robot to autonomously generate dialogues by detecting phase boundaries. To do so, we built phase estimation models using robot sensor data and designed system modules. Specifically, we collected data in the scenario of a robot assisting to lift the user up from bed. We segmented the continuous motion into three phases based on the user’s posture and whether the robot applied force to the human. The best phase estimation model achieved a macro F1 score of 0.894, demonstrating that phases can be estimated from sensor data. The evaluation results of our system demonstrated that the system accurately detects phase boundaries and generates appropriate dialogues corresponding to phases. Furthermore, we conducted simulations with a user agent to investigate system behaviors when the phase estimation was incorrect. The results suggested that explicitly stating the phase is important for avoiding misunderstandings and safety issues. Sixia Li, Tamon Miyake, Tetsuya Ogata, Shigeki Sugano, Shogo Okada |
RO-MAN | 3 |
| 2024 | Real-time Coordinated Motion Generation: A Hierarchical Deep Predictive Learning Model for Bimanual TasksabstractRobots that autonomously operate in human living environments require the ability to adapt to unpredictable changes and flexibly handle a variety of tasks. Particularly, coordinated bimanual motions are essential for enabling tasks that are difficult with just one hand, such as grasping bulky objects, transporting heavy loads, and precision work. Traditional methods of generating robot motions typically involve executing pre-programmed motions, making it challenging to adapt to complex and unpredictable environmental changes. To address this issue, our research focuses on generating diverse motions that can flexibly adapt to environmental changes based on Deep Predictive Learning from a small amount of real-world data. Previous Deep Predictive Learning models have generated the motions of a robot’s left and right arms by a single LSTM, making it difficult to operate them independently. Therefore, we propose a new Hierarchical Deep Predictive Learning model specialized for generating coordinated bimanual motions. This model comprises three components: a Left-LSTM, which learns the body and visual information on the robot’s left side, a Right-LSTM that performs a similar function for the right side, and a Union-LSTM which integrates this information at a higher level. To verify the effectiveness of the proposed model, we conducted bimanual grasping experiments with multiple different objects using two different robots. The experimental results showed that independent of hardware, our model demonstrated a higher success rate compared to the traditional approach, indicating its enhanced capability in coordinating bimanual motions. Genki Shikada, Simon Armleder, Gordon Cheng, Tetsuya Ogata |
IROS | 5 |
| 2024 | Sensorimotor Attention and Language-based Regressions in Shared Latent Variables for Integrating Robot Motion Learning and LLMabstractIn recent years, studies have been actively conducted on combining large language models (LLM) and robotics; however, most have not considered end-to-end feed-back in the robot-motion generation phase. The prediction of deep neural networks must contain errors, it is required to update the trained model to correspond to the real environment to generate robot motion adaptively. This study proposes an integration method that connects the robot-motion learning model and LLM using shared latent variables. When generating robot motion, the proposed method updates shared parameters based on prediction errors from both sensorimotor attention points and task language instructions given to the robot. This allows the model to search for latent parameters appropriate for the robot task efficiently. Through simulator experiments on multiple robot tasks, we demonstrated the effectiveness of our proposed method from two perspectives: position generalization and language instruction generalization abilities. Kanata Suzuki, Tetsuya Ogata |
IROS | 2 |
| 2024 | Multi-Fingered Dragging of Unknown Objects and Orientations Using Distributed Tactile Information Through Vision-Transformer and LSTMabstractMulti-fingered hands can be suitable for stable object manipulation. Furthermore, abundant tactile information can be acquired with multi-fingered hands, useful to recognize the object’s properties, which is beneficial to adapt the motion to the object. However, generating dexterous manipulation motions with multi-fingered hands with high density tactile sensors is challenging due to complex touch states. Hence, tasks that conventionally require a high level of active tactile sensing simultaneously with motion generation, such as pulling in the hand while recognizing the posture of an object are difficult to accomplish. In this letter, we propose a novel deep predictive learning approach using Vision-Transformer (ViT) and Long-Short Term Memory (LSTM). The ViT’s attention mechanism can spatially focus on specific fingers represented by distributed 3-axis tactile sensors (uSkin). The LSTM can preserve long time-series information of the manipulation which can realize changing the desired motion according to the initial touching position and orientation for the target object. Results showed that the ViT-LSTM is effective in performing adaptive finger movements according to the properties of the object, i.e. its hardness and relative posture. Takahisa Ueno, Satoshi Funabashi, Alexander Schmitz, Shardul Kulkarni, Tetsuya Ogata, Shigeki Sugano |
IROS | 6 |
| 2024 | Tactile Transfer Learning and Object Recognition With a Multifingered Hand Using Morphology Specific Convolutional Neural NetworksabstractMultifingered robot hands can be extremely effective in physically exploring and recognizing objects, especially if they are extensively covered with distributed tactile sensors. Convolutional neural networks (CNNs) have been proven successful in processing high dimensional data, such as camera images, and are, therefore, very well suited to analyze distributed tactile information as well. However, a major challenge is to organize tactile inputs coming from different locations on the hand in a coherent structure that could leverage the computational properties of the CNN. Therefore, we introduce a morphology-specific CNN (MS-CNN), in which hierarchical convolutional layers are formed following the physical configuration of the tactile sensors on the robot. We equipped a four-fingered Allegro robot hand with several uSkin tactile sensors; overall, the hand is covered with 240 sensitive elements, each one measuring three-axis contact force. The MS-CNN layers process the tactile data hierarchically: at the level of small local clusters first, then each finger, and then the entire hand. We show experimentally that, after training, the robot hand can successfully recognize objects by a single touch, with a recognition rate of over 95%. Interestingly, the learned MS-CNN representation transfers well to novel tasks: by adding a limited amount of data about new objects, the network can recognize nine types of physical properties. Satoshi Funabashi, Gang Yan 0003, Fei Hongyi, Alexander Schmitz, Lorenzo Jamone, Tetsuya Ogata, Shigeki Sugano |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Multi-Timestep-Ahead Prediction with Mixture of Experts for Embodied Question Answering
Kanata Suzuki, Yuya Kamiwano, Naoya Chiba, Hiroki Mori, Tetsuya Ogata |
ICANN (6) | 5 |
| 2023 | Multimodal Time Series Learning of Robots Based on Distributed and Integrated Modalities: Verification with a Simulator and Actual RobotsabstractWe have developed an autonomous robot motion generation model based on distributed and integrated multimodal learning. Since each modality used as a robot's senses, such as image, joint angle, and torque, has a different physical meaning and time characteristic, the generation of autonomous motions using multimodal learning has sometimes failed due to overlearning in one of the modalities. Inspired by the sensory processing of the human brain, our model is based on the processing of each sense performed in the primary somatosensory cortex and the integrated processing of multiple senses in the association cortex and the primary motor cortex. Specifically, the proposed model utilizes two types of recurrent neural networks: sensory RNNs, which learn each sense in a time series, and a union RNN, which communicates with sensory RNNs and learns sensory integration. The simulation results of multiple tasks showed that our model processes multiple modalities appropriately and generates smoother motions with lower jerk than the conventional model. We also demonstrated a chair assembly task by combining fixed motions and autonomous motions with our model. Hideyuki Ichiwara, Kenjiro Yamamoto, Hiroki Mori, Tetsuya Ogata |
ICRA | 5 |
| 2023 | Structured Motion Generation with Predictive Learning: Proposing Subgoal for Long-Horizon ManipulationabstractFor assisting humans in their daily lives, robots need to perform long-horizon tasks, such as tidying up a room or preparing a meal. One effective strategy for handling a long-horizon task is to break it down into short-horizon subgoals, that the robot can execute sequentially. In this paper, we propose extending a predictive learning model using deep neural networks (DNN) with a Subgoal Proposal Module (SPM), with the goal of making such tasks realizable. We evaluate our proposed model in a case-study of a long-horizon task, consisting of cutting and arranging a pizza. This task requires the robot to consider: (1) the order of the subtasks, (2) multiple subtask selection, (3) coordination of dual-arm, and (4) variations within a subtask. The results confirm that the model is able to generalize motion generation to unseen tools and objects arrangement combinations. Furthermore, it significantly reduces the prediction error of the generated motions compared to without the proposed SPM. Finally, we validate the generated motions on the dual-arm robot Nextage Open. See our accompanying video here: https://youtu.be/3hYS2knRm50 Namiko Saito, João Moura 0003, Tetsuya Ogata, Marina Y. Aoyama, Shingo Murata, Shigeki Sugano, Sethu Vijayakumar |
ICRA | 3 |
| 2023 | Force Map: Learning to Predict Contact Force Distribution from VisionabstractWhen humans see a scene, they can roughly imagine the forces applied to objects based on their expe-rience and use them to handle the objects properly. This paper considers transferring this “force-visualization” ability to robots. We hypothesize that a rough force distribution (named “force map”) can be utilized for object manipulation strategies even if accurate force estimation is impossible. Based on this hypothesis, we propose a training method to predict the force map from vision. To investigate this hypothesis, we generated scenes where objects were stacked in bulk through simulation and trained a model to predict the contact force from a single image. We further applied domain randomization to make the trained model function on real images. The experimental results showed that the model trained using only synthetic images could predict approximate patterns representing the contact areas of the objects even for real images. Then, we designed a simple algorithm to plan a lifting direction using the predicted force distribution. We confirmed that using the predicted force distribution contributes to finding natural lifting directions for typical real-world scenes. Furthermore, the evaluation through simulations showed that the disturbance caused to surrounding objects was reduced by 26 % (translation displacement) and by 39 % (angular displacement) for scenes where objects were overlapping. Ryo Hanai, Yukiyasu Domae, Ixchel G. Ramirez, Bruno Leme, Tetsuya Ogata |
IROS | 5 |
| 2022 | Point Cloud Pre-training with Natural 3D StructuresabstractThe construction of 3D point cloud datasets requires a great deal of human effort. Therefore, constructing a large-scale 3D point clouds dataset is difficult. In order to rem-edy this issue, we propose a newly developed point cloud fractal database (PC-FractalDB), which is a novel family of formula-driven supervised learning inspired by fractal geometry encountered in natural 3D structures. Our re-search is based on the hypothesis that we could learn rep-resentations from more real-world 3D patterns than con-ventional 3D datasets by learning fractal geometry. We show how the PC-FractalDB facilitates solving several re-cent dataset-related problems in 3D scene understanding, such as 3D model collection and labor-intensive annotation. The experimental section shows how we achieved the performance rate of up to 61.9% and 59.0% for the Scan-NetV2 and SUN RGB-D datasets, respectively, over the current highest scores obtained with the PointContrast, con-trastive scene contexts (CSC), and RandomRooms. More-over, the PC-FractalDB pre-trained model is especially ef-fective in training with limited data. For example, in 10% of training data on ScanNetV2, the PC-FractalDB pre-trained VoteNet performs at 38.3%, which is +14.8% higher accu-racy than CSC. Of particular note, we found that the pro-posed method achieves the highest results for 3D object de-tection pre-training in limited point cloud data.11Dataset release: https://ryosuke-yamada.github.io/PointCloud-FractalDataBase/ Ryosuke Yamada, Hirokatsu Kataoka, Naoya Chiba, Yukiyasu Domae, Tetsuya Ogata |
CVPR | 5 |
| 2022 | Contact-Rich Manipulation of a Flexible Object based on Deep Predictive Learning using Vision and TactilityabstractWe achieved contact-rich flexible object manipulation, which was difficult to control with vision alone. In the unzipping task we chose as a validation task, the gripper grasps the puller, which hides the bag state such as the direction and amount of deformation behind it, making it difficult to obtain information to perform the task by vision alone. Additionally, the flexible fabric bag state constantly changes during operation, so the robot needs to dynamically respond to the change. However, the appropriate robot behavior for all bag states is difficult to prepare in advance. To solve this problem, we developed a model that can perform contact-rich flexible object manipulation by real-time prediction of vision with tactility. We introduced a point-based attention mechanism for extracting image features, softmax transformation for predicting motions, and convolutional neural network for extracting tactile features. The results of experiments using a real robot arm revealed that our method can realize motions responding to the deformation of the bag while reducing the load on the zipper. Furthermore, using tactility improved the success rate from 56.7% to 93.3% compared with vision alone, demonstrating the effectiveness and high performance of our method. Hideyuki Ichiwara, Kenjiro Yamamoto, Hiroki Mori, Tetsuya Ogata |
ICRA | 5 |
| 2022 | Integrated Learning of Robot Motion and Sentences: Real-Time Prediction of Grasping Motion and Attention based on Language InstructionsabstractWe propose a motion generation model that can achieve robust behavior against environmental changes based on language instructions at a low cost. Conventional robots that communicate with humans use a restricted environment and language to build up a mapping between language and motion, and thus need to prepare a huge training set in order to achieve versatility. Our method trains pairs of language, visual, and motor information of the robot, and generates motions in real-time based on the “attention” of the language instructions. Specifically, the robot generates motions while focusing on the indicated objects by the human when multiple objects are in the field of view. In addition, since position recognition and motion generation of the indicated object are performed in real-time, robust motion generation is possible in response to changes in the object position and lighting conditions. We clarified that features related to the object name and its location are self-organized in the latent (PB: Parametric Bias) space by end-to-end learning of robot motion and sentences. These observations may indicate the importance of integrated learning of robot motion and sentences since such feature representations cannot be obtained by learning motions alone. Hideyuki Ichiwara, Kenjiro Yamamoto, Hiroki Mori, Tetsuya Ogata |
ICRA | 5 |
| 2022 | Guided Visual Attention Model Based on Interactions Between Top-down and Bottom-up Prediction for Robot Pose PredictionabstractDeep robot vision models are widely used for recognizing objects from camera images, but shows poor performance when detecting objects at untrained positions. Although such problem can be alleviated by training with large datasets, the dataset collection cost cannot be ignored. Existing visual attention models tackled the problem by employing a data efficient structure which learns to extract task relevant image areas. However, since the models cannot modify attention targets after training, it is difficult to apply to dynamically changing tasks. This paper proposed a novel Key-Query-Value formulated visual attention model. This model is capable of switching attention targets by externally modifying the Query representations, namely top-down attention. The proposed model is experimented on a simulator and a real-world environment. The model was compared to existing end-to-end robot vision models in the simulator experiments, showing higher performance and data efficiency. In the real-world robot experiments, the model showed high precision along with its scalability and extendibility. Hyogo Hiruma, Hiroki Mori, Tetsuya Ogata |
IECON | 4 |
| 2022 | Use of Action Label in Deep Predictive Learning for Robot ManipulationabstractVarious forms of human knowledge can be explicitly used to enhance deep robot learning from demonstrations. Annotation of subtasks from task segmentation is one type of human symbolism and knowledge. Annotated subtasks can be referred to as action labels, which are more primitive symbols that can be building blocks for more complex human reasoning, like language instructions. However, action labels are not widely used to boost learning processes because of problems that include (1) real-time annotation for online manipulation, (2) temporal inconsistency by annotators, (3) difference in data characteristics of motor commands and action labels, and (4) annotation cost. To address these problems, we propose the Gated Action Motor Predictive Learning (GAMPL) framework to leverage action labels for improved performance. GAMPL has two modules to obtain soft action labels compatible with motor commands and to generate motion. In this study, GAMPL is evaluated for towel-folding manipulation tasks in a real environment with a six degrees-of-freedom (6 DoF) robot and shows improved generalizability with action labels. Kei Kase, Chikara Utsumi, Yukiyasu Domae, Tetsuya Ogata |
IROS | 4 |
| 2021 | A Peg-in-hole Task Strategy for Holes in ConcreteabstractA method that enables an industrial robot to accomplish the peg-in-hole task for holes in concrete is proposed. The proposed method involves slightly detaching the peg from the wall, when moving between search positions, to avoid the negative influence of the concrete’s high friction coefficient. It uses a deep neural network (DNN), trained via reinforcement learning, to effectively find holes with variable shape and surface finish (due to the brittle nature of concrete) without analytical modeling or control parameter tuning. The method uses displacement of the peg toward the wall surface, in addition to force and torque, as one of the inputs of the DNN. Since the displacement increases as the peg gets closer to the hole (due to the chamfered shape of holes in concrete), it is a useful parameter for inputting in the DNN. The proposed method was evaluated by training the DNN on a hole 500 times and attempting to find 12 unknown holes. The results of the evaluation show the DNN enabled a robot to find the unknown holes with average success rate of 96.1% and average execution time of 12.5 seconds. Additional evaluations with random initial positions and a different type of peg demonstrate the trained DNN can generalize well to different conditions. Analyses of the influence of the peg displacement input showed the success rate of the DNN is increased by utilizing this parameter. These results validate the proposed method in terms of its effectiveness and applicability to the construction industry. André Yuji Yasutomi, Hiroki Mori, Tetsuya Ogata |
ICRA | 3 |
| 2021 | Comparison of Consolidation Methods for Predictive Learning of Time Series
Ryoichi Nakajo, Tetsuya Ogata |
IEA/AIE (1) | 2 |
| 2021 | Binary Neural Network in Robotic Manipulation: Flexible Object Manipulation for Humanoid Robot Using Partially Binarized Auto-Encoder on FPGAabstractA neural network based flexible object manipulation system for a humanoid robot on FPGA is proposed. Although the manipulations of flexible objects using robots attract ever increasing attention since these tasks are the basic and essential activities in our daily life, it has been put into practice only recently with the help of deep neural networks. However such systems have relied on GPU accelerators, which cannot be implemented into the space limited robotic body. Although field programmable gate arrays (FPGAs) are known to be energy efficient and suitable for embedded systems, the model size should be drastically reduced since FPGAs have limited on-chip memory. To this end, we propose "partially" binarized deep convolutional auto-encoder technique, where only an encoder part is binarized to compress model size without degrading the inference accuracy. The model implemented on Xilinx ZCU102 achieves 41.1 frames per second with a power consumption of 3.1 W, which corresponds to 10× and 3.7× improvements from the systems implemented on Core i7 6700K and RTX 2080 Ti, respectively. Satoshi Ohara, Tetsuya Ogata, Hiromitsu Awano |
IROS | 2 |
| 2021 | In-air Knotting of Rope using Dual-Arm Robot based on Deep LearningabstractIn this study, we report the successful execution of in-air knotting of rope using a dual-arm two-finger robot based on deep learning. Owing to its flexibility, the state of the rope was in constant flux during the operation of the robot. This required the robot control system to dynamically correspond to the state of the object at all times. However, a manual description of appropriate robot motions corresponding to all object states is difficult to be prepared in advance. To resolve this issue, we constructed a model that instructed the robot to perform bowknots and overhand knots based on two deep neural networks trained using the data gathered from its sensorimotor, including visual and proximity sensors. The resultant model was verified to be capable of predicting the appropriate robot motions based on the sensory information available online. In addition, we designed certain task motions based on the Ian knot method using the dual-arm two-fingers robot. The designed knotting motions do not require a dedicated workbench or robot hand, thereby enhancing the versatility of the proposed method. Finally, experiments were performed to estimate the knotting performance of the real robot while executing overhand knots and bowknots on rope and its success rate. The experimental results established the effectiveness and high performance of the proposed method. Kanata Suzuki, Momomi Kanamura, Yuki Suga, Hiroki Mori, Tetsuya Ogata |
IROS | 5 |
| 2021 | Paradoxical sensory reactivity induced by functional disconnection in a robot model of neurodevelopmental disorderabstractNeurodevelopmental disorders are characterized by heterogeneous and non-specific nature of their clinical symptoms. In particular, hyper- and hypo-reactivity to sensory stimuli are diagnostic features of autism spectrum disorder and are reported across many neurodevelopmental disorders. However, computational mechanisms underlying the unusual paradoxical behaviors remain unclear. In this study, using a robot controlled by a hierarchical recurrent neural network model with predictive processing and learning mechanism, we simulated how functional disconnection altered the learning process and subsequent behavioral reactivity to environmental change. The results show that, through the learning process, long-range functional disconnection between distinct network levels could simultaneously lower the precision of sensory information and higher-level prediction. The alteration caused a robot to exhibit sensory-dominated and sensory-ignoring behaviors ascribed to sensory hyper- and hypo-reactivity, respectively. As long-range functional disconnection became more severe, a frequency shift from hyporeactivity to hyperreactivity was observed, paralleling an early sign of autism spectrum disorder. Furthermore, local functional disconnection at the level of sensory processing similarly induced hyporeactivity due to low sensory precision. These findings suggest a computational explanation for paradoxical sensory behaviors in neurodevelopmental disorders, such as coexisting hyper- and hypo-reactivity to sensory stimulus. A neurorobotics approach may be useful for bridging various levels of understanding in neurodevelopmental disorders and providing insights into mechanisms underlying complex clinical symptoms. Hayato Idei, Shingo Murata, Yuichi Yamashita, Tetsuya Ogata |
Neural Networks | 4 |
| 2020 | Stable Deep Reinforcement Learning Method by Predicting Uncertainty in Rewards as a Subtask
Kanata Suzuki, Tetsuya Ogata |
ICONIP (2) | 2 |
| 2020 | Transferable Task Execution from Pixels through Deep Planning Domain LearningabstractWhile robots can learn models to solve many manipulation tasks from raw visual input, they cannot usually use these models to solve new problems. On the other hand, symbolic planning methods such as STRIPS have long been able to solve new problems given only a domain definition and a symbolic goal, but these approaches often struggle on the real world robotic tasks due to the challenges of grounding these symbols from sensor data in a partially-observable world. We propose Deep Planning Domain Learning (DPDL), an approach that combines the strengths of both methods to learn a hierarchical model. DPDL learns a high-level model which predicts values for a large set of logical predicates consisting of the current symbolic world state, and separately learns a low-level policy which translates symbolic operators into executable actions on the robot. This allows us to perform complex, multistep tasks even when the robot has not been explicitly trained on them. We show our method on manipulation tasks in a photorealistic kitchen scenario. Kei Kase, Chris Paxton 0001, Hammad Mazhar, Tetsuya Ogata, Dieter Fox |
ICRA | 4 |
| 2020 | Stable In-Grasp Manipulation with a Low-Cost Robot Hand by Using 3-Axis Tactile Sensors with a CNNabstractThe use of tactile information is one of the most important factors for achieving stable in-grasp manipulation. Especially with low-cost robotic hands that provide low-precision control, robust in-grasp manipulation is challenging. Abundant tactile information could provide the required feed-back to achieve reliable in-grasp manipulation also in such cases. In this research, soft distributed 3-axis skin sensors ("uSkin") and 6-axis F/T (force/torque) sensors were mounted on each fingertip of an Allegro Hand to provide rich tactile information. These sensors yielded 78 measurements for each fingertip (72 measurements from the uSkin and 6 measurements from the 6-axis F/T sensor). However, such high-dimensional tactile information can be difficult to process because of the complex contact states between the grasped object and the fingertips. Therefore, a convolutional neural network (CNN) was employed to process the tactile information. In this paper, we explored the importance of the different sensors for achieving in-grasp manipulation. Successful in-grasp manipulation with untrained daily objects was achieved when both 3-axis uSkin and 6-axis F/T information was provided and when the information was processed using a CNN. Satoshi Funabashi, Tomoki Isobe, Shun Ogasa, Tetsuya Ogata, Alexander Schmitz, Tito Pradhono Tomo, Shigeki Sugano |
IROS | 4 |
| 2020 | Variable In-Hand Manipulations for Tactile-Driven Robot Hand via CNN-LSTMabstractPerforming various in-hand manipulation tasks, without learning each individual task, would enable robots to act more versatile, while reducing the effort for training. However, in general it is difficult to achieve stable in-hand manipulation, because the contact state between the fingertips becomes difficult to model, especially for a robot hand with anthropomorphically shaped fingertips. Rich tactile feedback can aid the robust task execution, but on the other hand it is challenging to process high-dimensional tactile information. In the current paper we use two fingers of the Allegro hand, and each fingertip is anthropomorphically shaped and equipped not only with 6-axis force-torque (F/T) sensors, but also with uSkin tactile sensors, which provide 24 tri-axial measurements per fingertip. A convolutional neural network is used to process the high dimensional uSkin information, and a long short-term memory (LSTM) handles the time-series information. The network is trained to generate two different motions ("twist" and "push"). The desired motion is provided as a task-parameter to the network, with twist defined as -1 and push as +1. When values between -1 and +1 are used as the task parameter, the network is able to generate untrained motions in-between the two trained motions. Thereby, we can achieve multiple untrained manipulations, and can achieve robustness with high-dimensional tactile feedback. Satoshi Funabashi, Shun Ogasa, Tomoki Isobe, Tetsuya Ogata, Alexander Schmitz, Tito Pradhono Tomo, Shigeki Sugano |
IROS | 4 |
| 2020 | Wiping 3D-objects using Deep Learning Model based on Image/Force/Joint InformationabstractWe propose a deep learning model for a robot to wipe 3D-objects. Wiping of 3D-objects requires recognizing the shapes of objects and planning the motor angle adjustments for tracing the objects. Unlike previous research, our learning model does not require pre-designed computational models of target objects. The robot is able to wipe the objects to be placed by using image, force, and arm joint information. We evaluate the generalization ability of the model by confirming that the robot handles untrained cube and bowl shaped-objects. We also find that it is necessary to use both image and force information to recognize the shape of and wipe 3D objects consistently by comparing changes in the input sensor data to the model. To our knowledge, this is the first work enabling a robot to use learning sensorimotor information alone to trace various unknown 3D-shape. Namiko Saito, Tetsuya Ogata, Hiroki Mori, Shigeki Sugano |
IROS | 3 |
| 2020 | HATSUKI : An anime character like robot figure platform with anime-style expressions and imitation learning based action generationabstractJapanese character figurines are popular and have a pivot position in Otaku culture. Although numerous robots have been developed, few have focused on otaku-culture or on embodying anime character figurines. Therefore, we take the first steps to bridge this gap by developing Hatsuki, which is a humanoid robot platform with anime based design. Hatsuki's novelty lies in its aesthetic design, 2D facial expressions, and anime-style behaviors that allows Hatsuki to deliver rich interaction experiences resembling anime-characters. We explain our design implementation process of Hatsuki, followed by our evaluations. In order to explore user impressions and opinions towards Hatsuki, we conducted a questionnaire in the world's largest anime-figurine event. The results indicate that participants were generally very satisfied with Hatsuki's design, and proposed various use case scenarios and deployment contexts for Hatsuki. The second evaluation focused on imitation learning, as such a method can provide better interaction ability in the real world and generate rich, context-adaptive behaviors in different situations. We made Hatsuki learn 11 actions, combining voice, facial expressions and motions, through the neural network based policy model with our proposed interface. Results show our approach was successfully able to generate the actions through self-organized contexts, which shows the potential for generalizing our approach in further actions under different contexts. Lastly, we present our future research direction for Hatsuki and provide our conclusion. Pin-Chu Yang, Mohammed AlSada, Chang-Chieh Chiu, Kevin Kuo, Tito Pradhono Tomo, Kanata Suzuki, Nelson Enrique Yalta Soplin, Kuo-Hao Shu, Tetsuya Ogata |
RO-MAN | 9 |
| 2019 | Achieving Human-Robot Collaboration with Dynamic Goal Inference by Gradient Descent
Shingo Murata, Wataru Masuda, Hiroaki Arie, Tetsuya Ogata, Shigeki Sugano |
ICONIP (2) | 5 |
| 2019 | Morphology-Specific Convolutional Neural Networks for Tactile Object Recognition with a Multi-Fingered HandabstractDistributed tactile sensors on multi-fingered hands can provide high-dimensional information for grasping objects, but it is not clear how to optimally process such abundant tactile information. The current paper explores the possibility of using a morphology-specific convolutional neural network (MS-CNN). uSkin tactile sensors are mounted on an Allegro Hand, which provides 720 force measurements (15 patches of uSkin modules with 16 triaxial force sensors each) in addition to 16 joint angle measurements. Consecutive layers in the CNN get input from parts of one finger segment, one finger, and the whole hand. Since the sensors give 3D (x, y, z) vector tactile information, inputs with 3 channels (x, y and z) are used in the first layer, based on the idea of such inputs for RGB images from cameras. Overall, the layers are combined, resulting in the building of a tactile map based on the relative position of the tactile sensors on the hand. Seven different combination variations were evaluated, and an over-95% object recognition rate with 20 objects was achieved, even though only one random time instance from a repeated squeezing motion of an object in an unknown pose within the hand was used as input. Satoshi Funabashi, Gang Yan 0003, Andreas Geier, Alexander Schmitz, Tetsuya Ogata, Shigeki Sugano |
ICRA | 5 |
| 2019 | Weakly-Supervised Deep Recurrent Neural Networks for Basic Dance Step GenerationabstractSynthesizing human's movements such as dancing is a flourishing research field which has several applications in computer graphics. Recent studies have demonstrated the advantages of deep neural networks (DNNs) for achieving remarkable performance in motion and music tasks with little effort for feature pre-processing. However, applying DNNs for generating dance to a piece of music is nevertheless challenging, because of 1) DNNs need to generate large sequences while mapping the music input, 2) the DNN needs to constraint the motion beat to the music, and 3) DNNs require a considerable amount of hand-crafted data. In this study, we propose a weakly supervised deep recurrent method for real-time basic dance generation with audio power spectrum as input. The proposed model employs convolutional layers and a multilayered Long Short-Term memory (LSTM) to process the audio input. Then, another deep LSTM layer decodes the target dance sequence. Notably, this end-to-end approach has 1) an auto-conditioned decode configuration that reduces accumulation of feedback error of large dance sequence, 2) uses a contrastive cost function to regulate the mapping between the music and motion beat, and 3) trains with weak labels generated from the motion beat, reducing the amount of hand-crafted data. We evaluate the proposed network based on i) the similarities between generated and the baseline dancer motion with a cross entropy measure for large dance sequences, and ii) accurate timing between the music and motion beat with an F-measure. Experimental results revealed that, after training using a small dataset, the model generates basic dance steps with low cross entropy and maintains an F-measure score similar to that of a baseline dancer. Nelson Enrique Yalta Soplin, Shinji Watanabe 0001, Kazuhiro Nakadai, Tetsuya Ogata |
IJCNN | 4 |
| 2019 | A Bi-directional Multiple Timescales LSTM Model for Grounding of Actions and VerbsabstractIn this paper we present a neural architecture to learn a bi-directional mapping between actions and language. We implement a Multiple Timescale Long Short-Term Memory (MT-LSTM) network comprised of 7 layers with different timescale factors, to connect actions to language without explicitly learning an intermediate representation. Instead, the model self-organizes such representations at the level of a slow-varying latent layer, linking action branch and language branch at the center. We train the model in a bi-directional way, learning how to produce a sentence from a certain action sequence input and, simultaneously, how to generate an action sequence given a sentence as input. Furthermore we show this model preserves some of the generalization behaviour of Multiple Timescale Recurrent Neural Networks (MTRNN) in generating sentences and actions that were not explicitly trained. We compare this model with a number of different baseline models, confirming the importance of both the bi-directional training and the multiple timescales architecture. Finally, the network was evaluated on motor actions performed by an iCub robot and their corresponding letter-based description. The results of these experiments are presented at the end of the paper. Alexandre Antunes, Alban Laflaquière, Tetsuya Ogata, Angelo Cangelosi |
IROS | 3 |
| 2019 | Learning Multiple Sensorimotor Units to Complete Compound Tasks using an RNN with Multiple AttractorsabstractAs the complexity of the robot's tasks increases, we can consider many general tasks in a compound form that consists of shorter tasks. Therefore, for robots to generate various tasks, they need to be able to execute shorter tasks in succession, appropriately to the situation. With the design principle to construct the architecture for robots to execute complex tasks compounded with multiple subtasks, this study proposes a visuomotor-control framework with the characteristics of a state machine to train shorter tasks as sensorimotor units. The design procedure of training framework consists of 4 steps: (1) segment entire task into appropriate subtasks, (2) define subtasks as states and transitions in a state machine, (3) collect subtasks data, and (4) train neural networks: (a) autoencoder to extract visual features, (b) a single recurrent neural network to generate subtasks to realize a pseud-state-machine model with a constraint in hidden values. We implemented this framework on two different robots to allow their performance of repetitive tasks with error-recovery motion, subsequently, confirming the ability of the robot to switch the sensorimotor units from visual input at the attractors of the hidden values created by the constraint. Kei Kase, Ryoichi Nakajo, Hiroki Mori, Tetsuya Ogata |
IROS | 4 |
| 2019 | Large-scale Data Collection for Goal-directed Drawing Task with Self-report Psychiatric Symptom Questionnaires via CrowdsourcingabstractDrawing is a representative human cognitive ability and may mirror cognitive characteristics including those associated with psychiatric symptoms. Therefore, analysis of drawing data collected from various populations such as healthy people and psychiatric patients may be beneficial for better understanding human cognition. However, collecting such large-scale data about the relationship between drawing and cognitive/personality traits offline-in a laboratory-is a difficult issue. To overcome this issue, we devised a novel experimental paradigm involving a goal-directed drawing task conducted online-on the eb-with participants recruited via a crowdsourcing platform. With the assistance of 1155 participants with differing levels of psychiatric symptoms, we collected a total of 194,040 trajectory data and answers to seven different self-report psychiatric symptom questionnaires comprising 181 items. We visualized the collected trajectory data and performed an exploratory factor analysis on the correlation matrix of the psychiatric symptom questionnaire items. Our results suggest that there were associations between psychiatric symptoms represented by specific psychiatric factors and atypical behavior observed while performing the goal-directed drawing task. This indicates the efficacy of a dimensional approach to large-scale online experiments with respect to clinical psychiatry. Shingo Murata, Hikaru Yanagida, Kentaro Katahira, Shinsuke Suzuki, Tetsuya Ogata, Yuichi Yamashita |
SMC | 5 |
| 2018 | Deep 3D Pose Dictionary: 3D Human Pose Estimation from Single RGB Image Using Deep Convolutional Neural Network
Reda Elbasiony, Walid Gomaa 0001, Tetsuya Ogata |
ICANN (3) | 3 |
| 2018 | Put-in-Box Task Generated from Multiple Discrete Tasks by aHumanoid Robot Using Deep LearningabstractFor robots to have a wide range of applications, they must be able to execute numerous tasks. However, recent studies into robot manipulation using deep neural networks (DNN) have primarily focused on single tasks. Therefore, we investigate a robot manipulation model that uses DNNs and can execute long sequential dynamic tasks by performing multiple short sequential tasks at appropriate times. To generate compound tasks, we propose a model comprising two DNNs: a convolutional autoencoder that extracts image features and a multiple timescale recurrent neural network (MTRNN) to generate motion. The internal state of the MTRNN is constrained to have similar values at the initial and final motion steps; thus, motions can be differentiated based on the initial image input. As an example compound task, we demonstrate that the robot can generate a “Put-In-Box” task that is divided into three subtasks: open the box, grasp the object and put it into the box, and close the box. The subtasks were trained as discrete tasks, and the connections between each subtask were not trained. With the proposed model, the robot could perform the Put-In-Box task by switching among subtasks and could skip or repeat subtasks depending on the situation. Kei Kase, Kanata Suzuki, Pin-Chu Yang, Hiroki Mori, Tetsuya Ogata |
ICRA | 5 |
| 2018 | End-to-End Visuomotor Learning of Drawing Sequences using Recurrent Neural NetworksabstractDrawing is one of the complex cognitive abilities of humans. Cognitive neuropsychological studies have attempted to develop models that can explain the observations of the drawing behavior. These models exhibit limitations to reproduce the drawing behaviors because of individual factors that are related to the drawing style or non-reproducibility of motions. A constructive approach provides another methodology to investigate the complex systems by constructing models that can reproducibly replicate the behaviors. In this study, we focus on an ability to reuse the integrated visuomotor memory of drawing to associate the drawing motion from an image. Existing computational models of drawing have not considered the visual information in hand-drawn pictures. Therefore, we propose a dynamical model of the visuomotor process of drawing. The proposed model does not require any prior knowledge of the process such as the pre-designed shape primitives or the image processing algorithms. The proposed model is implemented by utilizing a recurrent neural network that learns the visuomotor transition of the drawing process. The association of the model's drawing motion by reusing the obtained memory can be obtained by minimizing the prediction error of the image. By performing simulator experiments, the proposed model demonstrates its association ability in case of pictures that comprise multiple lines. Kazuma Sasaki, Tetsuya Ogata |
IJCNN | 2 |
| 2018 | AFA-PredNet: The Action Modulation Within Predictive CodingabstractThe predictive processing (PP) hypothesizes that the predictive inference of our sensorimotor system is encoded implicitly in the regularities between perception and action. We propose a neural architecture in which such regularities of active inference are encoded hierarchically. We further suggest that this encoding emerges during the embodied learning process when the appropriate action is selected to minimize the prediction error in perception. Therefore, this predictive stream in the sensorimotor loop is generated in a top-down manner. Specifically, it is constantly modulated by the motor actions and is updated by the bottom-up prediction error signals. In this way, the top-down prediction originally comes from the prior experience from both perception and action representing the higher levels of this hierarchical cognition. In our proposed embodied model, we extend the PredNet Network, a hierarchical predictive coding network, with the motor action units implemented by a multi-layer perceptron network (MLP) to modulate the network top-down prediction. Two experiments, a minimalistic world experiment, and a mobile robot experiment are conducted to evaluate the proposed model in a qualitative way. In the neural representation, it can be observed that the causal inference of predictive percept from motor actions can be also observed while the agent is interacting with the environment. Junpei Zhong, Angelo Cangelosi, Xinzheng Zhang 0001, Tetsuya Ogata |
IJCNN | 4 |
| 2017 | Mixing Actual and Predicted Sensory States Based on Uncertainty Estimation for Flexible and Robust Robot Behavior
Shingo Murata, Wataru Masuda, Saki Tomioka, Tetsuya Ogata, Shigeki Sugano |
ICANN (1) | 4 |
| 2017 | Learning of Labeling Room Space for Mobile Robots Based on Visual Motor Experience
Tatsuro Yamada, Saki Ito, Hiroaki Arie, Tetsuya Ogata |
ICANN (1) | 4 |
| 2017 | Toward abstraction from multi-modal data: Empirical studies on multiple time-scale recurrent modelsabstractThe abstraction tasks are challenging for multi-modal sequences as they require a deeper semantic understanding and a novel text generation for the data. Although the recurrent neural networks (RNN) can be used to model the context of the time-sequences, in most cases the long-term dependencies of multi-modal data make the back-propagation through time training of RNN tend to vanish in the time domain. Recently, inspired from Multiple Time-scale Recurrent Neural Network (MTRNN) [1], an extension of Gated Recurrent Unit (GRU), called Multiple Time-scale Gated Recurrent Unit (MTGRU), has been proposed [2] to learn the long-term dependencies in natural language processing. Particularly it is also able to accomplish the abstraction task for paragraphs given that the time constants are well defined. In this paper, we compare the MTRNN and MTGRU in terms of its learning performances as well as their abstraction representation on higher level (with a slower neural activation). This was done by conducting two studies based on a smaller dataset (two-dimension time sequences from non-linear functions) and a relatively large data-set (43-dimension time sequences from iCub manipulation tasks with multi-modal data). We conclude that gated recurrent mechanisms may be necessary for learning long-term dependencies in large dimension multi-modal data-sets (e.g. learning of robot manipulation), even when natural language commands was not involved. But for smaller learning tasks with simple time-sequences, generic version of recurrent models, such as MTRNN, were sufficient to accomplish the abstraction task. Junpei Zhong, Angelo Cangelosi, Tetsuya Ogata |
IJCNN | 3 |
| 2017 | Learning to Perceive the World as Probabilistic or Deterministic via Interaction With Others: A Neuro-Robotics ExperimentabstractWe suggest that different behavior generation schemes, such as sensory reflex behavior and intentional proactive behavior, can be developed by a newly proposed dynamic neural network model, named stochastic multiple timescale recurrent neural network (S-MTRNN). The model learns to predict subsequent sensory inputs, generating both their means and their uncertainty levels in terms of variance (or inverse precision) by utilizing its multiple timescale property. This model was employed in robotics learning experiments in which one robot controlled by the S-MTRNN was required to interact with another robot under the condition of uncertainty about the other's behavior. The experimental results show that self-organized and sensory reflex behavior-based on probabilistic prediction-emerges when learning proceeds without a precise specification of initial conditions. In contrast, intentional proactive behavior with deterministic predictions emerges when precise initial conditions are available. The results also showed that, in situations where unanticipated behavior of the other robot was perceived, the behavioral context was revised adequately by adaptation of the internal neural dynamics to respond to sensory inputs during sensory reflex behavior generation. On the other hand, during intentional proactive behavior generation, an error regression scheme by which the internal neural activity was modified in the direction of minimizing prediction errors was needed for adequately revising the behavioral context. These results indicate that two different ways of treating uncertainty about perceptual events in learning, namely, probabilistic modeling and deterministic modeling, contribute to the development of different dynamic neuronal structures governing the two types of behavior generation schemes. Shingo Murata, Yuichi Yamashita, Hiroaki Arie, Tetsuya Ogata, Shigeki Sugano, Jun Tani |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2016 | Classification of Photo and Sketch Images Using Convolutional Neural Networks
Kazuma Sasaki, Madoka Yamakawa, Kana Sekiguchi, Tetsuya Ogata |
ICANN (2) | 4 |
| 2016 | Body Model Transition by Tool Grasping During Motor Babbling Using Deep Learning and RNN
Kuniyuki Takahashi, Hadi Tjandra, Tetsuya Ogata, Shigeki Sugano |
ICANN (1) | 3 |
| 2016 | Dynamical Linking of Positive and Negative Sentences to Goal-Oriented Robot Behavior by Hierarchical RNN
Tatsuro Yamada, Shingo Murata, Hiroaki Arie, Tetsuya Ogata |
ICANN (1) | 4 |
| 2016 | Self and Non-self Discrimination Mechanism Based on Predictive Learning with Estimation of Uncertainty
Ryoichi Nakajo, Maasa Takahashi, Shingo Murata, Hiroaki Arie, Tetsuya Ogata |
ICONIP (4) | 5 |
| 2015 | Efficient Motor Babbling Using Variance Predictions from a Recurrent Neural Network
Kuniyuki Takahashi, Kanata Suzuki, Tetsuya Ogata, Hadi Tjandra, Shigeki Sugano |
ICONIP (3) | 3 |
| 2015 | Neural network based model for visual-motor integration learning of robot's drawing behavior: Association of a drawing motion from a drawn imageabstractIn this study, we propose a neural network based model for learning a robot's drawing sequences in an unsupervised manner. We focus on the ability to learn visual-motor relationships, which can work as a reusable memory in association of drawing motion from a picture image. Assuming that a humanoid robot can draw a shape on a pen tablet, the proposed model learns drawing sequences, which comprises drawing motion and drawn picture image frames. To learn raw pixel data without any given specific features, we utilized a deep neural network for compressing large dimensional picture images and a continuous time recurrent neural network for integration of motion and picture images. To confirm the ability of the proposed model, we performed an experiment for learning 15 sequences comprising three types of shapes. The model successfully learns all the sequences and can associate a drawing motion from a not trained picture image and a trained picture with similar success. We also show that the proposed model self-organizes its behavior according to types shapes. Kazuma Sasaki, Hadi Tjandra, Kuniaki Noda, Kuniyuki Takahashi, Tetsuya Ogata |
IROS | 5 |
| 2015 | Effective motion learning for a flexible-joint robot using motor babblingabstractWe propose a method for realizing effective dynamic motion learning in a flexible-joint robot using motor babbling. Flexible-joint robots have recently attracted attention because of their adaptiveness, safety, and, in particular, dynamic motions. It is difficult to control robots that require dynamic motion. In past studies, attractors and oscillators were designed as motion primitives of an assumed task in advance. However, it is difficult to adapt to unintended environmental changes using such methods. To overcome this problem, we use a recurrent neural network (RNN) that does not require predetermined parameters. In this research, we propose a method for facilitating effective learning. First, a robot learns simple motions via motor babbling, acquiring body dynamics using a recurrent neural network (RNN). Motor babbling is the process of movement that infants use to acquire their own body dynamics during their early days. Next, the robot learns additional motions required for a target task using the acquired body dynamics. For acquiring these body dynamics, the robot uses motor babbling with its redundant flexible joints to learn motion primitives. This redundancy implies that there are numerous possible motion patterns. In comparison to a basic learning task, the motion primitives are simply modified to adjust to the task. Next, we focus on the types of motions used in motor babbling. We classify the motions into two motion types, passive motion and active motion. Passive motion involves inertia without any torque input, whereas active motion involves a torque input. The robot acquires body dynamics from the passive motion and a means of torque generation from the active motion. As a result, we demonstrate the importance of performing prior learning via motor babbling before learning a task. In addition, task learning is made more efficient by dividing the motion into two types of motor babbling patterns. Kuniyuki Takahashi, Tetsuya Ogata, Hiroki Yamada, Hadi Tjandra, Shigeki Sugano |
IROS | 2 |
| 2015 | Attractor representations of language-behavior structure in a recurrent neural network for human-robot interactionabstractIn recent years there has been increased interest in studies that explore integrative learning of language and other modalities by using neural network models. However, for practical application to human-robot interaction, the acquired semantic structure between language and meaning has to be available immediately and repeatably whenever necessary, just as in everyday communication. As a solution to this problem, this study proposes a method in which a recurrent neural network self-organizes cyclic attractors that reflect semantic structure and represent interaction flows in its internal dynamics. To evaluate this method we design a simple task in which a human verbally directs a robot, which responds appropriately. Training the network with training data that represent the interaction series, the cyclic attractors that reflect the semantic structure is self-organized. The network first receives a verbal direction, and its internal state moves according to the first half of the cyclic attractors with branch structures corresponding to semantics. After that, the internal state reaches a potential to generate appropriate behavior. Finally, the internal state moves to the second half and converges on the initial point of the cycle while generating the appropriate behavior. By self-organizing such an internal structure in its forward dynamics, the model achieves immediate and repeatable response to linguistic directions. Furthermore, the network self-organizes a fixed-point attractor, and so able to wait for directions. It can thus repeat the interaction flexibly without explicit turn-taking signs. Tatsuro Yamada, Shingo Murata, Hiroaki Arie, Tetsuya Ogata |
IROS | 4 |
| 2015 | Audio-visual speech recognition using deep learningabstractAudio-visual speech recognition (AVSR) system is thought to be one of the most promising solutions for reliable speech recognition, particularly when the audio is corrupted by noise. However, cautious selection of sensory features is crucial for attaining high recognition performance. In the machine-learning community, deep learning approaches have recently attracted increasing attention because deep neural networks can effectively extract robust latent features that enable various recognition algorithms to demonstrate revolutionary generalization capabilities under diverse application conditions. This study introduces a connectionist-hidden Markov model (HMM) system for noise-robust AVSR. First, a deep denoising autoencoder is utilized for acquiring noise-robust audio features. By preparing the training data for the network with pairs of consecutive multiple steps of deteriorated audio features and the corresponding clean features, the network is trained to output denoised audio features from the corresponding features deteriorated by noise. Second, a convolutional neural network (CNN) is utilized to extract visual features from raw mouth area images. By preparing the training data for the CNN as pairs of raw images and the corresponding phoneme label outputs, the network is trained to predict phoneme labels from the corresponding mouth area input images. Finally, a multi-stream HMM (MSHMM) is applied for integrating the acquired audio and visual HMMs independently trained with the respective features. By comparing the cases when normal and denoised mel-frequency cepstral coefficients (MFCCs) are utilized as audio features to the HMM, our unimodal isolated word recognition results demonstrate that approximately 65 % word recognition rate gain is attained with denoised MFCCs under 10 dB signal-to-noise-ratio (SNR) for the audio signal input. Moreover, our multimodal isolated word recognition results utilizing MSHMM with denoised MFCCs and acquired visual features demonstrate that an additional word recognition rate gain is attained for the SNR conditions below 10 dB. Kuniaki Noda, Yuki Yamaguchi, Kazuhiro Nakadai, Hiroshi G. Okuno, Tetsuya Ogata |
Appl. Intell. | 5 |
| 2014 | Learning and Recognition of Multiple Fluctuating Temporal Patterns Using S-CTRNN
Shingo Murata, Hiroaki Arie, Tetsuya Ogata, Jun Tani, Shigeki Sugano |
ICANN | 3 |
| 2014 | Tool-Body Assimilation Model Based on Body Babbling and a Neuro-Dynamical System for Motion Generation
Kuniyuki Takahashi, Tetsuya Ogata, Hadi Tjandra, Shingo Murata, Hiroaki Arie, Shigeki Sugano |
ICANN | 2 |
| 2014 | Insertion of pause in drawing from babbling for robot's developmental imitation learningabstractIn this paper, we present a method to improve a robot's imitation performance in a drawing scenario by inserting pauses in motion. Human's drawing skills are said to develop through five stages: 1) Scribbling, 2) Fortuitous Realism, 3) Failed Realism, 4) Intellectual Realism, and 5) Visual Realism. We focus on stages 1) and 3) for creating our system, each corresponding to body babbling and imitation learning, respectively. For stage 1), the robot randomly moves its arm to associate robot's arm dynamics with the drawing result. Presuming that the robot has no knowledge about its own dynamics, the robot learns its body dynamics in this stage. For stage 3), we consider a scenario where a robot would imitate a human's drawing motion. Upon creating the system, we focus on the motionese phenomenon, which is one of the key factors for discussing acquisition of a skill through a human parent-child interaction. In motionese, the parent would first show each action elaborately to the child, when teaching a skill. As the child starts to improve, the parent's actions would be simplified. Likewise in our scenario, the human would first insert pauses during the drawing motions where the direction of drawing changes (i.e. corners). As the robot's imitation learning of drawing converges, the human would change to drawing without pauses. The experimental results show that insertion of pause in drawing imitation scenarios greatly improves the robot's drawing performance. Shun Nishide, Keita Mochizuki, Hiroshi G. Okuno, Tetsuya Ogata |
ICRA | 4 |
| 2014 | Lipreading using convolutional neural network
Kuniaki Noda, Yuki Yamaguchi, Kazuhiro Nakadai, Hiroshi G. Okuno, Tetsuya Ogata |
INTERSPEECH | 5 |
| 2013 | Multimodal integration learning of object manipulation behaviors using deep neural networksabstractThis paper presents a novel computational approach for modeling and generating multiple object manipulation behaviors by a humanoid robot. The contribution of this paper is that deep learning methods are applied not only for multimodal sensor fusion but also for sensory-motor coordination. More specifically, a time-delay deep neural network is applied for modeling multiple behavior patterns represented with multi-dimensional visuomotor temporal sequences. By using the efficient training performance of Hessian-free optimization, the proposed mechanism successfully models six different object manipulation behaviors in a single network. The generalization capability of the learning mechanism enables the acquired model to perform the functions of cross-modal memory retrieval and temporal sequence prediction. The experimental results show that the motion patterns for object manipulation behaviors are successfully generated from the corresponding image sequence, and vice versa. Moreover, the temporal sequence prediction enables the robot to interactively switch multiple behaviors in accordance with changes in the displayed objects. Kuniaki Noda, Hiroaki Arie, Yuki Suga, Tetsuya Ogata |
IROS | 4 |
| 2013 | Developmental Human-Robot Imitation Learning of Drawing with a Neuro Dynamical SystemabstractThis paper mainly deals with robot developmental learning on drawing and discusses the influences of physical embodiment to the task. Humans are said to develop their drawing skills through five phases: 1) Scribbling, 2) Fortuitous Realism, 3) Failed Realism, 4) Intellectual Realism, 5) Visual Realism. We implement phases 1) and 3) into the humanoid robot NAO, holding a pen, using a neuro dynamical model, namely Multiple Timescales Recurrent Neural Network (MTRNN). For phase 1), we used random arm motion of the robot as body babbling to associate motor dynamics with pen position dynamics. For phase 3), we developed incremental imitation learning to imitate and develop the robot's drawing skill using basic shapes: circle, triangle, and rectangle. We confirmed two notable features from the experiment. First, the drawing was better performed for shapes requiring arm motions used in babbling. Second, performance of clockwise drawing of circle was good from beginning, which is a similar phenomenon that can be observed in human development. The results imply the capability of the model to create a developmental robot relating to human development. Keita Mochizuki, Shun Nishide, Hiroshi G. Okuno, Tetsuya Ogata |
SMC | 4 |
| 2013 | Intersensory Causality Modeling Using Deep Neural NetworksabstractOur brain is known to enhance perceptual precision and reduce ambiguity about sensory environment by integrating multiple sources of sensory information acquired from different modalities, such as vision, auditory and somatic sensation. From an engineering perspective, building a computational model that replicates this ability to integrate multimodal information and to self-organize the causal dependency among them, represents one of the central challenges in robotics. In this study, we propose such a model based on a deep learning framework and we evaluate the proposed model by conducting a bell ring task using a small humanoid robot. Our experimental results demonstrate that (1) the cross-modal memory retrieval function of the proposed method succeeds in generating visual sequence from the corresponding sound and bell ring motion, and (2) the proposed method leads to accurate causal dependencies among the sensory-motor sequence. Kuniaki Noda, Hiroaki Arie, Yuki Suga, Tetsuya Ogata |
SMC | 4 |
| 2012 | Initialization-robust multipitch estimation based on latent harmonic allocation using overtone corpusabstractWe present a new method for modeling the overtone structures of musical instruments that uses an overtone corpus generated using a MIDI synthesizer. Since multipitch estimation requires a joint estimation of F0's and their overtone structures, one of the most important problems is the overtone structure modeling. Latent harmonic allocation (LHA), a promising multipitch estimation method, is difficult to use for various applications because it requires appropriate prior distributions of the overtone structures, which cannot be determined from statistical evidence. Our method uses an overtone corpus to avoid the problem of setting prior distributions and instead restricts the lower and upper bounds of each overtone weight. The bounds are determined from reference signals generated by a MIDI synthesizer. Experimental results demonstrated that the overtone structures were stably and accurately estimated for a wide variety of initial settings. Daichi Sakaue, Katsutoshi Itoyama, Tetsuya Ogata, Hiroshi G. Okuno |
ICASSP | 3 |
| 2012 | Incremental probabilistic geometry estimation for robot scene understandingabstractOur goal is to give mobile robots a rich representation of their environment as fast as possible. Current mapping methods such as SLAM are often sparse, and scene reconstruction methods using tilting laser scanners are relatively slow. In this paper, we outline a new method for iterative construction of a geometric mesh using streaming time-of-flight range data. Our results show that our algorithm can produce a stable representation after 6 frames, with higher accuracy than raw time-of-flight data. Louis-Kenzo Cahier, Tetsuya Ogata, Hiroshi G. Okuno |
ICRA | 2 |
| 2012 | Rhythm-based adaptive localization in incomplete RFID landmark environmentsabstractThis paper proposes a novel hybrid-structured model for the adaptive localization of robots combining a stochastic localization model and a rhythmic action model, for avoiding vacant spaces of landmarks efficiently. In regularly arranged landmark environments, robots may not be able to detect any landmarks for a long time during a straight-like movement. Consequently, locally diverse and smooth movement patterns need to be generated to keep the position estimation stable. Conventional approaches aiming at the probabilistic optimization cannot rapidly generate the detailed movement pattern due to a huge computational cost; therefore a simple but diverse movement structure needs to be introduced as an alternative option. We solve this problem by combining a particle filter as the stochastic localization module and the dynamical action model generating a zig-zagging motion. The validation experiments, where virtual-line-tracing tasks are exhibited on a floor-installed RFID environment, show that introducing the proposed rhythm pattern can improve a minimum error boundary and a velocity performance for arbitrary tolerance errors can be improved by the rhythm amplitude adaptation fed back by the localization deviation. Kenri Kodaka, Tetsuya Ogata, Shigeki Sugano |
ICRA | 2 |
| 2012 | Automatic Chord Recognition Based on Probabilistic Integration of Acoustic Features, Bass Sounds, and Chord Transition
Katsutoshi Itoyama, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE | 2 |
| 2012 | Self-organization of object features representing motion using Multiple Timescales Recurrent Neural NetworkabstractAffordance theory suggests that humans recognize the environment based on invariants. Invariants are features that describe the environment offering behavioral information to humans. Two types of invariants exist, structural invariants and transformational invariants. In our previous paper, we developed a method that self-organizes transformational invariants, or motion features, from camera images based on robot's experiences. The model used a bi-directional technique combining a recurrent neural network for dynamics learning, namely Recurrent Neural Network with Parametric Bias (RNNPB), and a hierarchical neural network for feature extraction. The bi-directional training method developed in the previous work was effective in clustering the motion of objects, but the analysis did not give good segregation results of the self-organized features (transformational invariants) among different motion types. In this paper, we present a refined model which integrates dynamics learning and feature extraction in a single model. The refined model is comprised of Multiple Timescales Recurrent Neural Network (MTRNN), which possesses better learning capability than RNNPB. Self-organization result of four types of motions have proved the model's capability to create clusters of object motions. The analysis showed that the model extracted feature sequences with different characteristics for four object motion types. Shun Nishide, Jun Tani, Hiroshi G. Okuno, Tetsuya Ogata |
IJCNN | 4 |
| 2012 | Body area segmentation from visual scene based on predictability of neuro-dynamical systemabstractWe propose neural models for segmenting the area of a body from visual scene based on predictability. Neuroscience has shown that a prediction model in brain, which predicts sensory-feedback from motor command, can divide the sensory-feedback into the self-motion derived feedback and other derived feedback. The prediction model is important for prediction control of the body. Previous studies in robotics of the prediction model assumed that a robot can recognize the position of its body (e.g. its hand) and that the view contains only that body part. In our models, motor commands and visual feedback (pixel image that includes not only a hand but also object and background) are input into a neural network model and then the body area is segmented and prediction model of body is acquired. Our model contains two parts: 1) An object detection model obtains a conversion system between object positions and the pixel image. 2) A movement prediction model predicts hand-object positions from motor commands and identifies the body. We confirmed that our models can segment the body/object area based on their pixel textures and discriminate between them by using prediction error. Harumitsu Nobuta, Kenta Kawamoto, Kuniaki Noda, Kohtaro Sabe, Shun Nishide, Hiroshi G. Okuno, Tetsuya Ogata |
IJCNN | 7 |
| 2012 | Who is the leader in a multiperson ensemble? - Multiperson human-robot ensemble model with leaderness -abstractThis paper presents a state space model for a multiperson ensemble and an estimation method of the onset timings, tempos, and leaders. In a multiperson ensemble, determining one explicit leader is difficult because (1) participants' rhythms are mutually influenced and (2) they compete with each other. Most ensemble studies however assumed that one leader exists at a time and the others just follow the leader. To deal with the multiple and time-varying leaders, we define leaderness indicating the power to influence the others as the product of the tempo stability and the distance from the ensemble tempo. This definition means that a leader should have a strong desire to change the current tempo. Using the leaderness, we present a state space model of a multiperson ensemble and an unscented Kalman filter based estimation method. The model consists of the leaderness update, the ensemble tempo update, the individual tempo update, and the onset timing adaptation, each of which has a relationship to psychological results of an ensemble. We evaluate our method using simulation and human behavior. The simulation results show that our model is stable for various initial tempos and the number of participants. For the human behavior, pairs and triads of participants are asked to tap keys in synchronization with the others. The results show that the leaderness successfully indicate the dynamics of the leaders, and the onset errors are 181msec and 241msec for pairs and triads on average, respectively, which are comparable to those of humans (153msec and 227msec for pairs and triads, respectively.) Takeshi Mizumoto, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 2 |
| 2012 | Sound sources selection system by using onomatopoeic querries from multiple sound sourcesabstractOur motivation is to develop a robot that treats auditory information in real environment because auditory information is useful for animated communications or understanding our surroundings. Interactions by using sound information need an aquisition of it and a proper sound source reference between a user and a robot leads to it. Such sound source reference is difficult due to multiple sound sources generating in real environemnt, and we use onomatopoeic representations as a representation for the reference. This paper shows a system that selects a sound source specified by a user from multiple sound sources. Users use onomatopoeias in the specification, and our system separates a mixed sound and converts separated sounds into onomatopoeias for the selection. Onomatopoeais have the ambiguity that each user gives each expression to a certain sound and we create an original similarity based on Minimum Edit Distance and acoustic features for solving its problem. In experiments, our system receives a mixed sound consisting of three sounds and a user's query as inputs, and checks a count of a consistency of a sound source selected by a system and a sound source specified by a user in 100 tests. The result shows. Yusuke Yamamura, Toru Takahashi 0001, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 3 |
| 2012 | Efficient Blind Dereverberation and Echo Cancellation Based on Independent Component Analysis for Actual Acoustic SignalsabstractThis letter presents a new algorithm for blind dereverberation and echo cancellation based on independent component analysis (ICA) for actual acoustic signals. We focus on frequency domain ICA (FD-ICA) because its computational cost and speed of learning convergence are sufficiently reasonable for practical applications such as hands-free speech recognition. In applying conventional FD-ICA as a preprocessing of automatic speech recognition in noisy environments, one of the most critical problems is how to cope with reverberations. To extract a clean signal from the reverberant observation, we model the separation process in the short-time Fourier transform domain and apply the multiple input/output inverse-filtering theorem (MINT) to the FD-ICA separation model. A naive implementation of this method is computationally expensive, because its time complexity is the second order of reverberation time. Therefore, the main issue in dereverberation is to reduce the high computational cost of ICA. In this letter, we reduce the computational complexity to the linear order of the reverberation time by using two techniques: (1) a separation model based on the independence of delayed observed signals with MINT and (2) spatial sphering for preprocessing. Experiments show that the computational cost grows in proportion to the linear order of the reverberation time and that our method improves the word correctness of automatic speech recognition by 10 to 20 points in a RT₂₀= 670 ms reverberant environment. Ryu Takeda, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
Neural Comput. | 5 |
| 2011 | Cluster Self-organization of Known and Unknown Environmental Sounds Using Recurrent Neural Network
Shun Nishide, Toru Takahashi 0001, Hiroshi G. Okuno, Tetsuya Ogata |
ICANN (1) | 5 |
| 2011 | Simultaneous processing of sound source separation and musical instrument identification using Bayesian spectral modelingabstractThis paper presents a method of both separating audio mixtures into sound sources and identifying the musical instruments of the sources. A statistical tone model of the power spectrogram, called an integrated model, is defined and source separation and instrument identification are carried out on the basis of Bayesian inference. Since, the parameter distributions of the integrated model depend on each instrument, the instrument name is identified by selecting the one that has the maximum relative instrument weight. Experimental results showed correct instrument identification enables precise source separation even when many overtones overlap. Katsutoshi Itoyama, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ICASSP | 4 |
| 2011 | Polyphonic audio-to-score alignment based on Bayesian Latent Harmonic Allocation Hidden Markov ModelabstractThis paper presents a Bayesian method for temporally aligning a music score and an audio rendition. A critical problem in audio-to-score alignment is in dealing with the wide variety of timbre and volume of the audio rendition. In contrast with existing works that achieve this through ad-hoc feature design or careful training of tone models, we propose a Bayesian audio-to-score alignment method by modeling music performance as a Bayesian Hidden Markov Model, each state of which emits a Bayesian signal model based on Latent Harmonic Allocation. After attenuating reverberation, variational Bayes method is used to iteratively adapt the alignment, instrument tone model and the volume balance at each position of the score. The method is evaluated using sixty works of classical music of a variety of instrumentation ranging from solo piano to full orchestra. We verify that our method improves the alignment accuracy compared to dynamic time warping based on chroma vector for orchestral music, or our method employed in a maximum likelihood setting. Akira Maezawa, Hiroshi G. Okuno, Tetsuya Ogata, Masataka Goto |
ICASSP | 3 |
| 2011 | Use of a Sparse Structure to Improve Learning Performance of Recurrent Neural Networks
Hiromitsu Awano, Shun Nishide, Hiroaki Arie, Jun Tani, Toru Takahashi 0001, Hiroshi G. Okuno, Tetsuya Ogata |
ICONIP (3) | 7 |
| 2011 | Robot with Two Ears Listens to More than Two Simultaneous Utterances by Exploiting Harmonic Structures
Yasuharu Hirasawa, Toru Takahashi 0001, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE (1) | 3 |
| 2011 | Environmental Sound Recognition for Robot Audition Using Matching-Pursuit
Nobuhide Yamakawa, Toru Takahashi 0001, Tetsuro Kitahara, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE (2) | 4 |
| 2011 | Fast and Simple Iterative Algorithm of Lp-Norm Minimization for Under-Determined Speech SeparationabstractThis paper presents an efficient algorithm to solve Lp-norm minimization problem for under-determined speech separation; that is, for the case that there are more sound sources than microphones. We employ an auxiliary function method in order to derive update rules under the assumption that the amplitude of each sound source follows generalized Gaussian distribution. Experiments reveal that our method solves the L1-norm minimization problem ten times faster than a general solver, and also solves Lp-norm minimization problem efficiently, especially when the parameter p is small; when p is not more than 0.7, it runs in real-time without loss of separation quality. Index Terms: speech separation, under-determined condition, Lp-norm minimization, auxiliary function method Yasuharu Hirasawa, Naoki Yasuraoka, Toru Takahashi 0001, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 4 |
| 2011 | Bayesian Extension of MUSIC for Sound Source Localization and TrackingabstractThis paper presents a Bayesian extension of MUSIC-based sound source localization (SSL) and tracking method. SSL is important for distant speech enhancement and simultaneous speech separation for improving speech recognition, as well as for auditory scene analysis by mobile robots. One of the drawbacks of existing SSL methods is the necessity of careful parameter tunings, e.g., the sound source detection threshold depending on the reverberation time and the number of sources. Our contribution consists of (1) automatic parameter estimation in the variational Bayesian framework and (2) tracking of sound sources with reliability. Experimental results demonstrate our method robustly tracks multiple sound sources in a reverberant environment with RT20 = 840 (ms). Index Terms: simultaneous sound source localization, MUSIC algorithm, variational Bayes, particle filter Takuma Otsuka, Kazuhiro Nakadai, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 3 |
| 2011 | Particle-filter based audio-visual beat-tracking for music robot ensemble with human guitaristabstractThis paper presents an audio-visual beat-tracking method for ensemble robots with a human guitarist. Beat-tracking, or estimation of tempo and beat times of music, is critical to the high quality of musical ensemble performance. Since a human plays the guitar in out-beat in back beat and syncopation, the main problems of beat-tracking of a human's guitar playing are twofold: tempo changes and varying note lengths. Most conventional methods have not addressed human's guitar playing. Therefore, they lack the adaptation of either of the problems. To solve the problems simultaneously, our method uses not only audio but visual features. We extract audio features with Spectro-Temporal Pattern Matching (STPM) and visual features with optical flow, mean shift and Hough transform. Our beat-tracking estimates tempo and beat time using a particle filter; both acoustic feature of guitar sounds and visual features of arm motions are represented as particles. The particle is determined based on prior distribution of audio and visual features, respectively Experimental results confirm that our integrated audio-visual approach is robust against tempo changes and varying note lengths. In addition, they also show that estimation convergence rate depends only a little on the number of particles. The real-time factor is 0.88 when the number of particles is 200, and this shows out method works in real-time. Tatsuhiko Itohara, Takuma Otsuka, Takeshi Mizumoto, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 4 |
| 2011 | Improvement of speaker localization by considering multipath interference of sound wave for binaural robot auditionabstractThis paper presents an improved speaker localization method based on the generalized cross-correlation (GCC) method weighted by the phase transform (PHAT) for binaural robot audition. The problem with the conventional direction-of-arrival (DOA) estimation based on the GCC-PHAT method is a multipath interference whereby a sound wave travels to microphones via the front-head path and the back-head path in binaural robot audition. This paper describes a new time delay factor for the GCC-PHAT method to compensate multipath interference on the assumption of spherical robot head. In addition, the restriction of the time difference of arrival (TDOA) estimation by the sampling frequency is also solved by applying the maximum likelihood (ML) estimation in frequency domain. Experiments conducted in the SIG-2 humanoid robot show that the proposed method reduces localization errors by 17.8 degrees on average and by over 35 degrees in side directions comparing to the conventional DOA estimation. Ui-Hyun Kim, Takeshi Mizumoto, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 3 |
| 2011 | Handwriting prediction based character recognition using recurrent neural networkabstractHumans are said to unintentionally trace handwriting sequences in their brains based on handwriting experiences when recognizing written text. In this paper, we propose a model for predicting handwriting sequence for written text recognition based on handwriting experiences. The model is first trained using image sequences acquired while writing text. The image features of sequences are self-organized from the images using Self-Organizing Map. The feature sequences are used to train a neuro-dynamics learning model. For recognition, the text image is input into the model for predicting the handwriting sequence and recognition of the text. We conducted two experiments using ten Japanese characters. The results of the experiments show the effectivity of the model. Shun Nishide, Hiroshi G. Okuno, Tetsuya Ogata, Jun Tani |
SMC | 3 |
| 2011 | Emergence of hierarchical structure mirroring linguistic composition in a recurrent neural network
Wataru Hinoshita, Hiroaki Arie, Jun Tani, Hiroshi G. Okuno, Tetsuya Ogata |
Neural Networks | 5 |
| 2010 | Design and Implementation of Two-level Synchronization for Interactive Music RobotabstractOur goal is to develop an interactive music robot, i.e., a robot that presents a musical expression together with humans. A music interaction requires two important functions: synchronization with the music and musical expression, such as singing and dancing. Many instrument-performing robots are only capable of the latter function, they may have difficulty in playing live with human performers. The synchronization function is critical for the interaction. We classify synchronization and musical expression into two levels: (1) the rhythm level and (2) the melody level. Two issues in achieving two-layer synchronization and musical expression are: (1) simultaneous estimation of the rhythm structure and the current part of the music and (2) derivation of the estimation confidence to switch behavior between the rhythm level and the melody level. This paper presents a score following algorithm, incremental audio to score alignment, that conforms to the two-level synchronization design using a particle filter. Our method estimates the score position for the melody level and the tempo for the rhythm level. The reliability of the score position estimation is extracted from the probability distribution of the score position. Experiments are carried out using polyphonic jazz songs. The results confirm that our method switches levels in accordance with the difficulty of the score estimation. When the tempo of the music is less than 120 (beats per minute; bpm), the estimated score positions are accurate and reported; when the tempo is over 120 (bpm), the system tends to report only the tempo to suppress the error in the reported score position predictions. Takuma Otsuka, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
AAAI | 5 |
| 2010 | Improvement in listening capability for humanoid robot HRP-2abstractThis paper describes improvement of sound source separation for a simultaneous automatic speech recognition (ASR) system of a humanoid robot. A recognition error in the system is caused by a separation error and interferences of other sources. In separability, an original geometric source separation (GSS) is improved. Our GSS uses a measured robot's head related transfer function (HRTF) to estimate a separation matrix. As an original GSS uses a simulated HRTF calculated based on a distance between microphone and sound source, there is a large mismatch between the simulated and the measured transfer functions. The mismatch causes a severe degradation of recognition performance. Faster convergence speed of separation matrix reduces separation error. Our approach gives a nearer initial separation matrix based on a measured transfer function from an optimal separation matrix than a simulated one. As a result, we expect that our GSS improves the convergence speed. Our GSS is also able to handle an adaptive step-size parameter. These new features are added into open source robot audition software (OSS) called "HARK" which is newly updated as version 1.0.0. The HARK has been installed on a HRP-2 humanoid with an 8-element microphone array. The listening capability of HRP-2 is evaluated by recognizing a target speech signal which is separated from a simultaneous speech signal by three talkers. The word correct rate (WCR) of ASR improves by 5 points under normal acoustic environments and by 10 points under noisy environments. Experimental results show that HARK 1.0.0 improves the robustness against noises. Toru Takahashi 0001, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ICRA | 4 |
| 2010 | Upper-limit evaluation of robot audition based on ICA-BSS in multi-source, barge-in and highly reverberant conditionsabstractThis paper presents the upper-limit evaluation of robot audition based on ICA-BSS in multi-source, barge-in and highly reverberant conditions. The goal is that the robot can automatically distinguish a target speech from its own speech and other sound sources in a reverberant environment. We focus on the multi-channel semi-blind ICA (MCSB-ICA), which is one of the sound source separation methods with a microphone array, to achieve such an audition system because it can separate sound source signals including reverberations with few assumptions on environments. The evaluation of MCSB-ICA has been limited to robot's speech separation and reverberation separation. In this paper, we evaluate MCSB-ICA extensively by applying it to multi-source separation problems under common reverberant environments. Experimental results prove that MCSB-ICA outperforms conventional ICA by 30 points in automatic speech recognition performance. Ryu Takeda, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ICRA | 5 |
| 2010 | Recognition and Generation of Sentences through Self-organizing Linguistic Hierarchy Using MTRNN
Wataru Hinoshita, Hiroaki Arie, Jun Tani, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE (3) | 4 |
| 2010 | Violin Fingering Estimation Based on Violin Pedagogical Fingering Model Constrained by Bowed Sequence Estimation from Audio Input
Akira Maezawa, Katsutoshi Itoyama, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE (3) | 5 |
| 2010 | Improving Identification Accuracy by Extending Acceptable Utterances in Spoken Dialogue System Using Barge-in Timing
Kyoko Matsuyama, Kazunori Komatani, Toru Takahashi 0001, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE (2) | 4 |
| 2010 | Music-Ensemble Robot That Is Capable of Playing the Theremin While Listening to the Accompanied Music
Takuma Otsuka, Takeshi Mizumoto, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE (1) | 6 |
| 2010 | Analyzing user utterances in barge-in-able spoken dialogue system for improving identification accuracyabstractIn our barge-in-able spoken dialogue system, the user’s behaviors such as barge-in timing and utterance expressions vary according to his/her characteristics and situations. The system adapts to the behaviors by modeling them. We analyzed 1584 utterances collected by our systems of quiz and news-listing tasks and showed that ratio of using referential expressions depends on individual users and average lengths of listed items. This tendency was incorporated as a prior probability into our method and improved the identification accuracy of the user’s intended items. Index Terms: barge-in, spoken dialogue systems, utterance timing, user characteristics Kyoko Matsuyama, Kazunori Komatani, Ryu Takeda, Toru Takahashi 0001, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 5 |
| 2010 | Effects of modelling within- and between-frame temporal variations in power spectra on non-verbal sound recognitionabstractResearch on environmental sound recognition has not shown great development in comparison with that on speech and musical signals. One of the reasons is that the sound category of environmental sounds covers a broad range of acoustical natures. We classified them in order to explore suitable recognition techniques for each characteristic. We focus on impulsive sounds and their non-stationary feature within and between analytic frames. We used matching-pursuit as a framework to use wavelet analysis for extracting temporal variation of audio features inside a frame. We also investigated the validity of modeling decaying patterns of sounds using Hidden markov models. Experimental results indicate that sounds with multiple impulsive signals are recognized better by using time-frequency analyzing bases than by frequency domain analysis. Classification of sound classes with a long and clear decaying pattern improves when HMMs with multiple number of hidden states are applied. Index Terms: audio signal classification, non-speech sound recognition, environmental sound recognition, time-frequency analysis, Matching-Pursuit Nobuhide Yamakawa, Tetsuro Kitahara, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 5 |
| 2010 | Exploiting harmonic structures to improve separating simultaneous speech in under-determined conditionsabstractIn real-world situations, a robot may often encounter “under-determined” situation, where there are more sound sources than microphones. This paper presents a speech separation method using a new constraint on the harmonic structure for a simultaneous speech-recognition system in under-determined conditions. The requirements for a speech separation method in a simultaneous speech-recognition system are (1) ability to handle a large number of talkers, and (2) reduction of distortion in acoustic features. Conventional methods use a maximum likelihood estimation in sound source separation, which fulfills requirement (1). Since it is a general approach, the performance is limited when separating speech. This paper presents a two-stage method to improve the separation. The first stage uses maximum likelihood estimation and extracts the harmonic structure, and the second stage exploits the harmonic structure as a new constraint to achieve requirement (2). We carried out an experiment that simulated three simultaneous utterances using impulse responses recorded by two microphones in an anechoic chamber. The experimental results revealed that our method could improve speech recognition correctness by about four points. Yasuharu Hirasawa, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 4 |
| 2010 | Robot musical accompaniment: integrating audio and visual cues for real-time synchronization with a human flutistabstractMusicians often have the following problem: they have a music score that requires 2 or more players, but they have no one with whom to practice. So far, score-playing music robots exist, but they lack adaptive abilities to synchronize with fellow players' tempo variations. In other words, if the human speeds up their play, the robot should also increase its speed. However, computer accompaniment systems allow exactly this kind of adaptive ability. We present a first step towards giving these accompaniment abilities to a music robot. We introduce a new paradigm of beat tracking using 2 types of sensory input - visual and audio - using our own visual cue recognition system and state-of-the-art acoustic onset detection techniques. Preliminary experiments suggest that by coupling these two modalities, a robot accompanist can start and stop a performance in synchrony with a flutist, and detect tempo changes within half a second. Angelica Lim, Takeshi Mizumoto, Louis-Kenzo Cahier, Takuma Otsuka, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 7 |
| 2010 | Human-robot ensemble between robot thereminist and human percussionist using coupled oscillator modelabstractThis paper presents a novel synchronizing method for a human-robot ensemble using coupled oscillators. We define an ensemble as a synchronized performance produced through interactions between independent players. To attain better synchronized performance, the robot should predict the human's behavior to reduce the difference between the human's and robot's onset timings. Existing studies in such synchronization only adapts to onset intervals, thus, need a considerable time to synchronize. We use a coupled oscillator model to predict the human's behavior. Experimental results show that our method reduces the average of onset time errors; when we use a metronome, a tempo-varying metronome or a human drummer, errors are reduced by 38%, 10% or 14% on the average, respectively. These results mean that the prediction of human's behaviors is effective for the synchronized performance. Takeshi Mizumoto, Takuma Otsuka, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 6 |
| 2010 | Motion generation based on reliable predictability using self-organized object featuresabstractPredictability is an important factor for determining robot motions. This paper presents a model to generate robot motions based on reliable predictability evaluated through a dynamics learning model which self-organizes object features. The model is composed of a dynamics learning module, namely Recurrent Neural Network with Parametric Bias (RNNPB), and a hierarchical neural network as a feature extraction module. The model inputs raw object images and robot motions. Through bi-directional training of the two models, object features which describe the object motion are self-organized in the output of the hierarchical neural network, which is linked to the input of RNNPB. After training, the model searches for the robot motion with high reliable predictability of object motion. Experiments were performed with the robot's pushing motion with a variety of objects to generate sliding, falling over, bouncing, and rolling motions. For objects with single motion possibility, the robot tended to generate motions that induce the object motion. For objects with two motion possibilities, the robot evenly generated motions that induce the two object motions. Shun Nishide, Tetsuya Ogata, Jun Tani, Toru Takahashi 0001, Kazunori Komatani, Hiroshi G. Okuno |
IROS | 2 |
| 2010 | An improvement in automatic speech recognition using soft missing feature masks for robot auditionabstractWe describe integration of preprocessing and automatic speech recognition based on Missing-Feature-Theory (MFT) to recognize a highly interfered speech signal, such as the signal in a narrow angle between a desired and interfered speakers. As a speech signal separated from a mixture of speech signals includes the leakage from other speech signals, recognition performance of the separated speech degrades. An important problem is estimating the leakage in time-frequency components. Once the leakage is estimated, we can generate missing feature masks (MFM) automatically by using our method. A new weighted sigmoid function is introduced for our MFM generation method. An experiment shows that a word correct rate improves from 66 % to 74 % by using our MFM generation method tuned by a search base approach in the parameter space. Toru Takahashi 0001, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 4 |
| 2010 | Speedup and performance improvement of ICA-based robot audition by parallel and resampling-based block-wise processingabstractThis paper describes a speedup and performance improvement of multi-channel semi-blind ICA (MCSB-ICA) with parallel and resampling-based block-wise processing. MCSB-ICA is an integrated method of sound source separation that accomplishes blind source separation, blind dereverberation, and echo cancellation. This method enables robots to separate user's speech signals from observed signals including the robot's own speech, other speech and their reverberations without a priori information. The main problem when MCSB-ICA is applied to robot audition is its high computational cost. We tackle this by multi-threading programming, and the two main issues are 1) the design of parallel processing and 2) incremental implementation. These are solved by a) multiple-stack-based parallel implementation, and b) resampling-based overlaps and block-wise separation. The experimental results proved that our method reduced the real-time factor to less than 0.5 with an eight-core CPU, and it improves the performance of automatic speech recognition by 2-10 points compared with the single-stack-based parallel implementation without the resampling technique. Ryu Takeda, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 5 |
| 2010 | Human-robot cooperation in arrangement of objects using confidence measure of neuro-dynamical systemabstractThe objective of our study was to develop dynamic collaboration between a human and a robot. Most conventional studies have created pre-designed rule-based collaboration systems to determine the timing and behavior of robots to participate in tasks. Our aim is to introduce the confidence of the task as a criterion for robots to determine their timing and behavior. In this paper, we report the effectiveness of applying reproduction accuracy as a measure for quantitatively evaluating confidence in an object arrangement task. Our method is comprised of three phases. First, we obtain human-robot interaction data through the Wizard of OZ method. Second, the obtained data are trained using a neuro-dynamical system, namely, the Multiple Time-scales Recurrent Neural Network (MTRNN). Finally, the prediction error in MTRNN is applied as a confidence measure to determine the robot's behavior. The robot participated in the task when its confidence was high, while it just observed when its confidence was low. Training data were acquired using an actual robot platform, Hiro. The method was evaluated using a robot simulator. The results revealed that motion trajectories could be precisely reproduced with a high degree of confidence, demonstrating the effectiveness of the method. Hiromitsu Awano, Tetsuya Ogata, Shun Nishide, Toru Takahashi 0001, Kazunori Komatani, Hiroshi G. Okuno |
SMC | 2 |
| 2010 | Inter-modality mapping in robot with recurrent neural network
Tetsuya Ogata, Shun Nishide, Hideki Kozima, Kazunori Komatani, Hiroshi G. Okuno |
Pattern Recognit. Lett. | 1 |
| 2009 | ICA-based efficient blind dereverberation and echo cancellation method for barge-in-able robot auditionabstractThis paper describes a new method that allows ldquoBarge-Inrdquo in various environments for robot audition. ldquoBarge-inrdquo means that a user begins to speak simultaneously while a robot is speaking. To achieve the function, we must deal with problems on blind dereverberation and echo cancellation at the same time. We adopt Independent Component Analysis (ICA) because it essentially provides a natural framework for these two problems. To deal with reverberation, we apply a Multiple Input/Output INverse-filtering Theorem-based model of observation to the frequency domain ICA. The main problem is its high-computational cost of ICA. We reduce the computational complexity to the linear order of reverberation time by using two techniques: 1) a separation modelbased on observed signal independence, and 2) enforced spatial sphering for preprocessing. The experimental results revealed that our method improved word correctness of reverberant speech by 10-20 points. Ryu Takeda, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ICASSP | 5 |
| 2009 | Continuous vocal imitation with self-organized vowel spaces in Recurrent Neural NetworkabstractA continuous vocal imitation system was developed using a computational model that explains the process of phoneme acquisition by infants. Human infants perceive speech sounds not as discrete phoneme sequences but as continuous acoustic signals. One of critical problems in phoneme acquisition is the design for segmenting these continuous speech sounds. The key idea to solve this problem is that articulatory mechanisms such as the vocal tract help human beings to perceive speech sound units corresponding to phonemes. To segment acoustic signal with articulatory movement, we apply the segmenting method to our system by Recurrent Neural Network with Parametric Bias (RNNPB). This method determines the multiple segmentation boundaries in a temporal sequence using the prediction error of the RNNPB model, and the PB values obtained by the method can be encoded as kind of phonemes. Our system was implemented by using a physical vocal tract model, called the Maeda model. Experimental results demonstrated that our system can self-organize the same phonemes in different continuous sounds, and can imitate vocal sound involving arbitrary numbers of vowels using the vowel space in the RNNPB. This suggests that our model reflects the process of phoneme acquisition. Hisashi Kanda, Tetsuya Ogata, Toru Takahashi 0001, Kazunori Komatani, Hiroshi G. Okuno |
ICRA | 2 |
| 2009 | Prediction and imitation of other's motions by reusing own forward-inverse model in robotsabstractThis paper proposes a model that enables a robot to predict and imitate the motions of another by reusing its body forward-inverse model. Our model includes three approaches: (i) projection of a self-forward model for predicting phenomena in the external environment (other individuals), (ii) ldquotriadic relationrdquo that is mediation by a physical object between self and others, (iii) introduction of infant imitation by a parent. The recurrent neural network with parametric bias (RNNPB) model is used as the robot's self forward-inverse model. A group of hierarchical neural networks are attached to the RNNPB model as ldquoconversion modulesrdquo. Experiments demonstrated that a robot with our model could imitate a human's motions by translating the viewpoint. It could also discriminate known/unknown motions appropriately, and associate whole motion dynamics from only one motion snap image. Tetsuya Ogata, Ryunosuke Yokoya, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno |
ICRA | 1 |
| 2009 | Adjusting Occurrence Probabilities of Automatically-Generated Abbreviated Words in Spoken Dialogue Systems
Masaki Katsumaru, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE | 3 |
| 2009 | Improving speech understanding accuracy with limited training data using multiple language models and multiple understanding modelsabstractWe aim to improve a speech understanding module with a small amount of training data. A speech understanding module uses a language model (LM) and a language understanding model (LUM). A lot of training data are needed to improve the models. Such data collection is, however, difficult in an actual process of development. We therefore design and develop a new framework that uses multiple LMs and LUMs to improve speech understanding accuracy under various amounts of training data. Even if the amount of available training data is small, each LM and each LUM can deal well with different types of utterances and more utterances are understood by using multiple LM and LUM. As one implementation of the framework, we develop a method for selecting the most appropriate speech understanding result from several candidates. The selection is based on probabilities of correctness calculated by logistic regressions. We evaluate our framework with various amounts of training data. Index Terms: speech understanding, multiple language models and language understanding models, limited training data Masaki Katsumaru, Mikio Nakano, Kazunori Komatani, Kotaro Funakoshi, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 5 |
| 2009 | Enabling a user to specify an item at any time during system enumeration - item identification for barge-in-able conversational dialogue systemsabstractIn conversational dialogue systems, users prefer to speak at any time and to use natural expressions. We have developed an Independent Component Analysis (ICA) based semi-blind source separation method, which allows users to barge-in over system utterances at any time. We created a novel method from timing information derived from barge-in utterances to identify one item that a user indicates during system enumeration. First, we determine the timing distribution of user utterances containing referential expressions and then approximate it using a gamma distribution. Second, we represent both the utterance timing and automatic speech recognition (ASR) results as probabilities of the desired selection from the system’s enumeration. We then integrate these two probabilities to identify the item having the maximum likelihood of selection. Experimental results using 400 utterances indicated that our method outperformed two methods used as a baseline (one of ASR results only and one of utterance timing only) in identification accuracy. Index Terms: spoken dialogue system, conversational interaction, barge-in, utterance timing Kyoko Matsuyama, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 3 |
| 2009 | Emergence of evolutionary interaction with voice and motion between two robots using RNNabstractWe propose a model of evolutionary interaction between two robots where signs used for communication emerge through mutual adaptation. Signs used in human interaction, e.g., language, gestures and eye contact change and evolve in form and meaning through repeated use. To create flexible human-like interaction systems, it is necessary to deal with signs as a dynamic property and to construct a framework in which signs emerge from mutual adaptation by agents. Our target is multi-modal interaction using voice and motion between two robots where a voice/motion pattern is used as a sign referring to a motion/voice pattern. To enable evolutionary signs (voice and motion patterns) to be recognized and generated, we utilized a dynamics model: Multiple Timescale Recurrent Neural Network (MTRNN). To enable the robots to interpret signs, we utilized hierarchical neural networks, which transform dynamics model parameters of voice/motion into those of motion/voice. In our experiment, two robots modified their own interpretation of signs constantly through mutual adaptation in interaction where they responded to the other's voice with motion one after the other. As a result of the experiment, we found that the interaction kept evolving through the robots' repeated and alternate miscommunications and re-adaptations, and this induced the emergence of diverse new signs that depended on the robots' body dynamics through the generalization capability of MTRNN. Wataru Hinoshita, Tetsuya Ogata, Hideki Kozima, Hisashi Kanda, Toru Takahashi 0001, Hiroshi G. Okuno |
IROS | 2 |
| 2009 | Phoneme acquisition model based on vowel imitation using Recurrent Neural NetworkabstractA phoneme-acquisition system was developed using a computational model that explains the developmental process of human infants in the early period of acquiring language. There are two important findings in constructing an infant's acquisition of phonemes: (1) an infant's vowel like cooing tends to invoke utterances that are imitated by its caregiver, and (2) maternal imitation effectively reinforces infant vocalization. Therefore, we hypothesized that infants can acquire phonemes to imitate their caregivers' voices by trial and error, i. e., infants use self-vocalization experience to search for imitable and unimitable elements in their caregivers' voices. On the basis of this hypothesis, we constructed a phoneme acquisition process using interaction involving vowel imitation between a human and an infant model. Our infant model had a vocal tract system, called the Maeda model, and an auditory system implemented by using mel-frequency cepstral coefficients (MFCCs) through STRAIGHT analysis. We applied recurrent neural network with parametric bias (RNNPB) to learn the experience of self-vocalization, to recognize the human voice, and to produce the sound imitated by the infant model. To evaluate imitable and unimitable sounds, we used the prediction error of the RNNPB model. The experimental results revealed that as imitation interactions were repeated, the formants of sounds imitated by our system moved closer to those of human voices, and our system could self-organize the same vowels in different continuous sounds. This suggests that our system can reflect the process of phoneme acquisition. Hisashi Kanda, Tetsuya Ogata, Toru Takahashi 0001, Kazunori Komatani, Hiroshi G. Okuno |
IROS | 2 |
| 2009 | Thereminist robot: Development of a robot theremin player with feedforward and feedback arm control based on a Theremin's pitch modelabstractWe propose a Thereminist robot system that plays the Theremin based on a Theremin's pitch model. The theremin, which is a 1920s electronic musical instrument, is played by moving a player's hand position in the air without touching it. It is difficult to play the Theremin because the relationship between the hand position and Theremin's pitch (pitch characteristics) is non-linear and varies according to the electromagnetic field (hereafter called environment). These characteristics cause two problems: (1) adapting to the environment change is required and (2) a nai¿ve design tends to depend on robot's particular hardware. We implement the coarse-to-fine control system on the Thereminist robot using newly proposed two pitch models: parametric and nonparametric ones. The Thereminist robot works as below: first, the robot calibrates the pitch model by parameter fitting with the Levenberg-Marquardt method. Second, the robot moves its hand in a coarse manner by feedforward control based on the pitch model. Finally, the robot adjusts its position by feedback control (proportional-integral control). In these steps, the robot can play a required pitch quickly, because the robot moves its hand using the pitch model without listening to the Theremin's sound Thus, the time to play the exact pitch is shorter than when only feedback control is used. Three experiments were conducted to evaluate the robustness against the number of samples, environment change, and types of robots. The results revealed that our pitch model describes using only 12 samples of pitches for estimation of the parameters, and adapts if the environment changes. In addition, our system works on two different robots: HRP-2 and ASIMO. Takeshi Mizumoto, Hiroshi Tsujino, Toru Takahashi 0001, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 4 |
| 2009 | Modeling tool-body assimilation using second-order Recurrent Neural NetworkabstractTool-body assimilation is one of the intelligent human abilities. Through trial and experience, humans are capable of using tools as if they are part of their own bodies. This paper presents a method to apply a robot's active sensing experience for creating the tool-body assimilation model. The model is composed of a feature extraction module, dynamics learning module, and a tool recognition module. Self-Organizing Map (SOM) is used for the feature extraction module to extract object features from raw images. Multiple Time-scales Recurrent Neural Network (MTRNN) is used as the dynamics learning module. Parametric Bias (PB) nodes are attached to the weights of MTRNN as second-order network to modulate the behavior of MTRNN based on the tool. The generalization capability of neural networks provide the model the ability to deal with unknown tools. Experiments are performed with HRP-2 using no tool, I-shaped, T-shaped, and L-shaped tools. The distribution of PB values have shown that the model has learned that the robot's dynamic properties change when holding a tool. The results of the experiment show that the tool-body assimilation model is capable of applying to unknown objects to generate goal-oriented motions. Shun Nishide, Tatsuhiro Nakagawa, Tetsuya Ogata, Jun Tani, Toru Takahashi 0001, Hiroshi G. Okuno |
IROS | 3 |
| 2009 | Incremental polyphonic audio to score alignment using beat tracking for singer robotsabstractWe aim at developing a singer robot capable of listening to music with its own ¿ears¿ and interacting with a human's musical performance. Such a singer robot requires at least three functions: listening to the music, understanding what position in the music is being performed, and generating a singing voice. In this paper, we focus on the second function, that is, the capability to align an audio signal to its musical score represented symbolically. Issues underlying the score alignment problem are: (1) diversity in the sounds of various musical instruments, (2) difference between the audio signal and the musical score, (3) fluctuation in tempo of the musical performance. Our solutions to these issues are as follows: (1) the design of features based on a chroma vector in the 12-tone model and onset of the sound, (2) defining the rareness for each tone based on the idea that scarcely used tone is salient in the audio signal, and (3) the use of a switching Kalman filter for robust tempo estimation. The experimental result shows that our score alignment method improves the average of cumulative absolute errors in score alignment by 29% using 100 popular music tunes compared to the beat tracking without score alignment. Takuma Otsuka, Toru Takahashi 0001, Hiroshi G. Okuno, Kazunori Komatani, Tetsuya Ogata, Kazumasa Murata, Kazuhiro Nakadai |
IROS | 5 |
| 2009 | Missing-feature-theory-based robust simultaneous speech recognition system with non-clean speech acoustic modelabstractA humanoid robot must recognize a target speech signal while people around the robot chat with them in real-world. To recognize the target speech signal, robot has to separate the target speech signal among other speech signals and recognize the separated speech signal. As separated signal includes distortion, automatic speech recognition (ASR) performance degrades. To avoid the degradation, we trained an acoustic model from non-clean speech signals to adapt acoustic feature of distorted signal and adding white noise to separated speech signal before extracting acoustic feature. The issues are (1) To determine optimal noise level to add the training speech signals, and (2) To determine optimal noise level to add the separated signal. In this paper, we investigate how much noises should be added to clean speech data for training and how speech recognition performance improves for different positions of three talkers with soft masking. Experimental results show that the best performance is obtained by adding white noises of 30 dB. The ASR with the acoustic model outperforms with ASR with the clean acoustic model by 4 points. Toru Takahashi 0001, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 4 |
| 2009 | Step-size parameter adaptation of multi-channel semi-blind ICA with piecewise linear model for barge-in-able robot auditionabstractThis paper describes a step-size parameter adaptation technique of multi-channel semi-blind independent component analysis (MCSB-ICA) for a ¿barge-in-able¿ robot audition system. By ¿barge-in¿, we mean that the user can speak simultaneously when the robot is speaking.We focused on MCSB-ICA to achieve such an audition system because it can separate a user's and a robot's speech under reverberant environments. The problem with MCSB-ICA for robot audition is the slow speed of convergence in estimating a separation filter due to its step-size parameters. Many optimization methods cannot be adopted because their computational costs are proportional to the 2nd order of the reverberation time. Our method yields adaptive step-size parameters with MCSB-ICA at low computational costs. It is based on three techniques; (1) recursive expression of the separation process, (2) a piecewise linear model of the step-size of the separation filter, and (3) adaptive step-size parameters with a sub-ICA-filter. Experimental results show that our approach attains faster convergence speed and lower computational costs than those with a fixed step-size parameter. Ryu Takeda, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 5 |
| 2009 | Bowed String Sequence Estimation of a Violin Based on Adaptive Audio Signal Classification and Context-Dependent Error CorrectionabstractThe sequence of strings played on a bowed string instrument is essential to understanding of the fingering. Thus, its estimation is required for machine understanding of violin playing. Audio-based identification is the only viable way to realize this goal for existing music recordings. A naive implementation using audio classification alone, however, is inaccurate and is not robust against variations in string or instruments. We develop a bowed string sequence estimation method by combining audio-based bowed string classification and context-dependent error correction. The robustness against different setups of instruments improves by normalizing the F0-dependent features using the average feature of a recording. The performance of error correction is evaluated using an electric violin with two different brands of strings and an acoustic violin. By incorporating mean normalization, the recognition error of recognition accuracy due to changing the string alleviates by 8 points, and that due to change of instrument by 12 points. Error correction decreases the error due to change of string by 8 points and that due to different instrument by 9 points. Akira Maezawa, Katsutoshi Itoyama, Toru Takahashi 0001, Tetsuya Ogata, Hiroshi G. Okuno |
ISM | 4 |
| 2009 | Changing timbre and phrase in existing musical performances as you like: manipulations of single part using harmonic and inharmonic modelsabstractThis paper presents a new music manipulation method that can change the timbre and phrases of an existing instrumental performance in a polyphonic sound mixture. This method consists of three primitive functions: 1) extracting and analyzing of a single instrumental part from polyphonic music signals, 2) mixing the instrument timbre with another, and 3) rendering a new phrase expression for another given score. The resulting customized part is re-mixed with the remaining parts of the original performance to generate new polyphonic music signals. A single instrumental part is extracted by using an integrated tone model that consists of harmonic and inharmonic tone models with the aid of the score of the single instrumental part. The extraction incorporates a residual model for the single instrumental part in order to avoid crosstalk between instrumental parts. The extracted model parameters are classified into their averages and deviations. The former is treated as instrument timbre and is customized by mixing, while the latter is treated as phrase expression and is customized by rendering. We evaluated our method in three experiments. The first experiment focused on introduction of the residual model, and it showed that the model parameters are estimated more accurately by 35.0 points. The second focused on timbral customization, and it showed that our method is more robust by 42.9 points in spectral distance compared with a conventional sound analysis-synthesis method, STRAIGHT. The third focused on the acoustic fidelity of customizing performance, and it showed that rendering phrase expression according to the note sequence leads to more accurate performance by 9.2 points in spectral distance in comparison with a rendering method that ignores the note sequence. Naoki Yasuraoka, Takehiro Abe, Katsutoshi Itoyama, Toru Takahashi 0001, Tetsuya Ogata, Hiroshi G. Okuno |
ACM Multimedia | 5 |
| 2009 | Ranking Help Message Candidates Based on Robust Grammar Verification Results and Utterance History in Spoken Dialogue Systems
Kazunori Komatani, Satoshi Ikeda, Yuichiro Fukubayashi, Tetsuya Ogata, Hiroshi G. Okuno |
SIGDIAL Conference | 4 |
| 2008 | Two-channel-based voice activity detection for humanoid robots in noisy home environmentsabstractThe purpose of this research is to accurately classify the speech signals originating from the front even in noisy home environments. This ability can help robots to improve speech recognition and to spot keywords. We therefore developed a new voice activity detection (VAD) based on the complex spectrum circle centroid (CSCC) method. It can classify the speech signals that are received at the front of two microphones by comparing the spectral energy of observed signals with that of target signals estimated by CSCC. Also, it can work in real time without training filter coefficients beforehand even in noisy environments (SNR ≫ 0 dB) and can cope with speech noises generated by audio-visual equipments such as televisions and audio devices. Since the CSCC method requires the directions of the noise signals, we also developed a sound source localization system integrated with cross-power spectrum phase (CSP) analysis and an expectation-maximization (EM) algorithm. This system was demonstrated to enable a robot to cope with multiple sound sources using two microphones. Hyun-Don Kim, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ICRA | 3 |
| 2008 | Object dynamics prediction and motion generation based on reliable predictabilityabstractConsistency of object dynamics, which is related to reliable predictability, is an important factor for generating object manipulation motions. This paper proposes a technique to generate autonomous motions based on consistency of object dynamics. The technique resolves two issues: construction of an object dynamics prediction model and evaluation of consistency. The authors utilize Recurrent Neural Network with Parametric Bias to self-organize the dynamics, and link static images to the self-organized dynamics using a hierarchical neural network to deal with the first issue. For evaluation of consistency, the authors have set an evaluation function based on object dynamics relative to robot motor dynamics. Experiments have shown that the method is capable of predicting 90% of unknown object dynamics. Motion generation experiments have proved that the technique is capable of generating autonomous pushing motions that generate consistent rolling motions. Shun Nishide, Tetsuya Ogata, Ryunosuke Yokoya, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno |
ICRA | 2 |
| 2008 | Integrating Topic Estimation and Dialogue History for Domain Selection in Multi-domain Spoken Dialogue Systems
Satoshi Ikeda, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE | 3 |
| 2008 | Rapid Prototyping of Robust Language Understanding Modules for Spoken Dialogue Systems
Yuichiro Fukubayashi, Kazunori Komatani, Mikio Nakano, Kotaro Funakoshi, Hiroshi Tsujino, Tetsuya Ogata, Hiroshi G. Okuno |
IJCNLP | 6 |
| 2008 | Extensibility verification of robust domain selection against out-of-grammar utterances in multi-domain spoken dialogue systemabstractWe developed a robust domain selection method and verified its extensibility. An issue in domain selection is its robustness against out-of-grammar utterances. It is essential to generate correct system responses because such utterances often cause domain selection errors. We therefore integrated the topic estimation results and the dialogue history to construct a robust domain classifier. Another issue is that domain selection should be performed within an extensible framework, because the system is often modified and extended. That is, the classifier should still have high performance without reconstructing it after adding new domains. The extensibility of our method was not experimentally verified yet, because it requires a lot of effort to collect new dialogue data after extending the system. Therefore, we verified extensibility without collecting new data. We constructed the classifier by leaving out some domains in the dialogue data and then evaluated its accuracy as the classifier for the data where the left-out domains were virtually added. Index Terms: multi-domain spoken dialogue system, domain selection, out-of-grammar utterance Satoshi Ikeda, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 3 |
| 2008 | Expanding vocabulary for recognizing user's abbreviations of proper nouns without increasing ASR error rates in spoken dialogue systemsabstractUsers often abbreviate long words when using spoken dialogue systems, which results in automatic speech recognition (ASR) errors. We define abbreviated words as sub-words of the original word, and add them into an ASR dictionary. The first problem is that proper nouns cannot be correctly segmented by general morphological analyzers, although long and compounded words need to be segmented in agglutinative languages such as Japanese. The second is that, as vocabulary increases, adding many abbreviated words degrades the ASR accuracy. We develop two methods, (1) to segment words by using conjunction probabilities between characters, and (2) to manipulate occurrence probabilities of generated abbreviated words on the basis of the phonological similarities between abbreviated and original words. By our method, the ASR accuracy is improved by 24.2 points for utterances containing abbreviated words, and degraded by only a 0.1 point for those containing original words. Index Terms: spoken dialogue systems, abbreviated words, proper nouns, vocabulary expansion Masaki Katsumaru, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 3 |
| 2008 | Soft missing-feature mask generation for simultaneous speech recognition system in robots
Toru Takahashi 0001, Shun'ichi Yamamoto, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 5 |
| 2008 | Segmenting acoustic signal with articulatory movement using Recurrent Neural Network for phoneme acquisitionabstractThis paper proposes a computational model for phoneme acquisition by infants. Human infants perceive speech sounds not as discrete phoneme sequences but as continuous acoustic signals. One of critical problems in phoneme acquisition is the design for segmenting these continuous speech sounds. The key idea to solve this problem is that articulatory mechanisms such as the vocal tract help human beings to perceive speech sound units corresponding to phonemes. That is, the ability to distinguish phonemes is learned by recognizing unstable points in the dynamics of continuous sound with articulatory movement. We have developed a vocal imitation system embodying the relationship between articulatory movements and sounds produced by the movements. To segment acoustic signal with articulatory movement, we apply the segmenting method to our system by recurrent neural network with parametric bias (RNNPB). This method determines the multiple segmentation boundaries in a temporal sequence using the prediction error of the RNNPB model, and the PB values obtained by the method can be encoded as kind of phonemes. Our system was implemented by using a physical vocal tract model, called the Maeda model. Experimental results demonstrated that our system can self-organize the same phonemes in different continuous sounds. This suggests that our model reflects the process of phoneme acquisition. Hisashi Kanda, Tetsuya Ogata, Kazunori Komatani, Hiroshi G. Okuno |
IROS | 2 |
| 2008 | Target speech detection and separation for humanoid robots in sparse dialogue with noisy home environmentsabstractIn normal human communication, people face the speaker when listening and usually pay attention to the speaker’ face. Therefore, in robot audition, the recognition of the front talker is critical for smooth interactions. This paper presents an enhanced speech detection method for a humanoid robot that can separate and recognize speech signals originating from the front even in noisy home environments. The robot audition system consists of a new type of voice activity detection (VAD) based on the complex spectrum circle centroid (CSCC) method and a maximum signal-to-noise (Max-SNR) beamformer. This VAD based on CSCC can classify speech signals that are retrieved at the frontal region of two microphones embedded on the robot. The system works in real-time without needing training filter coefficients given in advance even in a noisy environment (SNR ≫ 0 dB). It can cope with speech noise generated from televisions and audio devices that does not originate from the center. Experiments using a humanoid robot, SIG2, with two microphones showed that our system enhanced extracted target speech signals more than 12 dB (SNR) and the success rate of automatic speech recognition for Japanese words was increased about 17 points. Hyun-Don Kim, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 4 |
| 2008 | Design and evaluation of two-channel-based sound source localization over entire azimuth range for moving talkersabstractWe propose a way to evaluate various sound localization systems for moving sounds under the same conditions. To construct a database for moving sounds, we developed a moving sound creation tool using the API library developed by the ARINIS Company. We developed a two-channel-based sound source localization system integrated with a cross-power spectrum phase (CSP) analysis and EM algorithm. The CSP of sound signals obtained with only two microphones is used to localize the sound source without having to use prior information such as impulse response data. The EM algorithm helps the system cope with several moving sound sources and reduce localization error. We evaluated our sound localization method using artificial moving sounds and confirmed that it can well localize moving sounds slower than 1.125 rad/sec. Finally, we solve the problem of distinguishing whether sounds are coming from the front or back by rotating a robotpsilas head equipped with only two microphones. Our system was applied to a humanoid robot called SIG2, and we confirmed its ability to localize sounds over the entire azimuth range. Hyun-Don Kim, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 3 |
| 2008 | A robot listens to music and counts its beats aloud by separating music from counting voiceabstractThis paper presents a beat-counting robot that can count musical beats aloud, i.e., speak ldquoone, two, three, four, one, two, ...rdquo along music, while listening to music by using its own ears. Music-understanding robots that interact with humans should be able not only to recognize music internally, but also to express their own internal states. To develop our beat-counting robot, we have tackled three issues: (1) recognition of hierarchical beat structures, (2) expression of these structures by counting beats, and (3) suppression of counting voice (self-generated sound) in sound mixtures recorded by ears. The main issue is (3) because the interference of counting voice in music causes the decrease of the beat recognition accuracy. So we designed the architecture for music-understanding robot that is capable of dealing with the issue of self-generated sounds. To solve these issues, we took the following approaches: (1) beat structure prediction based on musical knowledge on chords and drums, (2) speed control of counting voice according to music tempo via a vocoder called STRAIGHT, and (3) semi-blind separation of sound mixtures into music and counting voice via an adaptive filter based on ICA (independent component analysis) that uses the waveform of the counting voice as a prior knowledge. Experimental result showed that suppressing robotpsilas own voice improved music recognition capability. Takeshi Mizumoto, Ryu Takeda, Kazuyoshi Yoshii, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 5 |
| 2008 | Active sensing based dynamical object feature extractionabstractThis paper presents a method to autonomously extract object features that describe their dynamics from active sensing experiences. The model is composed of a dynamics learning module and a feature extraction module. Recurrent Neural Network with Parametric Bias (RNNPB) is utilized for the dynamics learning module, learning and self-organizing the sequences of robot and object motions. A hierarchical neural network is linked to the input of RNNPB as the feature extraction module for extracting object features that describe the object motions. The two modules are simultaneously trained using image and motion sequences acquired from the robotpsilas active sensing with objects. Experiments are performed with the robotpsilas pushing motion with a variety of objects to generate sliding, falling over, bouncing, and rolling motions. The results have shown that the model is capable of extracting features that distinguish the characteristics of object dynamics. Shun Nishide, Tetsuya Ogata, Ryunosuke Yokoya, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno |
IROS | 2 |
| 2008 | Barge-in-able robot audition based on ICA and missing feature theory under semi-blind situationabstractThis paper describes a robot audition system that allows the user to barge-in; that is, the user can speak simultaneously when the robot is speaking. Our ldquobarge-in-ablerdquo system consists of two stages: (1) cancellation of robot speech and (2) recognition of the separated user speech under the ldquosemi-blind situationrdquo. The semi-blind situation is where a robotpsilas speech signal is known but a userpsilas speech signal is not. The first stage is achieved by using an adaptive filter based on time-frequency domain Independent Component Analysis, because that can separate robot speech more robustly against noise than conventional echo cancellers. To improve performance in online processing, we utilized known source normalization and the exponentially weighted stepsize method. The second stage is achieved by automatic speech recognition (ASR) based on the missing feature theory which provides robust recognition by exploiting the reliability of speech features distorted due to noise and/or separation. The semi-blind situation simplifies the estimation of such reliabilities. Experiments demonstrated that our system improved word correctness of ASR by 10.0%. Ryu Takeda, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 4 |
| 2008 | Design and Implementation of 3D Auditory Scene Visualizer towards Auditory Awareness with Face TrackingabstractIf machine audition can recognize an auditory scene containing simultaneous and moving talkers, what kinds of awareness will people gain from an auditory scene visualizer? This paper presents the design and implementation of 3D Auditory Scene Visualizer based on the visual information seeking mantra, i.e., ldquooverview first, zoom and filter, then details on demandrdquo. The machine audition system called HARK captures 3D sounds with a microphone array, localizes and separates sounds, and recognizes separated sounds by automatic speech recognition (ASR). The 3D visualizer implemented in Java 3D displays each sound stream as a beam originating from the center of the microphones (overview mode), shows temporal snapshots with/without specifying focusing areas (zoom and filter mode), and shows detailed information about a particular sound stream (details on demand). In the details-ondemand mode, ASR results are displayed in a ldquokaraokerdquo manner, i.e., character-by-character. This three-mode visualization will give the user auditory awareness enhanced by HARK. In addition, a face-tracking system automatically changes the focus of attention by tracking the userpsilas face. The resulting system is portable and can be deployed in any place, so it is expected to give more vivid awareness than expensive high-fidelity auditory scene reproduction systems. Yuji Kubota, Masatoshi Yoshida, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ISM | 4 |
| 2008 | SalienceGraph: Visualizing Salience Dynamics of Written Discourse by Using Reference Probability and PLSA
Shun Shiramatsu, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
PRICAI | 3 |
| 2008 | Managing out-of-grammar utterances by topic estimation with domain extensibility in multi-domain spoken dialogue systems
Kazunori Komatani, Satoshi Ikeda, Tetsuya Ogata, Hiroshi G. Okuno |
Speech Commun. | 3 |
| 2008 | An Efficient Hybrid Music Recommender System Using an Incrementally Trainable Probabilistic Generative ModelabstractThis paper presents a hybrid music recommender system that ranks musical pieces while efficiently maintaining collaborative and content-based data, i.e., rating scores given by users and acoustic features of audio signals. This hybrid approach overcomes the conventional tradeoff between recommendation accuracy and variety of recommended artists. Collaborative filtering, which is used on e-commerce sites, cannot recommend nonbrated pieces and provides a narrow variety of artists. Content-based filtering does not have satisfactory accuracy because it is based on the heuristics that the user's favorite pieces will have similar musical content despite there being exceptions. To attain a higher recommendation accuracy along with a wider variety of artists, we use a probabilistic generative model that unifies the collaborative and content-based data in a principled way. This model can explain the generative mechanism of the observed data in the probability theory. The probability distribution over users, pieces, and features is decomposed into three conditionally independent ones by introducing latent variables. This decomposition enables us to efficiently and incrementally adapt the model for increasing numbers of users and rating scores. We evaluated our system by using audio signals of commercial CDs and their corresponding rating scores obtained from an e-commerce site. The results revealed that our system accurately recommended pieces including nonrated ones from a wide variety of artists and maintained a high degree of accuracy even when new users and rating scores were added. Kazuyoshi Yoshii, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IEEE Trans. Speech Audio Process. | 4 |
| 2007 | Design and implementation of a robot audition system for automatic speech recognition of simultaneous speechabstractThis paper addresses robot audition that can cope with speech that has a low signal-to-noise ratio (SNR) in real time by using robot-embedded microphones. To cope with such a noise, we exploited two key ideas; Preprocessing consisting of sound source localization and separation with a microphone array, and system integration based on missing feature theory (MFT). Preprocessing improves the SNR of a target sound signal using geometric source separation with multichannel post-filter. MFT uses only reliable acoustic features in speech recognition and masks unreliable parts caused by errors in preprocessing. MFT thus provides smooth integration between preprocessing and automatic speech recognition. A real-time robot audition system based on these two key ideas is constructed for Honda ASIMO and Humanoid SIG2 with 8-ch microphone arrays. The paper also reports the improvement of ASR performance by using two and three simultaneous speech signals. Shun'ichi Yamamoto, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ASRU | 7 |
| 2007 | Integration and Adaptation of Harmonic and Inharmonic Models for Separating Polyphonic Musical SignalsabstractThis paper describes a sound source separation method for polyphonic sound mixtures of music to build an instrument equalizer for remixing multiple tracks separated from compact-disc recordings by changing the volume level of each track. Although such mixtures usually include both harmonic and inharmonic sounds, the difficulties in dealing with both types of sounds together have not been addressed in most previous methods that have focused on either of the two types separately. We therefore developed an integrated weighted-mixture model consisting of both harmonic-structure and inharmonic-structure tone models (generative models for the power spectrogram). On the basis of the MAP estimation using the EM algorithm, we estimated all model parameters of this integrated model under several original constraints for preventing over-training and maintaining intra-instrument consistency. Using standard MIDI files as prior information of the model parameters, we applied this model to compact-disc recordings and achieved the instrument equalizer. Katsutoshi Itoyama, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ICASSP (1) | 4 |
| 2007 | Vowel Imitation Using Vocal Tract Model and Recurrent Neural Network
Hisashi Kanda, Tetsuya Ogata, Kazunori Komatani, Hiroshi G. Okuno |
ICONIP (2) | 2 |
| 2007 | Enhancement of Self Organizing Network Elements for Supervised LearningabstractWe have proposed self-organizing network elements (SONE) as a learning method for robots to meet the requirements of autonomous exploration of effective output, simple external parameters, and low calculation costs. SONE can be used as an algorithm for obtaining network topology by propagating reinforcement signals between the elements of a network. Traditionally, the analysis of fundamental features in SONE and their application to supervised learning tasks were difficult because the learning method of SONE was limited to reinforcement learning. Here the abilities of generalization, incremental learning, and temporal sequence learning were evaluated using a supervised learning method with SONE. Moreover, the proposed method enabled our SONE to be applied to a greater variety of tasks. Chyon Hae Kim, Tetsuya Ogata, Shigeki Sugano |
ICRA | 2 |
| 2007 | Predicting Object Dynamics from Visual Images through Active Sensing ExperiencesabstractPrediction of dynamic features is an important task for determining the manipulation strategies of an object. This paper presents a technique for predicting dynamics of objects relative to the robot's motion from visual images. During the learning phase, the authors use recurrent neural network with parametric bias (RNNPB) to self-organize the dynamics of objects manipulated by the robot into the PB space. The acquired PB values, static images of objects, and robot motor values are input into a hierarchical neural network to link the static images to dynamic features (PB values). The neural network extracts prominent features that induce each object dynamics. For prediction of the motion sequence of an unknown object, the static image of the object and robot motor value are input into the neural network to calculate the PB values. By inputting the PB values into the closed loop RNNPB, the predicted movements of the object relative to the robot motion are calculated sequentially. Experiments were conducted with the humanoid robot Robovie-IIs pushing objects at different heights. Reducted grayscale images and shoulder pitch angles were input into the neural network to predict the dynamics of target objects. The results of the experiment proved that the technique is efficient for predicting the dynamics of the objects. Shun Nishide, Tetsuya Ogata, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno |
ICRA | 2 |
| 2007 | Distance Estimation of Hidden Objects Based on Acoustical Holography by applying Acoustic Diffraction of Audible SoundabstractOcclusion is a problem for range finders; ranging systems using cameras or lasers cannot be used to estimate distance to an object (hidden object) that is occluded by another (obstacle). We developed a method to estimate the distance to the hidden object by applying acoustic diffraction of audible sound. Our method is based on time-of-flight (TOF), which has been used in ultrasound ranging systems. We determined the best frequency of audible sound and designed its optimal modulated signal for our system. We determined that the system estimates the distance to the hidden object as well as the obstacle. However, the measurement signal obtained from the hidden object was weak. Thus, interference from sound signals reflected from other objects or walls was not negligible. Therefore, we combined acoustical holography (AH) and TOF, which enabled a partial analysis of the reflection sound intensity field around the obstacle and hidden object. Our method was effective for ranging two objects of the same size within a 1.2 m depth range. The accuracy of our method was 3 cm for the obstacle, and 6 cm for the hidden object. Haruhiko Niwa, Tetsuya Ogata, Kazunori Komatani, Hiroshi G. Okuno |
ICRA | 2 |
| 2007 | Human-Robot Cooperation using Quasi-symbols Generated by RNNPB ModelabstractWe describe a means of human robot interaction based not on natural language but on "quasi symbols," which represent sensory-motor dynamics in the task and/or environment. It thus overcomes a key problem of using natural language for human-robot interaction - the need to understand the dynamic context. The quasi-symbols used are motion primitives corresponding to the attractor dynamics of the sensory-motor flow. These primitives are extracted from the observed data using the recurrent neural network with parametric bias (RNNPB) model. Binary representations based on the model parameters were implemented as quasi symbols in a humanoid robot, Robovie. The experiment task was robot-arm operation on a table. The quasi-symbols acquired by learning enabled the robot to perform novel motions. A person was able to control the arm through speech interaction using these quasi-symbols. These quasi symbols formed a hierarchical structure corresponding to the number of nodes in the model. The meaning of some of the quasi-symbols depended on the context, indicating that they are useful for human-robot interaction. Tetsuya Ogata, Shohei Matsumoto, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno |
ICRA | 1 |
| 2007 | Real-Time Auditory and Visual Talker Tracking Through Integrating EM Algorithm and Particle Filter
Hyun-Don Kim, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE | 3 |
| 2007 | Evaluation of Two Simultaneous Continuous Speech Recognition with ICA BSS and MFT-Based ASR
Ryu Takeda, Shun'ichi Yamamoto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE | 4 |
| 2007 | Topic estimation with domain extensibility for guiding user's out-of-grammar utterances in multi-domain spoken dialogue systemsabstractIn a multi-domain spoken dialogue system, a user’s utterances are more prone to be out-of-grammar, because this kind of system deals with more tasks than a single-domain system. We defined a topic as a domain about which users want to find more information, and we developed a method of recovering out-ofgrammar utterances based on topic estimation, i.e., by providing a help message in the estimated domain. Moreover, the domain extensibility, that is, to facilitate adding new domains, should be inherently retained in multi-domain systems. We therefore collected documents from the Web as training data for topic estimation. Because the data contained not a few noises, we used Latent Semantic Mapping (LSM), which enables robust topic estimation by removing the effect of noise from the data. The experimental results based on using 272 utterances collected with a Woz-like method showed that our method increased the topic estimation accuracy by 23.1 points from the baseline. Index Terms: multi-domain spoken dialogue system, topic estimation, out-of-grammar utterance 1. Satoshi Ikeda, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 3 |
| 2007 | Vocal imitation using physical vocal tract modelabstractA vocal imitation system was developed using a computational model that supports the motor theory of speech perception. A critical problem in vocal imitation is how to generate speech sounds produced by adults, whose vocal tracts have physical properties (i.e., articulatory motions) differing from those of infants’ vocal tracts. To solve this problem, a model based on the motor theory of speech perception, was constructed. This model suggests that infants simulate the speech generation by estimating their own articulatory motions in order to interpret the speech sounds of adults. Applying this model enables the vocal imitation system to estimate articulatory motions for unexperienced speech sounds that have not actually been generated by the system. The system was implemented by using Recurrent Neural Network with Parametric Bias (RNNPB) and a physical vocal tract model, called the Maeda model. Experimental results demonstrated that the system was sufficiently robust with respect to individual differences in speech sounds and could imitate unexperienced vowel sounds. Hisashi Kanda, Tetsuya Ogata, Kazunori Komatani, Hiroshi G. Okuno |
IROS | 2 |
| 2007 | Auditory and visual integration based localization and tracking of humans in daily-life environmentsabstractThe purpose of this research is to develop techniques that enable robots to choose and track a desired person for interaction in daily-life environments. Therefore, localizing multiple moving sounds and human faces is necessary so that robots can locate a desired person. For sound source localization, we used a cross-power spectrum phase analysis (CSP) method and showed that CSP can localize sound sources only using two microphones and does not need impulse response data. An expectation-maximization (EM) algorithm was shown to enable a robot to cope with multiple moving sound sources. For face localization, we developed a method that can reliably detect several faces using the skin color classification obtained by using the EM algorithm. To deal with a change in color state according to illumination condition and various skin colors, the robot can obtain new skin color features of faces detected by OpenCV, an open vision library, for detecting human faces. Finally, we developed a probability based method to integrate auditory and visual information and to produce a reliable tracking path in real time. Furthermore, the developed system chose and tracked people while dealing with various background noises that are considered loud, even in the daily-life environments. Hyun-Don Kim, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 3 |
| 2007 | Two-way translation of compound sentences and arm motions by recurrent neural networksabstractWe present a connectionist model that combines motions and language based on the behavioral experiences of a real robot. Two models of recurrent neural network with parametric bias (RNNPB) were trained using motion sequences and linguistic sequences. These sequences were combined using their respective parameters so that the robot could handle many-to-many relationships between motion sequences and linguistic sequences. Motion sequences were articulated into some primitives corresponding to given linguistic sequences using the prediction error of the RNNPB model. The experimental task in which a humanoid robot moved its arm on a table demonstrated that the robot could generate a motion sequence corresponding to given linguistic sequence even if the motions or sequences were not included in the training data, and vice versa. Tetsuya Ogata, Masamitsu Murase, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno |
IROS | 1 |
| 2007 | Exploiting known sound source signals to improve ICA-based robot audition in speech separation and recognitionabstractThis paper describes a new semi-blind source separation (semi-BSS) technique with independent component analysis (ICA) for enhancing a target source of interest and for suppressing other known interference sources. The semi BSS technique is necessary for double-talk free robot audition systems in order to utilize known sound source signals such as self speech, music, or TV-sound, through a line-in or ubiquitous network. Unlike the conventional semi-BSS with ICA, we use the time-frequency domain convolution model to describe the reflection of the sound and a new mixing process of sounds for ICA. In other words, we consider that reflected sounds during some delay time are different from the original. ICA then separates the reflections as other interference sources. The model enables us to eliminate the frame size limitations of the frequency-domain ICA, and ICA can separate the known sources under a highly reverberative environment. Experimental results show that our method outperformed the conventional semi-BSS using ICA under simulated normal and highly reverberative environments. Ryu Takeda, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 4 |
| 2007 | Discovery of other individuals by projecting a self-model through imitationabstractThis paper proposes a novel model which enables a humanoid robot infant to discover other individual (e.g. human parent). In this work, the authors define “other individual” as an actor which can be predicted by a self-model. For modeling the developmental process of discovering ability, the following three approaches are employed. (i) Projection of a selfmodel for predicting other individual’s actions. (ii) Mediation by a physical object between self and other individual. (iii) Introduction of infant imitation by parent. For creating the self-model of a robot, we apply Recurrent Neural Network with Parametric Bias (RNNPB) model which can learn the robot’s body dynamics. For the other-model of a human, conventional hierarchical neural networks are attached to the RNNPB model as “conversion modules”. Our target task is a moving an object. For evaluation of our model, human discovery experiments by the robot projecting its self-model were conducted. The results demonstrated that our method enabled the robot to predict the human’s motions, and to estimate the human’s position fairly accurately, which proved its adequacy. Ryunosuke Yokoya, Tetsuya Ogata, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno |
IROS | 2 |
| 2007 | A biped robot that keeps steps in time with musical beats while listening to music with its own earsabstractWe aim at enabling a biped robot to interact with humans through real-world music in daily-life environments, e.g., to autonomously keep its steps (stamps) in time with musical beats. To achieve this, the robot should be able to robustly predict the beat times in real time while listening to musical performance with its own ears (head-embedded microphones). However, this has not previously been addressed in most studies on music-synchronized robots due to the difficulty in predicting the beat times in real-world music. To solve this problem, we implemented a beat-tracking method developed in the field of music information processing. The predicted beat times are then used by a feedback-control method that adjusts the robot's step intervals to synchronize its steps in time with the beats. The experimental results show that the robot can adjust its steps in time with the beat times as the tempo changes. The resulting robot needed about 25 [s] to recognize the tempo change after it and then synchronize its steps. Kazuyoshi Yoshii, Kazuhiro Nakadai, Toyotaka Torii, Yuji Hasegawa, Hiroshi Tsujino, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 7 |
| 2007 | Auditory and Visual Integration based Localization and Tracking of Multiple Moving Sounds in Daily-life EnvironmentsabstractThis paper presents techniques that enable talker tracking for effective human-robot interaction. To track moving people in daily-life environments, localizing multiple moving sounds is necessary so that robots can locate talkers. However, the conventional method requires an array of microphones and impulse response data. Therefore, we propose a way to integrate a cross-power spectrum phase analysis (CSP) method and an expectation-maximization (EM) algorithm. The CSP can localize sound sources using only two microphones and does not need impulse response data. Moreover, the EM algorithm increases the system's effectiveness and allows it to cope with multiple sound sources. We confirmed that the proposed method performs better than the conventional method. In addition, we added a particle filter to the tracking process to produce a reliable tracking path and the particle filter is able to integrate audio-visual information effectively. Furthermore, the applied particle filter is able to track people while dealing with various noises that are even loud sounds in the daily-life environments. Hyun-Don Kim, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
RO-MAN | 3 |
| 2006 | F0 Estimation Method for Singing Voice in Polyphonic Audio Signal Based on Statistical Vocal Model and Viterbi SearchabstractThis paper describes a method for estimating F0s of vocal from polyphonic audio signals. Because melody is sung by a singer in many musical pieces, the estimation of F0s of the vocal part is useful for many applications. Based on existing multiple-F0 estimation method, we evaluate the vocal probabilities of the harmonic structure of each F0 candidate. In order to calculate the vocal probabilities of the harmonic structure, we extract and resynthesize the harmonic structure by using a sinusoidal model and extract feature vectors. Then, we evaluate the vocal probability by using vocal and non-vocal Gaussian mixture models (GMMs). Finally, we track F0 trajectories using these probabilities based on Viterbi search. Experimental results show that our method improves estimation accuracy from 78.1% to 84.3%, which is 28.3% reduction of misestimation Hiromasa Fujihara, Tetsuro Kitahara, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ICASSP (5) | 5 |
| 2006 | Instrogram: A New Musical Instrument Recognition Technique Without Using Onset Detection NOR F0 EstimationabstractThis paper describes a new technique for recognizing musical instruments in polyphonic music. Because the conventional framework for musical instrument recognition in polyphonic music had to estimate the onset time and fundamental frequency (F0) of each note, instrument recognition strictly suffered from errors of onset detection and F0 estimation. Unlike such a note-based processing framework, our technique calculates the temporal trajectory of instrument existence probabilities for every possible F0, and the results are visualized with a spectrogram-like graphical representation called instrogram. The instrument existence probability is defined as the product of a nonspecific instrument existence probability calculated using PreFEst and a conditional instrument existence probability calculated using the hidden Markov model. Experimental results show that the obtained instrograms reflect the actual instrumentations and facilitate instrument recognition Tetsuro Kitahara, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ICASSP (5) | 4 |
| 2006 | An Error Correction Framework Based on Drum Pattern Periodicity for Improving Drum Sound DetectionabstractThis paper presents a framework for correcting errors of automatic drum sound detection focusing on the periodicity of drum patterns. We define drum patterns as periodic structures found in onset sequences of bass and snare drum sounds. Our framework extracts periodic drum patterns from imperfect onset sequences of detected drum sounds (bottom-up processing) and corrects errors using the periodicity of the drum patterns (top-down processing). We implemented this framework on our drum-sound detection system. We first obtained onset sequences of the drum sounds with our system and extracted drum patterns. On the basis of our observation that the same drum patterns tend to be repeated, we detected time points which deviate from the periodicity as error candidates. Finally, we verified each error candidate to judge whether it is an actual onset or not. Experiments of drum sound detection for polyphonic audio signals of popular CD recordings showed that our correction framework improved the average detection accuracy from 77.4% to 80.7% Kazuyoshi Yoshii, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ICASSP (5) | 4 |
| 2006 | Reinforcement Learning Algorithm with CTRNN in Continuous Action Space
Hiroaki Arie, Jun Namikawa, Tetsuya Ogata, Jun Tani, Shigeki Sugano |
ICONIP (1) | 3 |
| 2006 | Genetic Algorithm-Based Improvement of Robot Hearing Capabilities in Separating and Recognizing Simultaneous Speech Signals
Shun'ichi Yamamoto, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Ryu Takeda, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE | 8 |
| 2006 | Speaker identification under noisy environments by using harmonic structure extraction and reliable frame weightingabstractWe present methods for automatic speaker identification in noisy environments. To improve noise robustness of speaker identification, we developed two methods, theharmonic structure extraction method and the reliable frame weighting method. The harmonic structure extraction method enables the speaker of input speech signals to be identified after environmental noise has been reduced. This method first extracts harmonic components of the speech from the sound mixtures and then resynthesizes a clean speech signal by using a sinusoidal model driven by harmonic components. The reliable frame weighting method then determines how each frame of the resynthesized speech is reliable (i.e. little influenced by environmental noises) by using two Gaussian mixture models for the speech and noise. The speaker can be robustly identified by attaching importance to reliable frames. Experimental results with thirty speakers showed that our method was able to reduce the influences of environmental noise and achieved an error rate of 10.7%, while the error rate for a conventional method was 18.9%. Index Terms: speaker identification, noise robustness, voice extraction, voice reliability, Gaussian mixture model. Hiromasa Fujihara, Tetsuro Kitahara, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 5 |
| 2006 | Dynamic help generation by estimating user²s mental model in spoken dialogue systemsabstractIn a speech interface, a gap between a user’s mental model and actual structures of systems tends to be large because the amount of information conveyed by speech is limited. We address dynamic help generation adapted to users, to decrease the gap between them. We defined a domain concept tree as an expression of a system’s actual structure. We estimated and maintained user’s knowledge about the system on the tree. Every node in the tree has values representing the degree to which a user understands the concepts corresponding to the nodes. The values are updated based on the content of user’s utterances and help messages the system gives. Help messages provided for users are determined by referring to the domain concept tree and identifying concepts the user does not understand. We evaluated our method by testing twelve novice subjects. Both the average time to complete tasks and the number of utterances significantly decreased because of the help messages provided by our method. Index Terms: spoken dialogue system, adaptive help generation, novice user Yuichiro Fukubayashi, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 3 |
| 2006 | Improving speech recognition of two simultaneous speech signals by integrating ICA BSS and automatic missing feature mask generation
Ryu Takeda, Shun'ichi Yamamoto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 4 |
| 2006 | Efficient Organization of Network Topology based on Reinforcement SignalsabstractWe developed a learning system for autonomous robots that allows for autonomous exploration of the effective output, and has simple external parameters and a low calculation cost. We propose the concept of self-organizing network elements (SONE) for creating learning systems with these characteristics. We created and evaluated a self-organizing logic circuit by using this concept. Our results indicated this learning system had the characteristics Chyon Hae Kim, Shigeki Sugano, Tetsuya Ogata |
IROS | 3 |
| 2006 | Multiple Acoustical Holography Method for Localization of Objects in Broad Range using Audible SoundabstractThis paper describes a new acoustic localization method using audible sound, which can be applied over a broader range of search directions. In the field of robotics, most conventional indoor localization systems based on sonar range finders use ultrasound to obtain a highly accurate distance. Because ultrasound has high directivity, many measurements are required to localize objects in a large space. To achieve localization with one-time measurement, we use audible sounds. We then calculate an intensity field of the reflection sound to estimate object positions. Although acoustical holography (AH) is a well-known technique to do this, it has problems in that it generates false images. We propose multiple AH (MAH) to solve this problem. The method is used to divide a measurement plane into sub-planes and to apply AH to each sub-plane. By integrating the results of applying AH to the sub-planes, false images can be suppressed because the positions of the false images differ depending on the position of the sub-plane. In addition, we use multiple frequencies to advance an accuracy of the localization based on MAH in a real environment. We constructed a localization system with only one speaker and a microphone array. In both simulation and actual experiments, we confirmed that MAH was effective method for the suppression of false image and could be used within the range of an angle view of 120 deg Haruhiko Niwa, Tetsuya Ogata, Kazunori Komatani, Hiroshi G. Okuno |
IROS | 2 |
| 2006 | Adaptive Human-Robot Interaction System using Interactive ECabstractWe created a human-robot communication system that can adapt to user preferences that can easily change through communication. Even if any learning algorithms are used, evaluating the human-robot interaction is indispensable and difficult. To solve this problem, we installed a machine learning algorithm called interactive evolutionary computation (IEC) into a communication robot named WAMOEBA-3. IEC is a kind of evolutionary computation like a genetic algorithm. With IEC, the fitness function is performed by each user. We carried out experiments on the communication learning system using an advanced IEC system named HMHE. Before the experiments, we did not tell the subjects anything about the robot, so the interaction differed among the experimental subjects. We could observe mutual adaptation, because some subjects noticed the robot's functions and changed their interaction. From the results, we confirmed that, in spite of the changes of the preferences, the system can adapt to the interaction of multiple users Yuki Suga, Chihiro Endo, Daizo Kobayashi, Takeshi Matsumoto, Shigeki Sugano, Tetsuya Ogata |
IROS | 6 |
| 2006 | Missing-Feature based Speech Recognition for Two Simultaneous Speech Signals Separated by ICA with a pair of Humanoid EarsabstractRobot audition is a critical technology in making robots symbiosis with people. Since we hear a mixture of sounds in our daily lives, sound source localization and separation, and recognition of separated sounds are three essential capabilities. Sound source localization has been recently studied well for robots, while the other capabilities still need extensive studies. This paper reports the robot audition system with a pair of omni-directional microphones embedded in a humanoid to recognize two simultaneous talkers. It first separates sound sources by independent component analysis (ICA) with single-input multiple-output (SIMO) model. Then, spectral distortion for separated sounds is estimated to identify reliable and unreliable components of the spectrogram. This estimation generates the missing feature masks as spectrographic masks. These masks are then used to avoid influences caused by spectral distortion in automatic speech recognition based on missing-feature method. The novel ideas of our system reside in estimates of spectral distortion of temporal-frequency domain in terms of feature vectors. In addition, we point out that the voice-activity detection (VAD) is effective to overcome the weak point of ICA against the changing number of talkers. The resulting system outperformed the baseline robot audition system by 15% Ryu Takeda, Shun'ichi Yamamoto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 4 |
| 2006 | Real-Time Robot Audition System That Recognizes Simultaneous Speech in The Real WorldabstractThis paper presents a robot audition system that recognizes simultaneous speech in the real world by using robot-embedded microphones. We have previously reported missing feature theory (MFT) based integration of sound source separation (SSS) and automatic speech recognition (ASR) for building robust robot audition. We demonstrated that a MFT-based prototype system drastically improved the performance of speech recognition even when three speakers talked to a robot simultaneously. However, the prototype system had three problems; being offline, hand-tuning of system parameters, and failure in voice activity detection (VAD). To attain online processing, we introduced FlowDesigner-based architecture to integrate sound source localization (SSL), SSS and ASR. This architecture brings fast processing and easy implementation because it provides a simple framework of shared-object-based integration. To optimize the parameters, we developed genetic algorithm (GA) based parameter optimization, because it is difficult to build an analytical optimization model for mutually dependent system parameters. To improve VAD, we integrated new VAD based on a power spectrum and location of a sound source into the system, since conventional VAD relying only on power often fails due to low signal-to-noise ratio of simultaneous speech. We, then, constructed a robot audition system for Honda ASIMO. As a result, we showed that the system worked online and fast, and had a better performance in robustness and accuracy through experiments on recognition of simultaneous speech in a noisy and echoic environment Shun'ichi Yamamoto, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 7 |
| 2006 | Experience Based Imitation Using RNNPBabstractRobot imitation is a useful and promising alternative to robot programming. Robot imitation involves two crucial issues. The first is how a robot can imitate a human whose physical structure and properties differ greatly from its own. The second is how the robot can generate various motions from finite programmable patterns (generalization). This paper describes a novel approach to robot imitation based on its own physical experiences. Let us consider a target task of moving an object on a table. For imitation, we focused on an active sensing process in which the robot acquires the relation between the object's motion and its own arm motion. For generalization, we applied a recurrent neural network with parametric bias (RNNPB) model to enable recognition/generation of imitation motions. The robot associates the arm motion which reproduces the observed object's motion presented by a human operator. Experimental results demonstrated that our method enabled the robot to imitate not only motion it has experienced but also unknown motion, which proved its capability for generalization Ryunosuke Yokoya, Tetsuya Ogata, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno |
IROS | 2 |
| 2006 | Automatic Synchronization between Lyrics and Music CD Recordings Based on Viterbi Alignment of Segregated Vocal SignalsabstractThis paper describes a system that can automatically synchronize between polyphonic musical audio signals and corresponding lyrics. Although there were methods that can synchronize between monophonic speech signals and corresponding text transcriptions by using Viterbi alignment techniques, they cannot be applied to vocals in CD recordings because accompaniment sounds often overlap with vocals. To align lyrics with such vocals, we therefore developed three methods: a method for segregating vocals from polyphonic sound mixtures, a method for detecting vocal sections, and a method for adapting a speech-recognizer phone model to segregated vocal signals. Experimental results for 10 Japanese popular-music songs showed that our system can synchronize between music and lyrics with satisfactory accuracy for 8 songs Hiromasa Fujihara, Masataka Goto, Jun Ogata, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ISM | 5 |
| 2006 | Musical Instrument Recognizer "Instrogram" and Its Application to Music Retrieval Based on Instrumentation SimilarityabstractInstrumentation is an important cue in retrieving musical content. Conventional methods for instrument recognition performing notewise require accurate estimation of the onset time and fundamental frequency (FO) for each note, which is not easy in polyphonic music. This paper presents a non-notewise method for instrument recognition in polyphonic musical audio signals. Instead of such note-wise estimation, our method calculates the temporal trajectory of instrument existence probabilities for every FO and visualizes it as a spectrogram-like graphical representation, called an instrogram. This method can avoid the influence by errors of onset detection and FO estimation because it does not use them. We also present methods for MPEG-7-based instrument annotation and music information retrieval based on the similarity between instrograms. Experimental results with realistic music show the average accuracy of 76.2% for the instrument annotation and that the instrogram-based similarity measure represents the actual instrumentation similarity better than an MFCC-based one Tetsuro Kitahara, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ISM | 4 |
| 2006 | Recognition of Simultaneous Speech by Estimating Reliability of Separated Signals for Robot Audition
Shun'ichi Yamamoto, Ryu Takeda, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
PRICAI | 8 |
| 2005 | Enhanced Robot Speech Recognition Based on Microphone Array Source Separation and Missing Feature TheoryabstractA humanoid robot under real-world environments usually hears mixtures of sounds, and thus three capabilities are essential for robot audition; sound source localization, separation, and recognition of separated sounds. While the first two are frequently addressed, the last one has not been studied so much. We present a system that gives a humanoid robot the ability to localize, separate and recognize simultaneous sound sources. A microphone array is used along with a real-time dedicated implementation of Geometric Source Separation (GSS) and a multi-channel post-filter that gives us a further reduction of interferences from other sources. An automatic speech recognizer (ASR) based on the Missing Feature Theory (MFT) recognizes separated sounds in real-time by generating missing feature masks automatically from the post-filtering step. The main advantage of this approach for humanoid robots resides in the fact that the ASR with a clean acoustic model can adapt the distortion of separated sound by consulting the post-filter feature masks. Recognition rates are presented for three simultaneous speakers located at 2m from the robot. Use of both the post-filter and the missing feature mask results in an average reduction in error rate of 42% (relative). Shun'ichi Yamamoto, Jean-Marc Valin, Kazuhiro Nakadai, Jean Rouat, François Michaud, Tetsuya Ogata, Hiroshi G. Okuno |
ICRA | 6 |
| 2005 | Distance-Based Dynamic Interaction of Humanoid Robot with Multiple People
Tsuyoshi Tasaki, Shohei Matsumoto, Hayato Ohba, Mitsuhiko Toda, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE | 6 |
| 2005 | Contextual constraints based on dialogue models in database search task for spoken dialogue systemsabstractThis paper describes the incorporation of contextual information into spoken dialogue systems in the database search task. Appropriatedialoguemodeling is requiredto manageautomatic speech recognition (ASR) errors using dialogue-level information. We define two dialogue models: a model for dialogue flow and a model of structured dialogue history. The model for dialogueflowassumesdialoguesin the databasesearchtaskconsist of only two modes. In the structured dialogue history model, query conditions are maintained as a tree structure, taking into consideration their inputted order. The constraints derived from these models are integrated by using a decision tree learning, so that the system candeterminea dialogueact of the utteranceand whether each content word should be accepted or rejected, even when it contains ASR errors. The experimental result showed that our method could interpret content words better than conventional one without the contextual information. Furthermore, it was also shown that our method was domain-independent because it achieved equivalent accuracy in another domain without any more training. Kazunori Komatani, Naoyuki Kanda, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 3 |
| 2005 | Multiple moving speaker tracking by microphone array on mobile robotabstractReal-world applications often require tracking multiple moving speakers for improving human-robot interactions and/or sound source separation. This paper presents multiple moving speaker tracking using an 8ch microphone array system installed on a mobile robot. This problem is difficult because the system does not assume that sound sources and/or the microphone array are fixed. Our solutions consist of two key ideas – time delay of arrival estimation, and multiple Kalman filters. The former localizes multiple sound sources based on beamforming in real time. Non-linear movements are tracked by using a set of Kalman filters with different history lengths in order to reduce errors in tracking multiple moving speakers under noisy and echoic environments. For quantitative evaluation of the tracking, motion references of sound sources and a mobile robot, called SIG2, were measured accurately by ultrasonic 3D tag sensors. As a result, we showed that the system tracked three simultaneous sound sources even when SIG2 moved in a room with large reverberation due to glass walls. 1. Masamitsu Murase, Shun'ichi Yamamoto, Jean-Marc Valin, Kazuhiro Nakadai, Kentaro Yamada, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 7 |
| 2005 | Extracting multi-modal dynamics of objects using RNNPBabstractDynamic features play an important role in recognizing objects that have similar static features in colors and or shapes. This paper focuses on active sensing that exploits dynamic feature of an object. An extended version of the robot, Robovie-IIs, moves an object by its arm to obtain its dynamic features. Its issue is how to extract symbols from various kinds of temporal states of the object. We use the recurrent neural network with parametric bias (RNNPB) that generates self-organized nodes in the parametric bias space. The RNNPB with 42 neurons was trained with the data of sounds, trajectories, and tactile sensors generated while the robot was moving/hitting an object with its own arm. The clusters of 20 kinds of objects were successfully self-organized. The experiments with unknown (not trained) objects demonstrated that our method configured them in the PB space appropriately, which proves its generalization capability. Tetsuya Ogata, Hayato Ohba, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno |
IROS | 1 |
| 2005 | Interactive evolution of human-robot communication in real worldabstractThis paper describes how to implement interactive evolutionary computation (IEC) into a human-robot communication system. IEC is an evolutionary computation (EC) in which the fitness function is performed by human assessors. We used IEC to configure the human-robot communication system. We have already simulated IEC's application. In this paper, we implemented IEC into a real robot. Since this experiment leads considerable burdens on both the robot and experimental subjects, we propose the human-machine hybrid evaluation (HMHE) to increase the diversity within the genetic pool without increasing the number of interactions. We used a communication robot, WAMOEBA-3 (Waseda artificial mind on emotion base), which is appropriate for this experiment. In the experiment, human assessors interacted with WAMOEBA-3 in various ways. The fitness values increased gradually, and assessors felt the robot learnt the motions they desired. Therefore, it was confirmed that the IEC is most suitable as the communication learning system. Yuki Suga, Yoshinori Ikuma, Daisuke Nagao, Shigeki Sugano, Tetsuya Ogata |
IROS | 5 |
| 2005 | Spatially mapping of friendliness for human-robot interactionabstractIt is important that robots interact with multiple people. However, most research has dealt with only interaction between one robot and one person and assumed that the distance between them does not change. This paper focuses on the spatial relationships between a robot and multiple people during interaction. Based on the distance between them, our robot selects appropriate functions to use. It does this using a method we developed for spatially mapping the friendliness of each space around the robot. The robot interacts with the highest friendliness spaces (people) selectively, thereby enabling interaction between the robot and multiple people. Our humanoid robot, SIG2 which the proposed method was implemented into, interacted with about 30 visitors, at the Kyoto University Museum. The results obtained using questionnaires after interaction showed that the actions of SIG2 were easy to understand even when it interacted with multiple people at the same time and that SIG2 behaved in a friendly manner. Tsuyoshi Tasaki, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 3 |
| 2005 | Making a robot recognize three simultaneous sentences in real-timeabstractA humanoid robot under real-world environments usually hears mixtures of sounds, and thus three capabilities are essential for robot audition; sound source localization, separation, and recognition of separated sounds. We have adopted the missing feature theory (MFT) for automatic recognition of separated speech, and developed the robot audition system. A microphone array is used along with a real-time dedicated implementation of geometric source separation (GSS) and a multi-channel post-filter that gives us a further reduction of interferences from other sources. The automatic speech recognition based on MFT recognizes separated sounds by generating missing feature masks automatically from the post-filtering step. The main advantage of this approach for humanoid robots resides in the fact that the ASR with a clean acoustic model can adapt the distortion of separated sound by consulting the post-filter feature masks. In this paper, we used the improved Julius as an MFT-based automatic speech recognizer (ASR). The Julius is a real-time large vocabulary continuous speech recognition (LVCSR) system. We performed the experiment to evaluate our robot audition system. In this experiment, the system recognizes a sentence, not an isolated word. We showed the improvement in the system performance through three simultaneous speech recognition on the humanoid SIG2. Shun'ichi Yamamoto, Kazuhiro Nakadai, Jean-Marc Valin, Jean Rouat, François Michaud, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 7 |
| 2005 | Walking with body-sense in virtual space using the nonlinear oscillatorabstractThis paper presents a novel construction of a locomotion system that compensates for walking-sense with simple interfaces using hands or fingers. The realization of moving with body-sense in virtual space has required high-quality system designs including walking cancellations, haptic feedback, and high-resolution displays. However, such approaches result in increasing costs of calculation and space, which obstruct the spread of VR technology. We therefore propose a new framework of locomotion systems with simple interfaces that give users "passivity and restraint", which are essential components of walking-sense. They are realized with mutual entrainment between a nonlinear oscillator and users' input. Two experiments were conducted for evaluation. The first one showed that our system gives users sense of distance based on a body standard and sense of rhythm with stable input. The second one demonstrated that the users of our system can experience a subjective body-sense and a sense of velocity. Kenri Kodaka, Tetsuya Ogata, Hiroshi G. Okuno |
SMC | 2 |
| 2004 | Open-End Human Robot Interaction from the Dynamical Systems Perspective: Mutual Adaptation and Incremental LearningabstractIn this paper, we experimentally investigated the open-end interaction generated by the mutual adaptation between humans and robot. Its essential characteristic, incremental learning, is examined using the dynamical systems approach. Our research concentrated on the navigation system of a specially developed humanoid robot called Robovie and seven human subjects whose eyes were covered, making them dependent on the robot for directions. We used the usual feed-forward neural network (FFNN) without recursive connections and the recurrent neural network (RNN) for the robot control. Although the performances obtained with both the RNN and the FFNN improved in the early stages of learning, as the subject changed the operation by learning on its own, all performances gradually became unstable and failed. Next, we used a 'consolidation-learning algorithm' as a model of the hippocampus in the brain. In this method, the RNN was trained by both new data and the rehearsal outputs of the RNN not to damage the contents of current memory. The proposed method enabled the robot to improve performance even when learning continued for a long time (open-end). The dynamical systems analysis of RNNs supports these differences and also showed that the collaboration scheme was developed dynamically along with succeeding phase transitions. Tetsuya Ogata, Shigeki Sugano, Jun Tani |
IEA/AIE | 1 |
| 2004 | Disambiguation in determining phonemes of sound-imitation words for environmental sound recognitionabstractOnomatopoeia, or sound-imitation words (SIWs) are important in informing sound events in human-computer communication. One problem is listener-dependency in recognizing environmental sounds by means of SIWs, that is, different listener hears the same environmental sound as a different SIW even under the same condition. Therefore, the use of usual Japanese phonemes is not adequate to express SIWs. To cope with this ambiguity problem of phoneme determination, we designed a set of new phonemes, referred to as the basic phoneme-groups, to represent environmental sounds. The basic phonemegroup consists of one or more Japanese phonemes, and thus the ambiguity problem is resolved based on it by generating one or more SIWs for a sound event. An HMM-based scheme is adopted to recognize SIWs using the phoneme-groups. Listening experiments with seven subjects showed that automatic SIW recognition based on the basic phoneme-groups outperformed ones based on the other types of phonemes. The recall and precision rate were 56.4% and 72.2%, respectively. Kazushi Ishihara, Yuya Hattori, Tomohiro Nakatani, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 5 |
| 2004 | Robot motion control using listener's back-channels and head gesture informationabstractA novel method is described for robot gestures and utterances during a dialogue based on the listener’s understanding and interest, which are recognized from back-channels and head gestures. “Back-channels” are defined as sounds like ‘uhhuh’ uttered by a listener during a dialogue, and “head gestures” are defined as nod and tilt motions of the listener’s head. The back-channels are recognized using sound features such as power and fundamental frequency. The head gestures are recognized using the movement of the skin-color area and the optical flow data. Based on the estimated understanding and interest of the listener, the speed and size of robot motions are changed. This method was implemented in a humanoid robot called SIG2. Experiments with six participants demonstrated that the proposed method enabled the robot to increase the listener’s level of interest against the dialogue. Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno, Tsuyoshi Tasaki, Takeshi Yamaguchi |
INTERSPEECH | 2 |
| 2004 | Human-robot collaboration using behavioral primitivesabstractA novel approach to human-robot collaboration based on quasi-symbolic expressions is proposed. The target task is navigation in which a person with his or her covered and a humanoid robot collaborate in a context-dependent manner. The robot uses a recurrent neural net with parametric bias (RNNPB) model to acquire the behavioral primitives, which are sensory-motor units, composing the whole task. The robot expresses the PB dynamics as primitives using symbolic sounds, and the person influences these dynamics through tactile sensors attached to the robot. Experiments with six participants demonstrated that the level of influence the person has on the PB dynamics is strongly related to task performance, the person's subjective impressions, and the prediction error of the RNNPB model (task stability). Simulation experiments demonstrated that the subjective impressions of the correspondence between the utterance sounds (the PB values) and the motions were well reproduced by the rehearsal of the RNNPB model. Tetsuya Ogata, Masaki Matsunaga, Shigeki Sugano, Jun Tani |
IROS | 1 |
| 2004 | Human-robot communication using multiple recurrent neural networksabstractOn the methodology of robotic design from the traditional view of communication which is assumed as a symbol process, robots are forced to confront the symbol grounding problem. However, if communication is assumed as the analog dynamics and robots are driven by it, robots can avoid the problem and be situated in the environment and to other agents. In this paper we introduce a new communication system constructed from the view of dynamical systems to achieve the situatedness. This system is that there is a robot in a virtual environment and the control of the robot is shared by human operation using a joystick and a robot controller. As the controller, we adopt multiple recurrent neural networks (MRNN) which are able to cope with complex environments and broad communication that single recurrent net cannot cope with. We conduct two experiments in order to evaluate the effectiveness of MRNN to a low level communication task such as nonverbal interaction. First, we examine the effect of the number of RNNs contained in MRNN. Second, we examine the effect of the context dependency of MRNN. These experiments show the capability of MRNN as a new-type controller of communication robot. Yoshihiro Sakamoto, Tetsuya Ogata, Shigeki Sugano |
IROS | 2 |
| 2004 | Acquisition of reactive motion for communication robots using interactive ECabstractWe developed an emotional communication robot, WAMOEBA, using behavior-based techniques. We also proposed motor-agent (MA) model, which is an autonomous distributed-control algorithm constructed of simple sensor motor coordination. Though it enables WAMOEBA to behave in various ways, the weight of the combinations between different motor agents is influenced by the preferences of the developer. We usually use machine-learning algorithms to automatically configure these parameters for communication robots. However, this makes it difficult to define the quantitative evaluation required for communication. We therefore used the method of interactive evolutionary computation (IEC), which can be applied to problems involving quantitative evaluation. IEC does not require to define a fitness function; this task is performed by users. But the biggest problem with using IEC is human fatigue, which causes insufficiency of individuals and generations for convergence of EC. To fix this problem, we use the prediction function that automatically calculates the fitness values of genes from some samples that have received the human subjective evaluation. Then, we carried out the behavior acquisition experiment using the IEC simulation system with the prediction function. As the results of experiments, it is confirmed that diversifying the genetic pool is an efficient way for generating a variety of behavior. Yuki Suga, Tetsuya Ogata, Shigeki Sugano |
IROS | 2 |
| 2004 | Automatic Sound-Imitation Word Recognition from Environmental Sounds Focusing on Ambiguity Problem in Determining Phonemes
Kazushi Ishihara, Tomohiro Nakatani, Tetsuya Ogata, Hiroshi G. Okuno |
PRICAI | 3 |
| 2003 | Robust modeling of dynamic environment based on robot embodimentabstractRecent studies on embodied cognitive science have shown us the possibility of emergence of more complex and nontrivial behaviors with quite simple designs if the designer takes the dynamics of the system-environment interaction into account properly. In this paper, we report our tentative classification experiments of several objects using the human-like autonomous robot, "WAMOEBA-2Ri". As modeling the environment, we focus on not only static aspects of the environment but also dynamic aspects of it including that of the system own. The visualized result of this experiment shows the integration of multimodal sensor dataset acquired by the system-environment interaction ("grasping") enable robust categorization of several objects. Finally, in discussion, we demonstrate a possible application to making "invariance in motion" emerge consequently by extending this approach. Kuniaki Noda, Mototaka Suzuki, Naofumi Tsuchiya, Yuki Suga, Tetsuya Ogata, Shigeki Sugano |
ICRA | 5 |
| 2003 | Interactive learning in human-robot collaborationabstractIn this paper, we investigated interactive learning between human subjects and robot experimentally, and its essential characteristics are examined using the dynamical systems approach. Our research concentrated on the navigation system of a specially developed humanoid robot called Robovie and seven human subjects whose eyes were covered, making them dependent on the robot for directions. We compared the usual feed-forward neural network (FFNN) without recursive connections and the recurrent neural network (RNN). Although the performances obtained with both the RNN and the FFNN improved in the early stages of learning, as the subject changed the operation by learning on its own, all performances gradually became unstable and failed. Results of a questionnaire given to the subjects confirmed that the FFNN gives better mental impressions, especially from the aspect of operability. When the robot used a consolidation-learning algorithm using the rehearsal outputs of the RNN, the performance improved even when interactive learning continued for a long time. The questionnaire results then also confirmed that the subject's mental impressions of the RNN improved significantly. The dynamical systems analysis of RNNs supports these differences. Tetsuya Ogata, Noritaka Masago, Shigeki Sugano, Jun Tani |
IROS | 1 |
| 2002 | Development of Passive Elements with Variable Mechanical Impedance for Wearable RobotsabstractA new type of passive element with variable mechanical impedance is proposed. Since the proposed elements are lighter, smaller and softer than the previous passive elements, these elements have the possibility to develop wearable robots. In the paper, the principles of the variable mechanical impedance are explained and the basic experimental results demonstrate the performance of the proposed element. Moreover, a wearable robot for virtual reality is developed by using the proposed element and it is shown that the developed wearable robot is utilized as a force display in the virtual world. Sadao Kawamura, D. Ishida, Tetsuya Ogata, Y. Nakayama, Osamu Tabata, Susumu Sugiyama |
ICRA | 4 |
| 2001 | Motion generation of the autonomous robot based on body structureabstractAims to investigate the intelligence which can make robots adapt to the human environment. The paper points out the problems of the behavior-based robot, and proposes the methods which can generate whole body motions based on body structure and integrates the reflection motions to make the behaviors continuous. The motion performances are compared in two kinds of environments, such as a dynamic environment and a static environment, by using a simulator of the autonomous robot WAMOEBA-2Ri developed in this research. Finally, we show that the integration parameters of the proposed method reflect the body structure of the robot and environmental structures. Tetsuya Ogata, Takaaki Komiya, Shigeki Sugano |
IROS | 1 |
| 2000 | A Robotic Co-Operation System Based on a Self-Organization Approached Human Work ModelabstractThis study presents a method of human cooperating systems, which can determine when support behavior is necessary by a human work model. We focus on assembly work as the target and propose a self-organizing approach of human work models by sampled human information by vision sensors. The support is determined according to the work model. Such a system would realize provision of support without a strict model of the assembly target and enable support in cases where neither the assembly process, nor the final form of the completed task is known to the system in advance. First, a method of measuring human information and extracting states where support is necessary, from the human work model is presented. Next a support system for assembly work cooperation according to the work model, with physical interaction capabilities is described. Experiments were carried out to evaluate and verify the effectiveness of the system. The results show that the constructed assembly support system is effective in both improving performance and increasing friendliness. Yasuhisa Hayakawa, Tetsuya Ogata, Shigeki Sugano |
ICRA | 2 |
| 2000 | Development of emotional communication robot: WAMOEBA-2R-experimental evaluation of the emotional communication between robots and humansabstractThis paper aims to clarify the cooperation intelligence of robots. This paper describes the autonomous robot named WAMOEBA-2R which can communicate with humans by both an informational and physical way. WAMOEBA-2R has two arms of which each joint has a torque sensor to realize the physical interaction with humans, the function of the voice recognition and the face recognition. The arms are controlled by a distributed agent network system. The network architecture is acquired in a neural network by the feedback-error-learning algorithm. We surveyed 150 visitors at the '99 International Robot Exhibition held in Tokyo (Oct. 1999) to evaluate their psychological impressions of WAMOEBA-2R. As a result, some factors of the human-robot emotional communication were discovered. Tetsuya Ogata, Yoshihiro Matsuyama, Takaaki Komiya, Masataka Ida, Kuniaki Noda, Shigeki Sugano |
IROS | 1 |
| 2000 | A violin playing algorithm considering the change of phrase impressionabstractThe study focuses on the dynamics of KANSEI information and aims to propose an algorithm of motion planning using KANSEI. Concretely, the violin playing is regarded as the target motion which will be greatly influenced by KANSEI. The study introduces a multi-agent algorithm in which four physical bowing parameters are agents to adapt the impression transition smoothly, while maintaining the relationships between the parameters. We realized the violin performance suitable for the timbre words by introducing the proposed agent algorithm into the bowing machine developed in this research. As a result of the experiments, it was confirmed that there were various playing performances according to a single impression transition. Tetsuya Ogata, Akitoshi Shimura, Koji Shibuya, Shigeki Sugano |
SMC | 1 |
| 1999 | Emotional Communication Between Humans and the Autonomous Robot Which Has the Emotion ModelabstractDiscusses the communication between autonomous robots and humans through the development of a robot (WAMOEBA-2) which has an emotion model. The model refers to the internal secretion system of humans and it has four kinds of the hormone parameters to use to adjust various internal conditions such as motor output, cooling fan output and sensor gain. We surveyed 126 visitors at '97 International Robot Exhibition held in Tokyo, Japan (Oct. 1997) in order to evaluate psychological impressions of the robot. As a result, the human friendliness of the robot was confirmed and some factors of the human-robot emotional communication were discovered. Tetsuya Ogata, Shigeki Sugano |
ICRA | 1 |
| 1999 | Emotional communication between humans and robots - consideration of primitive language in robotsabstractThis research aims to clarify the behavior intelligence and the human cooperation intelligence of robots by the emotion models which is based on the robot's hardware structure. In this paper human's mental images and language are given consideration as a method for emotional expression. The hypothesis model for the acquisition of the internal expressions of robots and the experimental results using a real autonomous robot are described. Tetsuya Ogata, Shigeki Sugano |
IROS | 1 |
| 1998 | Communication between behavior-based robots with emotion model and humansabstractThis study discusses the communication between autonomous robots and humans through the development of a robot which has an emotion model. The model refers to the internal secretion system of humans, and it has four kinds of hormone parameters to be used for adjusting various internal conditions such as motor output cooling fan output and sensor gain. We surveyed 126 visitors at '97 International Robot Exhibition held in Tokyo, Japan, in order to evaluate psychological impressions of the robot. As a result, the human friendliness of the robot was confirmed and some factors of the human-robot emotional communication were discovered. Tetsuya Ogata, Shigeki Sugano |
SMC | 1 |
| 1997 | Generation of behavior automaton on neural networkabstractTo plan behavior procedures, it is necessary for an agent to have a world model concerning the temporal sequences information. In this paper, a temporal information learning algorithm is proposed with a three layer neural network implementing the "effectiveness of simulation accumulation" algorithm. This algorithm can construct a "behavior automaton" in the neural network. From the results of some learning experiments using a mobile robot simulation, the generated automaton expresses the complexity of the simulation environments. The robot agent acquires a behavior automaton for obstacle avoidance behavior which is influenced by the simulation environment. Tetsuya Ogata, Kazuki Hayashi, Ikuo Kitagishi, Shigeki Sugano |
IROS | 1 |
| 1996 | Emergence of mind in robots for human interface - research methodology and robot modelabstractThe objective of this work is to develop the technology for human-machine communication through the research of the emergence of mind in mechanical systems. In this paper, the hypothesis about the emergence of mind is proposed. First, a system chart expressing the human brain information processing and the development of an autonomous mobile robot "WAMOEBA-IR" (Waseda artificial mind on emotion base) are described. The conception of the WAMOEBA-IR design is that robots should have a self-presentation evaluation function. Further more, the method to evaluate the whole system is described from the viewpoint of the animal psychology. As a result of the experiments, WAMOEBA-IR showed specific emotional reactions with color appearances to some situations. WAMOEBA-IR has the sense of values about colors and sounds based on self-preservation as the first step of the emergence of mind. Shigeki Sugano, Tetsuya Ogata |
ICRA | 2 |