Takamitsu Matsubara

dblp:97/5838 · DBLP profile ↗
← Back
53ranked-venue papers
12as first author
18since 2021 · last 2025
0000-0003-3545-4814ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 52 · 12 first-author · 17 since 2021Systems, architecture and hardware · 30 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2
YearPublicationVenuePosition
2025 Feasibility-Aware Imitation Learning from Observations Through a Hand-Mounted Demonstration Interface
abstract
Imitation learning through a demonstration interface is expected to learn policies for robot automation from intuitive human demonstrations. However, due to the differences in human and robot movement characteristics, a human expert might unintentionally demonstrate an action that the robot cannot execute. We propose feasibility-aware behavior cloning from observation (FABCO). In the FABCO framework, the feasibility of each demonstration is assessed using the robot's pre-trained forward and inverse dynamics models. This feasibility information is provided as visual feedback to the demonstrators, encouraging them to refine their demonstrations. During policy learning, estimated feasibility serves as a weight for the demonstration data, improving both the data efficiency and the robustness of the learned policy. We experimentally validated FABCO's effectiveness by applying it to a pipette insertion task involving a pipette and a vial. Four participants assessed the impact of the feasibility feedback and the weighted policy learning in FABCO. Additionally, we used the NASA Task Load Index (NASA-TLX) to evaluate the workload induced by demonstrations with visual feedback.
Kei Takahashi, Hikaru Sasaki, Takamitsu Matsubara
ICRA3
2025 ICCO: Learning an Instruction-conditioned Coordinator for Language-guided Task-aligned Multi-robot Control
abstract
Recent advances in Large Language Models (LLMs) have permitted the development of language-guided multi-robot systems, which allow robots to execute tasks based on natural language instructions. However, achieving effective coordination in distributed multi-agent environments remains challenging due to (1) misalignment between instructions and task requirements and (2) inconsistency in robot behaviors when they independently interpret ambiguous instructions. To address these challenges, we propose Instruction-Conditioned Coordinator (ICCO), a Multi-Agent Reinforcement Learning (MARL) framework designed to enhance coordination in language-guided multi-robot systems. ICCO consists of a Coordinator agent and multiple Local Agents, where the Coordinator generates Task-Aligned and Consistent Instructions (TACI) by integrating language instructions with environmental states, ensuring task alignment and behavioral consistency. The Coordinator and Local Agents are jointly trained to optimize a reward function that balances task efficiency and instruction following. A Consistency Enhancement Term is added to the learning objective to maximize mutual information between instructions and robot behaviors, further improving coordination. Simulation and real-world experiments validate the effectiveness of ICCO in achieving language-guided task-aligned multi-robot control. The demonstration can be found at https://yanoyoshiki.github.io/ICCO/.
Yoshiki Yano, Kazuki Shibata, Maarten Kokshoorn, Takamitsu Matsubara
IROS4
2025 ASBI: Leveraging informative real-world data for active black-box simulator tuning
Gahee Kim, Takamitsu Matsubara
Appl. Intell.2
2025 Composite Gaussian processes flows for learning discontinuous multimodal policies
Shu-yuan Wang, Hikaru Sasaki, Takamitsu Matsubara
Appl. Intell.3
2025 Progressive-Resolution Policy Distillation: Leveraging Coarse-Resolution Simulations for Time-Efficient Fine-Resolution Policy Learning
abstract
In earthwork and construction, excavators often encounter large rocks mixed with various soil conditions, requiring skilled operators. This paper presents a framework for achieving autonomous excavation using reinforcement learning (RL) through a rock excavation simulator. In the simulation, resolution can be defined by the particle size/number in the whole soil space. Fine-resolution simulations closely mimic real-world behavior but demand significant calculation time and challenging sample collection, while coarse-resolution simulations enable faster sample collection but deviate from real-world behavior. To combine the advantages of both resolutions, we explore using policies developed in coarse-resolution simulations for pre-training in fine-resolution simulations. To this end, we propose a novel policy learning framework called Progressive-Resolution Policy Distillation (PRPD), which progressively transfers policies through some middle-resolution simulations with conservative policy transfer to avoid domain gaps that could lead to policy transfer failure. Validation in a rock excavation simulator and nine real-world rock environments demonstrated that PRPD reduced sampling time to less than 1/7 while maintaining task success rates comparable to those achieved through policy learning in a fine-resolution simulation. Note to Practitioners—This paper is motivated by the issue of computation time in excavation simulation using soil particles. The behavior of real soil is highly complex, and approximating it at high resolution requires enormous computational costs. Therefore, existing soil simulators have focused on improving simulation accuracy while maintaining reduced computation time. This paper takes a different approach by focusing on the learning of control policies in excavation simulators and proposes a framework for reducing calculation time in such use cases. In this framework, a control policy is first learned in a low-resolution simulation, significantly reducing computation time. The learned policy is then transferred to a high-resolution simulation for retraining, thereby achieving an overall reduction in simulation time. Furthermore, to enable robust policy transfer across different resolutions, this paper discusses a stable policy distillation scheme and insights into resolution design. This approach enables the development of autonomous excavation systems without relying on expensive real-world data collection, improving the scalability and adaptability of autonomous excavation. Simulation experiments suggest that this framework significantly reduces training time compared to conventional policy learning approaches. However, real-world validation has so far been limited to simple excavation robots. Future research will explore applications to excavators and other machinery more suitable for real-world operations. Although this paper focuses on autonomous excavation, the proposed approach can also be extended to environments where increased simulation resolution critically impacts computation time, such as liquid and soft object manipulation.
Yuki Kadokawa, Hirotaka Tahara, Takamitsu Matsubara
IEEE Trans Autom. Sci. Eng.3
2024 Domain Randomization-free Sim-to-Real : An Attention-Augmented Memory Approach for Robotic Tasks
abstract
The sim-to-real gap, a long-standing challenge in the field of robotics, has garnered significant attention. Essentially, it is important to learn robust representation models that can be seamlessly applied in both simulation and real world. Traditional approaches like domain randomization have demonstrated success in zero-short setting, by creating representations that are resilient and adaptable through the augmentation of diversity within simulations. However, they suffer from the need for extensive training across a range of parameter variances, and dependency on heuristic approaches. In this work, we present a novel reinforcement learning architecture named Soft Attention-Augmented Actor-Critic (Soft3AC) for sim-to-real robotic tasks without the need for heuristic domain randomization. Our approach achieves the learning of semantically task-relevant feature representations that exhibit resilience against appearance gaps. This is realized by employing an architectural design that separates current perceptions from historical perceptions in memory, fostering abstract spatial-temporal understanding. Simultaneously, the introduction of an attention mechanism enables a more contextual processing. We validated our method through conducting a valve rotation task with a robotic hand, under both sim-to-sim and sim-to-real conditions. The results indicate that our model adeptly bridges the appearance gap observed in sim-to-sim and sim-to-real transfers. Our method demonstrated its ability to be deployed directly into the real world in a domain randomization free zero-shot manner.
Shun Otsubo, Tomoya Yamanokuchi, Takamitsu Matsubara, Shotaro Miwa
IROS4
2023 Domains as Objectives: Multi-Domain Reinforcement Learning with Convex-Coverage Set Learning for Domain Uncertainty Awareness
abstract
Domain randomization (DR) is a powerful framework that has allowed the transfer of policies from randomized domain (a.k.a. simulation) to real robots with little to no retraining requirement. However, because the policy has to perform well for many different domain conditions, DR tends to produce sub-optimal policies that can be too conservative on the target real system. This problem is further exacerbated the larger the randomized domain is. To tackle this issue, recent works have proposed to learn universal policies (UP) with domain knowledge such that they can adapt their behavior to each domain when paired with an online system identifier (OSI). However, in most applications, perfect identifications of the target domain can be impossible. In this paper, by drawing similarities between DR as a multi-domain reinforcement learning and multi-objective reinforcement learning (MORL), we propose to learn a UP over the convex coverage set borrowed from the MORL theory. Thanks to this, our method learns a UP that effectively captures different sub-domains of the uncertainty set and can therefore adapt its behavior based on an OSI uncertainty, unlocking the power of stochastic system identification with no retraining requirement. This pseudo-MORL framework also contains previous works in DR and robust reinforcement learning. We conduct simulations on Mujoco tasks and experiments on a real D'Claw robot, revealing the effectiveness of our domain-uncertainty-aware UP for sim-to-real transfer.
Wendyam Eric Lionel Ilboudo, Taisuke Kobayashi, Takamitsu Matsubara
IROS3
2023 AdaTerm: Adaptive T-distribution estimated robust moments for Noise-Robust stochastic gradient optimization
Wendyam Eric Lionel Ilboudo, Taisuke Kobayashi, Takamitsu Matsubara
Neurocomputing3
2023 Cautious policy programming: exploiting KL regularization for monotonic policy improvement in reinforcement learning
abstract
Abstract In this paper, we propose cautious policy programming (CPP), a novel value-based reinforcement learning (RL) algorithm that exploits the idea of monotonic policy improvement during learning. Based on the nature of entropy-regularized RL, we derive a new entropy-regularization-aware lower bound of policy improvement that depends on the expected policy advantage function but not on state-action-space-wise maximization as in prior work. CPP leverages this lower bound as a criterion for adjusting the degree of a policy update for alleviating policy oscillation. Different from similar algorithms that are mostly theory-oriented, we also propose a novel interpolation scheme that makes CPP better scale in high dimensional control problems. We demonstrate that the proposed algorithm can trade off performance and stability in both didactic classic control problems and challenging high-dimensional Atari games.
Lingwei Zhu, Takamitsu Matsubara
Mach. Learn.2
2023 Bayesian Disturbance Injection: Robust imitation learning of flexible policies for robot manipulation
Hanbit Oh, Hikaru Sasaki, Brendan Michael, Takamitsu Matsubara
Neural Networks4
2022 Gaussian Process Self-triggered Policy Search in Weakly Observable Environments
abstract
The environments of such large industrial machines as waste cranes in waste incineration plants are often weakly observable, where little information about the environ-mental state is contained in the observations due to technical difficulty or maintenance cost (e.g., no sensors for observing the state of the garbage to be handled). Based on the findings that skilled operators in such environments choose predetermined control strategies (e.g., grasping and scattering) and their durations based on sensor values, we propose a novel non-parametric policy search algorithm: Gaussian process self-triggered policy search (GPSTPS). GPSTPS has two types of control policies: action and duration. A gating mechanism either maintains the action selected by the action policy for the duration specified by the duration policy or updates the action and duration by passing new observations to the policy; therefore, it is categorized as self-triggered. GPSTPS simultaneously learns both policies by trial and error based on sparse GP priors and variational learning to maximize the return. To verify the performance of our proposed method, we conducted experiments on garbage-grasping-scattering task for a waste crane with weak observations using a simulation and a robotic waste crane system. As experimental results, the proposed method acquired suitable policies to determine the action and duration based on the garbage's characteristics.
Hikaru Sasaki, Terushi Hirabayashi, Kaoru Kawabata, Takamitsu Matsubara
ICRA4
2022 Disturbance-injected Robust Imitation Learning with Task Achievement
abstract
Robust imitation learning using disturbance injections overcomes issues of limited variation in demonstrations. However, these methods assume demonstrations are optimal, and that policy stabilization can be learned via simple augmentations. In real-world scenarios, demonstrations are often of diverse-quality, and disturbance injection instead learns sub-optimal policies that fail to replicate desired behavior. To address this issue, this paper proposes a novel imitation learning framework that combines both policy robustification and optimal demonstration learning. Specifically, this combinatorial approach forces policy learning and disturbance injection optimization to focus on mainly learning from high task achievement demonstrations, while utilizing low achievement ones to decrease the number of samples needed. The effectiveness of the proposed method is verified through experiments using an excavation task in both simulations and a real robot, resulting in high-achieving policies that are more stable and robust to diverse-quality demonstrations. In addition, this method utilizes all of the weighted sub-optimal demonstrations without eliminating them, resulting in practical data efficiency benefits.
Hirotaka Tahara, Hikaru Sasaki, Hanbit Oh, Brendan Michael, Takamitsu Matsubara
ICRA5
2021 Geometric Value Iteration: Dynamic Error-Aware KL Regularization for Reinforcement Learning
abstract
The recent boom in the literature on entropy-regularized reinforcement learning (RL) approaches reveals that Kullback-Leibler (KL) regularization brings advantages to RL algorithms by canceling out errors under mild assumptions. However, existing analyses focus on fixed regularization with a constant weighting coefficient and do not consider cases where the coefficient is allowed to change dynamically. In this paper, we study the dynamic coefficient scheme and present the first asymptotic error bound. Based on the dynamic coefficient error bound, we propose an effective scheme to tune the coefficient according to the magnitude of error in favor of more robust learning. Complementing this development, we propose a novel algorithm, Geometric Value Iteration (GVI), that features a dynamic error-aware KL coefficient design with the aim of mitigating the impact of errors on performance. Our experiments demonstrate that GVI can effectively exploit the trade-off between learning speed and robustness over uniform averaging of a constant KL coefficient. The combination of GVI and deep networks shows stable learning behavior even in the absence of a target network, where algorithms with a constant KL coefficient would greatly oscillate or even fail to converge.
Toshinori Kitamura, Lingwei Zhu, Takamitsu Matsubara
ACML3
2021 Cautious Actor-Critic
abstract
The oscillating performance of off-policy learning and persisting errors in the actor-critic(AC) setting call for algorithms that can conservatively learn to suit the stability-critical applications better. In this paper, we propose a novel off-policy AC algorithm cautious actor-critic (CAC). The name cautious comes from the doubly conservative nature that we exploit the classic policy interpolation from conservative policy iteration for the actor and the entropy-regularization of conservative value iteration for the critic. Our key observation is the entropy-regularized critic facilitates and simplifies the unwieldy interpolated actor update while still ensuring robust policy improvement. We compare CAC to state-of-the-art AC methods on a set of challenging continuous control problems and demonstrate thatCAC achieves comparable performance while significantly stabilizes learning.
Lingwei Zhu, Toshinori Kitamura, Takamitsu Matsubara
ACML3
2021 Bayesian Disturbance Injection: Robust Imitation Learning of Flexible Policies
abstract
Scenarios requiring humans to choose from multiple seemingly optimal actions are commonplace, however standard imitation learning often fails to capture this behavior. Instead, an over-reliance on replicating expert actions induces inflexible and unstable policies, leading to poor generalizability in an application. To address the problem, this paper presents the first imitation learning framework that incorporates Bayesian variational inference for learning flexible nonparametric multi-action policies, while simultaneously robustifying the policies against sources of error, by introducing and optimizing disturbances to create a richer demonstration dataset. This combinatorial approach forces the policy to adapt to challenging situations, enabling stable multi-action policies to be learned efficiently. The effectiveness of our proposed method is evaluated through simulations and real-robot experiments for a table-sweep task using the UR3 6-DOF robotic arm. Results show that, through improved flexibility and robustness, the learning performance and control safety are better than comparison methods.
Hanbit Oh, Hikaru Sasaki, Brendan Michael, Takamitsu Matsubara
ICRA4
2021 Deep reinforcement learning of event-triggered communication and control for multi-agent cooperative transport
abstract
In this paper, we explore a multi-agent reinforcement learning approach to address the design problem of communication and control strategies for multi-agent cooperative transport. Typical end-to-end deep neural network policies may be insufficient for covering communication and control; these methods cannot decide the timing of communication and can only work with fixed-rate communications. Therefore, our framework exploits event-triggered architecture, namely, a feedback controller that computes the communication input and a triggering mechanism that determines when the input has to be updated again. Such event-triggered control policies are efficiently optimized using a multi-agent deep deterministic policy gradient. We confirmed that our approach could balance the transport performance and communication savings through numerical simulations.
Kazuki Shibata, Tomohiko Jimbo, Takamitsu Matsubara
ICRA3
2021 Learning Robotic Contact Juggling
abstract
Robotic contact juggling is a challenging task in which robots must control the movement of a ball rapidly and indirectly without holding it while keeping the ball in and sometimes out of contact with the robot’s body. In this work, we address the problem of learning such robotic contact juggling from trial and error via model-based reinforcement learning (MBRL). The key insight is that complex robot-ball interactions of the contact juggling actually consist of a small set of simple dynamics that each corresponds to a distinct interaction "primitive" such as touching and releasing the ball. Accordingly, we develop a tailored MBRL method that incrementally fits a set of simple dynamics models to the movements of a robot and a ball while also learning a switching model that can select a proper dynamics model depending on the current state and action. The learned model can then be used in an MBRL framework to seek optimal juggling control. We demonstrated the effectiveness of our approach on a simulator of contact juggling performed by a robotic arm.
Kazutoshi Tanaka, Masashi Hamaya, Devwrat Joshi, Felix von Drigalski, Ryo Yonetani, Takamitsu Matsubara, Yoshihisa Ijiri
IROS6
2021 Variational policy search using sparse Gaussian process priors for learning multimodal optimal actions
Hikaru Sasaki, Takamitsu Matsubara
Neural Networks2
2020 Contact-based in-hand pose estimation using Bayesian state estimation and particle filtering
abstract
In industrial assembly tasks, the position of an object grasped by the robot has to be known with high precision in order to insert or place it. In real applications, this problem is commonly solved by jigs that are specially produced for each part. However, they significantly limit flexibility and are prohibitive when the target parts change often, so a flexible method to localize parts with high accuracy after grasping is desired. To solve this problem, we propose a method that can estimate the position of an object in the robot's hand to sub-millimeter precision, and can improve its estimate incrementally, using only minimal calibration and a force sensor. Our method is applicable to any robotic gripper and any rigid object that the gripper can hold, and requires only a force sensor. We demonstrate that the method can determine the position of an object to a precision of under 1 mm without using any part-specific jigs or equipment.
Felix von Drigalski, Shohei Taniguchi, Robert Lee, Takamitsu Matsubara, Masashi Hamaya, Kazutoshi Tanaka, Yoshihisa Ijiri
ICRA4
2020 Sample-and-computation-efficient Probabilistic Model Predictive Control with Random Features
abstract
Gaussian processes (GPs) based Reinforcement Learning (RL) methods with Model Predictive Control (MPC) have demonstrated their excellent sample efficiency. However, since the computational cost of GPs largely depends on the training sample size, learning an accurate dynamics using GPs result in low control frequency in MPC. To alleviate this trade-off and achieve a sample-and-computation-efficient nature, we propose a novel model-based RL method with MPC. Our approach employs a linear Gaussian model with randomized features using the Fastfood as an approximated GP dynamics. Then, we derive an analytic moment-matching scheme in state prediction with the model and uncertain inputs. As a result, the computational cost of the MPC in our RL method does not depend on the training sample size and can improve the control frequency over previous methods. Through experiments with simulated and real robot control tasks, the sample efficiency, as well as the computation efficiency of our model-based RL method, are demonstrated.
Cheng-Yu Kuo, Yunduan Cui, Takamitsu Matsubara
ICRA3
2020 Dynamic Actor-Advisor Programming for Scalable Safe Reinforcement Learning
abstract
Real-world robots have complex strict constraints. Therefore, safe reinforcement learning algorithms that can simultaneously minimize the total cost and the risk of constraint violation are crucial. However, almost no algorithms exist that can scale to high-dimensional systems to the best of our knowledge. In this paper, we propose Dynamic Actor-Advisor Programming (DAAP), as an algorithm for sample-efficient and scalable safe reinforcement learning. DAAP employs two control policies, actor and advisor. They are updated to minimize total cost and risk of constraint violation intertwiningly and smoothly towards each other's direction by using the other as the baseline policy in the Kullback-Leibler divergence of Dynamic Policy Programming framework. We demonstrate the scalability and sample efficiency of DAAP through its application on simulated robot arm control tasks with performance comparisons to baselines.
Lingwei Zhu, Yunduan Cui, Takamitsu Matsubara
ICRA3
2020 Learning Soft Robotic Assembly Strategies from Successful and Failed Demonstrations
abstract
Physically soft robots are promising for robotic assembly tasks as they allow stable contacts with the environment. In this study, we propose a novel learning system for soft robotic assembly strategies. We formulate this problem as a reinforcement learning task and design the reward function from human demonstrations. Our key insight is that the failed demonstrations can be used as constraints to avoid failed behaviors. To this end, we developed a teaching device with which humans can intuitively provide various demonstrations. Moreover, we leverage Physically-Consistent Gaussian Mixture Models to clearly assign Gaussian components to the successful and failed trials. We then create the reference trajectories via Gaussian Mixture Regressions, which fit the successful demonstrations while considering the failed ones. Finally, we apply a sample- efficient deep model-based reinforcement learning method to obtain robust strategies with a few interactions. To validate our method, we developed a real-robot experimental system composed of a rigid collaborative robot arm with a compliant wrist and the teaching device. Our results demonstrated that our method learned the assembly strategies with a higher success rate than when using only successful demonstrations.
Masashi Hamaya, Felix von Drigalski, Takamitsu Matsubara, Kazutoshi Tanaka, Robert Lee, Chisato Nakashima, Yoshiya Shibata, Yoshihisa Ijiri
IROS3
2020 Probabilistic active filtering with gaussian processes for occluded object search in clutter
Yunduan Cui, Junichiro Ooga, Akihito Ogawa, Takamitsu Matsubara
Appl. Intell.4
2019 Exploiting Human and Robot Muscle Synergies for Human-in-the-loop Optimization of EMG-based Assistive Strategies
abstract
In this study, we propose a novel human-in-the-loop optimization approach for exoskeleton robot control. We develop a method to optimize widely-used Electromyography (EMG)-based assistive strategies. If we use multiple EMG channels to control multi-DoF robots, optimization process becomes complex and requires a large amount of data. To make the optimization tractable, we exploit the synergies both of the human muscles and artificial muscles of the exoskeleton robots to reduce the number of parameters of the assistive strategies. We show that we can extract the synergies not only from the user's muscle activities but from pneumatic artificial muscle (PAMs) contractions of the exoskeleton robot. Then, we adopt a Bayesian optimization method to acquire the parameters for assisting human movements by iteratively identifying the user's preferences of the assistive strategies. We conducted experiments to evaluate our proposed method with a PAMs-driven upper-limb exoskeleton robot. Our method successfully learned assistive strategies from the human-in-theloop optimization with a practicable number of interactions.
Masashi Hamaya, Takamitsu Matsubara, Jun-ichiro Furukawa, Satoshi Yagi, Tatsuya Teramae, Tomoyuki Noda, Jun Morimoto
ICRA2
2019 Probabilistic Active Filtering for Object Search in Clutter
abstract
This paper proposes a probabilistic approach for object search in clutter. Due to heavy occlusions, it is vital for an agent to be able to gradually reduce uncertainty in observations of the objects in its workspace by systematically rearranging them. Probabilistic methodologies present a promising sample-efficient alternative to handle the massively complex state-action space that inherently comes with this problem, avoiding the need for both exhaustive training samples and the accompanying heuristics for traversing a large-scale model during runtime. We approach the object search problem by extending a Gaussian Process active filtering strategy with an additional model for capturing state dynamics as the objects are moved over the course of the activity. This allows viable models to be built upon relatively scarce training data, while the complexity of the action space is also reduced by shifting objects over relatively short distances. Validation in both simulation and with a real Baxter robot with a limited number of training samples demonstrates the efficacy of the proposed approach.
James Poon, Yunduan Cui, Junichiro Ooga, Akihito Ogawa, Takamitsu Matsubara
ICRA5
2019 Multimodal Policy Search using Overlapping Mixtures of Sparse Gaussian Process Prior
abstract
In this paper, we present a novel policy search reinforcement learning algorithm that can deal with multimodality in control policies based on Gaussian processes. Our approach employs Overlapping Mixtures of Gaussian Processes (OMGPs) for a control policy, in which all the GPs in the mixture are global and overlapped in the input space. We first extend the OMGPs by combing sparse pseudo-input GPs as OMSGPs to reduce its computational cost of learning and prediction suitable for policy search. Then, we derive a novel multimodal policy search algorithm based on variational Bayesian inference by placing the OMSGPs as the prior of the multimodal control policy. To validate the effectiveness of our algorithm, we applied it to two typical robotic tasks in simulation: 1) object grasping and 2) table-sweep tasks since they both require the multimodality in the optimal policies. Simulation results demonstrate that our algorithm can efficiently learn multimodal policies even with high dimensional observations.
Hikaru Sasaki, Takamitsu Matsubara
ICRA2
2019 Reinforcement Learning Boat Autopilot: A Sample-efficient and Model Predictive Control based Approach
abstract
In this research we focus on developing a reinforcement learning system for a challenging task: autonomous control of a real-sized boat, with difficulties arising from large uncertainties in the challenging ocean environment and the extremely high cost of exploring and sampling with a real boat. To this end, we explore a novel Gaussian processes (GP) based reinforcement learning approach that combines sample-efficient model-based reinforcement learning and model predictive control (MPC). Our approach, sample-efficient probabilistic model predictive control (SPMPC), iteratively learns a Gaussian process dynamics model and uses it to efficiently update control signals within the MPC closed control loop. A system using SPMPC is built to efficiently learn an autopilot task. After investigating its performance in a simulation modeled upon real boat driving data, the proposed system successfully learns to drive a real-sized boat equipped with a single engine and sensors measuring GPS, speed, direction, and wind in an autopilot task without human demonstration.
Yunduan Cui, Shigeki Osaki, Takamitsu Matsubara
IROS3
2018 Learning Mobility Aid Assistance via Decoupled Observation Models
abstract
This paper presents an active assistance framework for mobility systems, such as Power Mobility Devices (PMD), with the distinctive goal of being able to operate within a local moving window, as opposed to the common reliance upon persistent global environments and objectives. Demonstration data from able experts driving a simulated mobility aid in a representative indoor setting is used off-line to build behavioral models of navigation postulated separately upon user joystick inputs and on-board sensor data. These models are built respectively via Gaussian Processes for the joystick signals, and a Deep Convolutional Neural Network for the sensor data; in this case a planar LIDAR. Their combined outputs form a continuous distribution of estimated traversal likelihood within the user's immediate space, allowing for real-time stochastic optimal path planning to guide a user to its intended local destination. Moreover, the computational efficiency of the decoupled models permits rapid replanning on-the-fly for a smooth assistive action. On-line and off-line evaluations substantiate the advantages of the framework in generalising intelligent navigational assistance, of particular relevance for users who experience difficulty in safe mobility.
James Poon, Yunduan Cui, Jaime Valls Miró, Takamitsu Matsubara
ICARCV4
2017 Learning task-parametrized assistive strategies for exoskeleton robots by multi-task reinforcement learning
abstract
Recent studies suggest that reinforcement learning has great potential for generating assistive strategies in exoskeletons through physical interactions between a user and a robot. Previous methods focused on a task-specific assistive strategy, where for every single task (situation/context), the user needs to interact with a robot to learn an appropriate assistive strategy. Therefore, the learned strategies cannot be generalized for a new task. Since the sampling cost is expensive for such human-in-the-loop systems as exoskeletons, generalization must be enabled. In this paper, we propose to learn task-parametrized assistive strategies for exoskeleton robots. Our method employs an assistive strategy, which depends on the task parameter and the state variable, that can be learned from multiple sets of human-robot interaction data across different tasks and generalized even for an unseen task, given the task parameter without additional learning. To alleviate the user's burden in the learning process across multiple tasks, we exploit a data-efficient multi-task reinforcement learning framework. To verify the effectiveness of our method, we developed an experimental platform with an exoskeleton robot. We conducted a series of experiments whose experimental results show that our method can learn such a task-parametrized assistive strategy and be generalized for unseen tasks to reduce the user's electromyography signals (EMGs) during tasks.
Masashi Hamaya, Takamitsu Matsubara, Tomoyuki Noda, Tatsuya Teramae, Jun Morimoto
ICRA2
2017 Local driving assistance from demonstration for mobility aids
abstract
Active assistive mobility systems are largely limited to a-priori mapped environments, whereas their reactive assistive counterparts are in general location independent and focus on the provision of collision avoidance in the immediate space surrounding the platform. This paper presents a framework capable of providing active short-term navigation, combining the intelligence of active assistance with the freedom of location independence. Demonstration data from an able expert while driving the mobility aid in a standard indoor setting is used off-line to learn reference behavioral models of navigation given perceptual information from the platform surroundings and the input controls exerted by the user while navigating. These serve as the foundation for on-line probabilistic short-term destination inference using the instantaneously available data from the user and on-board sensors. This is coupled with a real-time stochastic optimal path generation able to exploit the same short term demonstration paths from the expert with the belief they capture both the driver's awareness of the platform's physical geometry and appropriate behaviors for their surroundings. Experimental results with users of varying proficiency in a setting unvisited in training data show promise in using the framework in assisting users experiencing difficulty in safe power mobility aid use.
James Poon, Yunduan Cui, Jaime Valls Miró, Takamitsu Matsubara, Kenji Sugimoto
ICRA4
2017 User-robot collaborative excitation for PAM model identification in exoskeleton robots
abstract
Pneumatic Artificial Muscle (PAM) actuators have been used as exoskeletons because of their inherited compliance and high power-weight ratio. However, creating accurate models remains difficult mainly due to the compliance issue; the model can be changed by the force applied by the user. Therefore, both user and robot actions need to be considered for sufficient excitation of PAMs that are equipped in exoskeleton robots, unlike typical rigid actuators that can only be sufficiently excited by robot actions. In this paper, we propose a user-robot collaborative excitation approach for PAM model identification as an active learning framework for sequentially collecting data by deriving and executing optimal user and robot actions at each step with Gaussian processes. The optimal actions, which are executed by the robot, are displayed on a monitor that enables the user to execute them. We conducted experiments using a powered elbow exoskeleton with a PAM actuator. Experimental results show that our method can more efficiently identify the PAM model than a standard model identification method that does not use any data acquired through user-robot collaboration.
Masashi Hamaya, Takamitsu Matsubara, Tomoyuki Noda, Tatsuya Teramae, Jun Morimoto
IROS2
2017 Deep dynamic policy programming for robot control with raw images
abstract
Deep reinforcement learning has drawn much attention in robot control since it enables agents to learn control policies from very high dimensional states such as raw images. On the other hand, its dependency upon the availability of a significant quantity of training samples and its fragility in learning makes it difficult to apply for real world robot tasks. To alleviate these issues we propose Deep Dynamic Policy Programming (DDPP), which combines the sample efficiency and smooth policy updates of dynamic policy programming with the contemporary deep reinforcement learning framework. The effectiveness of the proposed method is first demonstrated in a simulation of the robot arm control problem, with comparison to Deep Q-Networks. As validation on a real robot system, DDPP also successfully learned the flipping of a handkerchief with a NEXTAGE humanoid robot using a reduced number of learning samples, whereas Deep Q-Networks failed to learn the task.
Yoshihisa Tsurumine, Yunduan Cui, Eiji Uchibe, Takamitsu Matsubara
IROS4
2017 Kernel dynamic policy programming: Applicable reinforcement learning to robot systems with high dimensional states
Yunduan Cui, Takamitsu Matsubara, Kenji Sugimoto
Neural Networks2
2017 Learning assistive strategies for exoskeleton robots from user-robot physical interaction
abstract
Social demand for exoskeleton robots that physically assist humans has been increasing in various situations due to the demographic trends of aging populations. With exoskeleton robots, an assistive strategy is a key ingredient. Since interactions between users and exoskeleton robots are bidirectional, the assistive strategy design problem is complex and challenging. In this paper, we explore a data-driven learning approach for designing assistive strategies for exoskeletons from user-robot physical interaction. We formulate the learning problem of assistive strategies as a policy search problem and exploit a data-efficient model-based reinforcement learning framework. Instead of explicitly providing the desired trajectories in the cost function, our cost function only considers the user’s muscular effort measured by electromyography signals (EMGs) to learn the assistive strategies. The key underlying assumption is that the user is instructed to perform the task by his/her own intended movements. Since the EMGs are observed when the intended movements are achieved by the user’s own muscle efforts rather than the robot’s assistance, EMGs can be interpreted as the “cost” of the current assistance. We applied our method to a 1-DoF exoskeleton robot and conducted a series of experiments with human subjects. Our experimental results demonstrated that our method learned proper assistive strategies that explicitly considered the bidirectional interactions between a user and a robot with only 60 seconds of interaction. We also showed that our proposed method can cope with changes in both the robot dynamics and movement trajectories.
Masashi Hamaya, Takamitsu Matsubara, Tomoyuki Noda, Tatsuya Teramae, Jun Morimoto
Pattern Recognit. Lett.2
2016 Learning assistive strategies from a few user-robot interactions: Model-based reinforcement learning approach
abstract
Designing an assistive strategy for exoskeletons is a key ingredient in movement assistance and rehabilitation. While several approaches have been explored, most studies are based on mechanical models of the human user, i.e., rigid-body dynamics or Center of Mass (CoM)-Zero Moment Point (ZMP) inverted pendulum moECenter of Massdel, or only focus on periodic movements with using oscillator models. On the other hand, the interactions between the user and the robot are often not considered explicitly because of its difficulty in modeling. In this paper, we propose to learn the assistive strategies directly from interactions between the user and the robot. We formulate the learning problem of assistive strategies as a policy search problem. To alleviate heavy burdens to the user for data acquisition, we exploit a data-efficient model-based reinforcement learning framework. To validate the effectiveness of our approach, an experimental platform composed of a real subject, an electromyography (EMG)-measurement system, and a simulated robot arm is developed. Then, a learning experiment with the assistive control task of the robot arm is conducted. As a result, proper assistive strategies that can achieve the robot control task and reduce EMG signals of the user are acquired only by 30 seconds interactions.
Masashi Hamaya, Takamitsu Matsubara, Tomoyuki Noda, Tatsuya Teramae, Jun Morimoto
ICRA2
2015 Reinforcement learning of shared control for dexterous telemanipulation: Application to a page turning skill
abstract
The ultimate goal of this study is to develop a method that can accomplish dexterous manipulation of various non-rigid objects by a robotic hand. In this paper, we propose a novel model-free approach using reinforcement learning to learn a shared control policy for dexterous telemanipulation by a human operator. A shared control policy is a probabilistic mapping from the human operator's (master) action and complementary sensor data to the robot (slave) control input for robot actuators. Through the learning process, our method can optimize the shared control policy so that it cooperates to the operator's policy and compensates the lack of sensory information of the operator using complementary sensor data to enhance the dexterity. To validate our method, we adopted a page turning task by telemanipulation and developed an experimental platform with a paper page model and a robot fingertip in simulation. Since the human operator cannot perceive the tactile information of the robot, it may not be as easy as humans do directly. Experimental results suggest that our method is able to learn task-relevant shared control for flexible and enhanced dexterous manipulation by a teleoperated robotic fingertip without tactile feedback to the operator.
Takamitsu Matsubara, Takahiro Hasegawa, Kenji Sugimoto
RO-MAN1
2015 Sequential intention estimation of a mobility aid user for intelligent navigational assistance
abstract
This paper proposes an intelligent mobility aid framework aimed at mitigating the impact of cognitive and/or physical user deficiencies by performing suitable mobility assistance with minimum interference. To this end, a user action model using Gaussian Process Regression (GPR) is proposed to encapsulate the probabilistic and nonlinear relationships among user action, state of the environment and user intention. Moreover, exploiting the analytical tractability of the predictive distribution allows a sequential Bayesian process for user intention estimation to take place. The proposed scheme is validated on data obtained in an indoor setting with an instrumented robotic wheelchair augmented with sensorial feedback from the environment and user commands as well as proprioceptive information from the actual vehicle, achieving accuracy in near real-time of ~80%. The initial results are promising and indicating the suitability of the process to infer user driving behaviors within the context of ambulatory robots designed to provide assistance to users with mobility impairments while carrying out regular daily activities.
Takamitsu Matsubara, Jaime Valls Miró, Daisuke Tanaka, James Poon, Kenji Sugimoto
RO-MAN1
2014 Object manifold learning with action features for active tactile object recognition
abstract
In this paper, we consider an object recognition problem based on tactile information using a robot hand. The robot performs an exploratory action to the object to obtain the tactile information, however, poorly designed actions may not be sufficiently informative. In contrast, if we could collect sample data by sequentially performing informative actions, i.e., active learning, the required time would be drastically reduced. To this end, we propose a novel approach for active tactile object recognition. Our approach combines both an active learning scheme and a nonlinear dimensionality reduction method. We first extracts the object manifold, each coordinate of which represents an object, from tactile sensor data and action features using Gaussian Process Latent Variable Models. At the same time, a probabilistic model of the observed data related to the action and the object are learned. Then, with the learned model, optimally-informative exploratory actions can be computed sequentially, and performed to efficiently collect the data for recognition. We show experimental results that verify the effectiveness of our proposed method with synthetic data and a real robot.
Daisuke Tanaka, Takamitsu Matsubara, Kentaro Ichien, Kenji Sugimoto
IROS2
2014 Real-time estimation of Human-Cloth topological relationship using depth sensor for robotic clothing assistance
abstract
In this study, we propose a novel method for the real-time estimation of Human-Cloth relationship, which is crucial for efficient motor skill learning in Robotic Clothing Assistance. This system relies on the use of low cost depth sensor, which provides color and depth images without requiring an elaborate setup making it suitable for real-world applications. We present an efficient algorithm to estimate the parameters that represent the topological relationship between human and the clothing article. At the core of our approach are low dimensional representation of Human-Cloth relationship using topology coordinates for fast learning of motor skills and a unified ellipse fitting algorithm for the compact representation of the state of clothing articles. We conducted experiments that illustrate the robustness of these feature representations. Furthermore, we evaluated the performance of our proposed method by applying it to real-time clothing assistance tasks and compared the estimates provided by our method with the ground truth.
Nishanth Koganti, Tomoya Tamei, Takamitsu Matsubara, Tomohiro Shibata
RO-MAN3
2014 Latent Kullback Leibler Control for Continuous-State Systems using Probabilistic Graphical Models
Takamitsu Matsubara, Vicenç Gómez, Hilbert J. Kappen
UAI1
2012 Spatio-temporal synchronization of periodic movements by style-phase adaptation: Application to biped walking
abstract
In this paper, we propose a framework for generating coordinated periodic movements of robotic systems with external inputs. We developed an adaptive pattern generator model that is composed of a two-factor observation model with a style parameter and phase dynamics with a phase variable. The style parameter controls the spatial patterns of the generated trajectories, and the phase variable controls its temporal profiles. To validate the effectiveness of our proposed method, we applied it to a simulated humanoid model to perform biped walking behaviors coordinated with observed walking patterns and the environment. The robot successfully performed stable biped walking behaviors even when the style of the observed walking pattern and the period were suddenly changed.
Takamitsu Matsubara, Akimasa Uchikata, Jun Morimoto
ICRA1
2012 Full-body exoskeleton robot control for walking assistance by style-phase adaptive pattern generation
abstract
We propose an adaptive walking assistance strategy to control an exoskeleton robot. In our proposed framework, we explicitly consider the following: 1) the diversity of user motions (style) and 2) the interactions among a user, a robot, and an environment. To spatially coordinate a wide variety of user motions and robot behaviors, we estimated style parameters from observed user movements. To temporally coordinate the interactions among the user, the robot, and the environment, we synchronized the phases of these three systems with a coupled oscillator model. The estimated style parameters and the phase of the user motion can be used to predict future user movements. We investigated how movement prediction and phase synchronization can be beneficial to control an exoskeleton robot. To evaluate our adaptive walking assistance strategy, we developed simulated user and exoskeleton models. The physical interactions among the user, the exoskeleton, and the ground models are introduced in the simulated system. We show that the necessary torque for the user walking movement was reduced around 40% by using our proposed method to control the exoskeleton model.
Takamitsu Matsubara, Akimasa Uchikata, Jun Morimoto
IROS1
2012 Adaptive choreography for user's preferences on personal robots
abstract
We propose an approach to efficiently adapt the non-verbal information of personal robots to a user's preferences. As non-verbal information focuses on gestures except meaningful gestures, in our method the adaptation problem is formalized as choreography; the action (as one unit of the non-verbal information) is assigned to each state as preferred by the user. Since the state is defined based on the classification of language focusing on the shallow discourse structure proposed by [1], our method is constructed with independence of tasks and situations unlike previous methods. Furthermore, using a preferences database obtained from multiple users, we produce a User Preference Model (UPM) that can represent various user preferences by a small number of parameters. A new user is asked to assign actions on a few states to adapt the UPM for the user preference. After the adaptation, the UPM can be used to predict actions on other states as preferred by the user, that accomplishes adaptive choreography. We implemented the proposed method with a personal robot, and validated it through experiments on multiple tasks with human subjects. As a result, we confirmed that the user's impressions of the robot were greatly improved by our method for multiple tasks.
Takamitsu Matsubara, Shizuko Matsuzoe, Masatsugu Kidode
RO-MAN1
2012 Real-time stylistic prediction for whole-body human motions
Takamitsu Matsubara, Sang-Ho Hyon, Jun Morimoto
Neural Networks1
2011 XoR: Hybrid drive exoskeleton robot that can balance
abstract
We propose a novel exoskeleton robot prototype aimed at a brain-machine interface and rehabilitation for postural control for elderly people, people with spinal cord injury, stroke patients, and others with similar needs. By arranging pneumatic muscles with electric motors in a optimal way, one can achieve both weight-reduction and torque-controllability. Its anthropomorphic design and torque-controllability enable users to implement and test various rehabilitation/compensation programs consistent with human motor control and learning mechanism. Hybrid drive itself is not new, but its specialized application to lightweight exoskeleton is novel. This paper reports the design and development of the robot, particularly addressing a hybrid drive for load-bearing tasks such as standing and postural maintenance. The experimental data as well as the attached videos demonstrate the effectiveness of the proposed system.
Sang-Ho Hyon, Jun Morimoto, Takamitsu Matsubara, Tomoyuki Noda, Mitsuo Kawato
IROS3
2011 Learning parametric dynamic movement primitives from multiple demonstrations
Takamitsu Matsubara, Sang-Ho Hyon, Jun Morimoto
Neural Networks1
2010 Learning Basis Representations of Inverse Dynamics Models for Real-Time Adaptive Control
Yasuhito Horiguchi, Takamitsu Matsubara, Masatsugu Kidode
ICONIP (2)2
2010 Learning Parametric Dynamic Movement Primitives from Multiple Demonstrations
Takamitsu Matsubara, Sang-Ho Hyon, Jun Morimoto
ICONIP (1)1
2010 Optimal Feedback Control for anthropomorphic manipulators
abstract
We study target reaching tasks of redundant anthropomorphic manipulators under the premise of minimal energy consumption and compliance during motion. We formulate this motor control problem in the framework of Optimal Feedback Control (OFC) by introducing a specific cost function that accounts for the physical constraints of the controlled plant. Using an approximative computational optimal control method we can optimally control a high-dimensional anthropomorphic robot without having to specify an explicit inverse kinematics, inverse dynamics or feedback control law. We highlight the benefits of this biologically plausible motor control strategy over traditional (open loop) optimal controllers: The presented approach proves to be significantly more energy efficient and compliant, while being accurate with respect to the task at hand. These properties are crucial for the control of mobile anthropomorphic robots, that are designed to interact safely in a human environment. To the best of our knowledge this is the first OFC implementation on a high-dimensional (redundant) manipulator.
Djordje Mitrovic, Sho Nagashima, Stefan Klanke, Takamitsu Matsubara, Sethu Vijayakumar
ICRA4
2010 Learning Stylistic Dynamic Movement Primitives from multiple demonstrations
abstract
In this paper, we propose a novel concept of movement primitives called Stylistic Dynamic Movement Primitives (SDMPs) for motor learning and control in humanoid robotics. In the SDMPs, a diversity of styles in human motion observed through multiple demonstrations can be compactly encoded in a movement primitive, and this allows style manipulation of motion sequences generated from the movement primitive by a control variable called a style parameter. Focusing on discrete movements, a model of the SDMPs is presented as an extension of Dynamic Movement Primitives (DMPs) proposed by Ijspeert et al.. A novel learning procedure of the SDMPs from multiple demonstrations, including a diversity of motion styles, is also described. We present two practical applications of the SDMPs, i.e., stylistic table tennis swings and obstacle avoidance with an anthropomorphic manipulator.
Takamitsu Matsubara, Sang-Ho Hyon, Jun Morimoto
IROS1
2007 Learning to acquire whole-body humanoid CoM movements to achieve dynamic tasks
abstract
This paper presents a novel approach to acquire dynamic whole-body movements on humanoid robots focused on learning a control policy for the center of mass. A policy-gradient method is used to acquire a CoM movement as a control policy for achieving a desired dynamic task. A CoM-Jacobian-based redundancy resolution is then used to compute angular velocities for all joints in order to achieve a whole-body movement consistent with the CoM movement acquired through learning. To demonstrate the effectiveness of our method, we apply it in simulation to the learning of a strong punching movement on the Fujitsu humanoid robot, Hoap-2.
Takamitsu Matsubara, Jun Morimoto, Jun Nakanishi, Sang-Ho Hyon, Joshua G. Hale, Gordon Cheng
ICRA1
2005 Learning CPG Sensory Feedback with Policy Gradient for Biped Locomotion for a Full-Body Humanoid
Gen Endo, Jun Morimoto, Takamitsu Matsubara, Jun Nakanishi, Gordon Cheng
AAAI3
2005 Learning Sensory Feedback to CPG with Policy Gradient for Biped Locomotion
abstract
This paper proposes a learning framework for a CPG-based biped locomotion controller using a policy gradient method. Our goal in this study is to develop an efficient learning algorithm by reducing the dimensionality of the state space used for learning. We demonstrate that an appropriate feedback controller in the CPG-based controller can be acquired using the proposed method within a few thousand trials by numerical simulations. Furthermore, we implement the learned controller on the physical biped robot to experimentally show that the learned controller successfully works in the real environment.
Takamitsu Matsubara, Jun Morimoto, Jun Nakanishi, Masa-aki Sato, Kenji Doya
ICRA1