EDBT 2026 Demo / reviewers in the wild / expert
Masashi Hamaya
dblp:164/8431
· DBLP profile ↗
34ranked-venue papers
7as first author
22since 2021 · last 2025
0000-0003-4189-8219ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 7 first-author · 22 since 2021Systems, architecture and hardware · 30 · 6 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph FormabstractDesigning a safe policy for uncertain environments is crucial in real-world control systems. However, this challenge remains inadequately addressed within the Markov decision process (MDP) framework. This paper presents the first algorithm guaranteed to identify a near-optimal policy in a robust constrained MDP (RCMDP), where an optimal policy minimizes cumulative cost while satisfying constraints in the worst-case scenario across a set of environments. We first prove that the conventional policy gradient approach to the Lagrangian max-min formulation can become trapped in suboptimal solutions. This occurs when its inner minimization encounters a sum of conflicting gradients from the objective and constraint functions. To address this, we leverage the epigraph form of the RCMDP problem, which resolves the conflict by selecting a single gradient from either the objective or the constraints. Building on the epigraph form, we propose a bisection search algorithm with a policy gradient subroutine and prove that it identifies an $\varepsilon$-optimal policy in an RCMDP with $\widetilde{\mathcal{O}}(\varepsilon^{-4})$ robust policy evaluations. Toshinori Kitamura, Tadashi Kozuno, Wataru Kumagai, Kenta Hoshino, Yohei Hosoe, Kazumi Kasaura, Masashi Hamaya, Paavo Parmas, Yutaka Matsuo |
ICLR | 7 |
| 2025 | SCU-Hand: Soft Conical Universal Robotic Hand for Scooping Granular Media from Containers of Various SizesabstractAutomating small-scale experiments in materials science presents challenges due to the heterogeneous nature of experimental setups. This study introduces the SCU-Hand (Soft Conical Universal Robot Hand), a novel end-effector designed to automate the task of scooping powdered samples from various container sizes using a robotic arm. The SCU-Hand employs a flexible, conical structure that adapts to different container geometries through deformation, maintaining consistent contact without complex force sensing or machine learning-based control methods. Its reconfigurable mechanism allows for size adjustment, enabling efficient scooping from diverse container types. By combining soft robotics principles with a sheet-morphing design, our end-effector achieves high flexibility while retaining the necessary stiffness for effective powder manipulation. We detail the design principles, fabrication process, and experimental validation of the SCU-Hand. Experimental validation showed that the scooping capacity is about 20% higher than that of a commercial tool, with a scooping performance of more than 95% for containers of sizes between 67 mm to 110 mm. This research contributes to laboratory automation by offering a cost-effective, easily implementable solution for automating tasks such as materials synthesis and characterization processes. Tomoya Takahashi, Cristian C. Beltran-Hernandez, Yuki Kuroda, Kazutoshi Tanaka, Masashi Hamaya, Yoshitaka Ushiku |
ICRA | 5 |
| 2025 | Pose Estimation of a Cable-Driven Serpentine Manipulator Utilizing Intrinsic Dynamics via Physical Reservoir ComputingabstractCable-driven serpentine manipulators hold great potential in unstructured environments, offering obstacle avoidance, multi-directional force application, and a lightweight design. By placing all motors and sensors at the base and employing plastic links, we can further reduce the arm’s weight. To demonstrate this concept, we developed a 9-degree-of-freedom cable-driven serpentine manipulator with an arm length of 545 mm and a total mass of only 308 g. However, this design introduces flexibility-induced variations, such as cable slack, elongation, and link deformation. These variations result in discrepancies between analytical predictions and actual link positions, making pose estimation more challenging. To address this challenge, we propose a physical reservoir computing based pose estimation method that exploits the manipulator’s intrinsic nonlinear dynamics as a high-dimensional reservoir. Experimental results show a mean pose error of 4.3 mm using our method, compared to 4.4 mm with a baseline long short-term memory network and 39.5 mm with an analytical approach. This work provides a new direction for control and perception strategies in lightweight cable-driven serpentine manipulators leveraging their intrinsic dynamics. Kazutoshi Tanaka, Tomoya Takahashi, Masashi Hamaya |
IROS | 3 |
| 2024 | SliceIt! - A Dual Simulator Framework for Learning Robot Food SlicingabstractCooking robots can enhance the home experience by reducing the burden of daily chores. However, these robots must perform their tasks dexterously and safely in shared human environments, especially when handling dangerous tools such as kitchen knives. This study focuses on enabling a robot to autonomously and safely learn food-cutting tasks. More specifically, our goal is to enable a collaborative robot or industrial robot arm to perform food-slicing tasks by adapting to varying material properties using compliance control. Our approach involves using Reinforcement Learning (RL) to train a robot to compliantly manipulate a knife, by reducing the contact forces exerted by the food items and by the cutting board. However, training the robot in the real world can be inefficient, and dangerous, and result in a lot of food waste. Therefore, we proposed SliceIt!, a framework for safely and efficiently learning robot food-slicing tasks in simulation. Following a real2sim2real approach, our framework consists of collecting a few real food slicing data, calibrating our dual simulation environment (a high-fidelity cutting simulator and a robotic simulator), learning compliant control policies on the calibrated simulation environment, and finally, deploying the policies on the real robot. Cristian C. Beltran-Hernandez, Nicolas Erbetti, Masashi Hamaya |
ICRA | 3 |
| 2024 | An Electromagnetism-Inspired Method for Estimating In-Grasp Torque from Visuotactile SensorsabstractTactile sensing has become a popular sensing modality for robot manipulators, due to the promise of providing robots with the ability to measure the rich contact information that gets transmitted through its sense of touch. Among the diverse range of information accessible from tactile sensors, torques transmitted from the grasped object to the fingers through extrinsic environmental contact may be particularly important for tasks such as object insertion. However, tactile torque estimation has received relatively little attention when compared to other sensing modalities, such as force, texture, or slip identification. In this work, we introduce the notion of the Tactile Dipole Moment, which we use to estimate tilt torques from gel-based visuotactile sensors. This method does not rely on deep learning, sensor-specific mechanical, or optical modeling, and instead takes inspiration from electromechanics to analyze the vector field produced from 2D marker displacements. Despite the simplicity of our technique, we demonstrate its ability to provide accurate torque readings over two different tactile sensors and three object geometries, and highlight its practicality for the task of USB stick insertion with a compliant robot arm. These results suggest that simple analytical calculations based on dipole moments can sufficiently extract physical quantities from visuotactile sensors. Yuni Fuchioka, Masashi Hamaya |
ICRA | 2 |
| 2024 | Symmetry-aware Reinforcement Learning for Robotic Assembly under Partial Observability with a Soft WristabstractThis study tackles the representative yet challenging contact-rich peg-in-hole task of robotic assembly, using a soft wrist that can operate more safely and tolerate lower-frequency control signals than a rigid one. Previous studies often use a fully observable formulation, requiring external setups or estimators for the peg-to-hole pose. In contrast, we use a partially observable formulation and deep reinforcement learning from demonstrations to learn a memory-based agent that acts purely on haptic and proprioceptive signals. Moreover, previous works do not incorporate potential domain symmetry and thus must search for solutions in a bigger space. Instead, we propose to leverage the symmetry for sample efficiency by augmenting the training data and constructing auxiliary losses to force the agent to adhere to the symmetry. Results in simulation with five different symmetric peg shapes show that our proposed agent can be comparable to or even outperform a state-based agent. In particular, the sample efficiency also allows us to learn directly on the real robot within 3 hours. Tadashi Kozuno, Cristian C. Beltran-Hernandez, Masashi Hamaya |
ICRA | 4 |
| 2024 | Vision-Language Interpreter for Robot Task PlanningabstractLarge language models (LLMs) are accelerating the development of language-guided robot planners. Meanwhile, symbolic planners offer the advantage of interpretability. This paper proposes a new task that bridges these two trends, namely, multimodal planning problem specification. The aim is to generate a problem description (PD), a machine-readable file used by the planners to find a plan. By generating PDs from language instruction and scene observation, we can drive symbolic planners in a language-guided framework. We propose a Vision-Language Interpreter (ViLaIn), a new framework that generates PDs using state-of-the-art LLM and vision-language models. ViLaIn can refine generated PDs via error message feedback from the symbolic planner. Our aim is to answer the question: How accurately can ViLaIn and the symbolic planner generate valid robot plans? To evaluate ViLaIn, we introduce a novel dataset called the problem description generation (ProDG) dataset. The framework is evaluated with four new evaluation metrics. Experimental results show that ViLaIn can generate syntactically correct problems with more than 99% accuracy and valid plans with more than 58% accuracy. Our code and dataset are available at https://github.com/omron-sinicx/ViLaIn. Keisuke Shirai, Cristian C. Beltran-Hernandez, Masashi Hamaya, Atsushi Hashimoto 0001, Shohei Tanaka, Kento Kawaharazuka, Kazutoshi Tanaka, Yoshitaka Ushiku, Shinsuke Mori |
ICRA | 3 |
| 2024 | Robotic Object Insertion with a Soft Wrist through Sim-to-Real Privileged TrainingabstractThis study addresses contact-rich object insertion tasks under unstructured environments using a robot with a soft wrist, enabling safe contact interactions. For the unstructured environments, we assume that there are uncertainties in object grasp and hole pose and that the soft wrist pose cannot be directly measured. Recent methods employ learning approaches and force/torque sensors for contact localization; however, they require data collection in the real world. This study proposes a sim-to-real approach using a privileged training strategy. This method has two steps. 1) The teacher policy is trained to complete the task with sensor inputs and ground truth privileged information such as the peg pose, and then 2) the student encoder is trained with data produced from teacher policy rollouts to estimate the privileged information from sensor history. We performed sim-to-real experiments under grasp and hole pose uncertainties. This resulted in 100%, 95%, and 80% success rates for circular peg insertion with 0°, +5°, and -5° peg misalignments, respectively, and start positions randomly shifted ± 10 mm from a default position. Also, we tested the proposed method with a square peg that was never seen during training. Additional simulation evaluations revealed that using the privileged strategy improved success rates compared to training with only simulated sensor data. Our results demonstrate the advantage of using sim-to-real privileged training for soft robots, which has the potential to alleviate human engineering efforts for robotic assembly. Yuni Fuchioka, Cristian C. Beltran-Hernandez, Masashi Hamaya |
IROS | 4 |
| 2024 | Learning Variable Compliance Control From a Few Demonstrations for Bimanual Robot with Haptic Feedback Teleoperation SystemabstractAutomating dexterous, contact-rich manipulation tasks using rigid robots is a significant challenge in robotics. Rigid robots, defined by their actuation through position commands, face issues of excessive contact forces due to their inability to adapt to contact with the environment, potentially causing damage. While compliance control schemes have been introduced to mitigate these issues by controlling forces via external sensors, they are hampered by the need for fine-tuning task-specific controller parameters. Learning from Demonstrations (LfD) offers an intuitive alternative, allowing robots to learn manipulations through observed actions. In this work, we introduce a novel system to enhance the teaching of dexterous, contact-rich manipulations to rigid robots. Our system is twofold: firstly, it incorporates a teleoperation interface utilizing Virtual Reality (VR) controllers, designed to provide an intuitive and cost-effective method for task demonstration with haptic feedback. Secondly, we present Comp-ACT (Compliance Control via Action Chunking with Transformers), a method that leverages the demonstrations to learn variable compliance control from a few demonstrations. Our methods have been validated across various complex contact-rich manipulation tasks using single-arm and bimanual robot setups in simulated and real-world environments, demonstrating the effectiveness of our system in teaching robots dexterous manipulations with enhanced adaptability and safety. Code available at https://github.com/omron-sinicx/CompACT. Tatsuya Kamijo, Cristian C. Beltran-Hernandez, Masashi Hamaya |
IROS | 3 |
| 2024 | Low-Cost Air Hockey Robot Using a Five-Bar Linkage Mechanism Driven by Position-Control ServomotorsabstractIn human-robot interaction (HRI) research, ball games pose significant challenges that demand robotic solutions that are both cost-effective and user-friendly for non-experts. Air hockey, characterized by safe, non-direct-contact play and a simplified state-action space, emerges as an ideal platform for such research. Despite the availability of various air hockey robots, their high cost and complexity have limited widespread use among researchers requiring robotics expertise. Addressing this gap, we introduce a low-cost, accessible air hockey robot designed to facilitate HRI studies. Featuring a lightweight five-bar linkage mechanism powered by low-cost servomotors for position control, this robot combines efficiency with ease of use. The complete robot’s cost is estimated at $346.8, with the arm weighing a mere 19 grams. The robot precisely returns the puck by intermittently adjusting its target joint positions, achieving a play with an average return error of 42.6 mm. These characteristics affirm the robot’s potential as a valuable tool for advancing HRI research. Mirai Shinjo, Cristian C. Beltran-Hernandez, Masashi Hamaya, Kazutoshi Tanaka |
IROS | 3 |
| 2024 | Visuo-Tactile Zero-Shot Object Recognition with Vision-Language ModelabstractTactile perception is vital, especially when distinguishing visually similar objects. We propose an approach to incorporate tactile data into a Vision-Language Model (VLM) for visuo-tactile zero-shot object recognition. Our approach leverages the zero-shot capability of VLMs to infer tactile properties from the names of tactilely similar objects. The proposed method translates tactile data into a textual description solely by annotating object names for each tactile sequence during training, making it adaptable to various contexts with low training costs. The proposed method was evaluated on the FoodReplica and Cube datasets, demonstrating its effectiveness in recognizing objects that are difficult to distinguish by vision alone. Shiori Ueda, Atsushi Hashimoto 0001, Masashi Hamaya, Kazutoshi Tanaka, Hideo Saito 0001 |
IROS | 3 |
| 2023 | Twist Snake: Plastic table-top cable-driven robotic arm with all motors located at the base linkabstractTable-top robotic arms for education and research must be low-cost for availability and lightweight and soft for safety. Therefore, as such a robot, this study focuses on designing a plastic table-top cable-driven robotic arm with all motors located at the base link. However, locating all motors at the base link results in a significant distance between a driving motor and driven joint, increases the number of parts for the force transmission, and increases the risk of a cable loosening and coming off of a pulley. To overcome these issues, this study proposed a novel cable-driven robotic arm named Twist Snake. We designed a joint composition of Twist Snake to minimize the number of parts for the force transmission. In addition, it has a compact cable-pretension/termination-mechanism and covering parts to prevent the cable from loosening and coming off of the pulley. The arm comprised 475 mm long moving links with an 802 g. The feasibility of the arm was experimentally demonstrated by contact rich tasks, the insertion of a toy peg into a hole and swiping a whiteboard with a cleaner. The optimization of the proposed design and the development of a learning method for the arm that leverages contact will be investigated in future work. Kazutoshi Tanaka, Masashi Hamaya |
ICRA | 2 |
| 2023 | Learning Food Picking without Food: Fracture Anticipation by Breaking Reusable Fragile ObjectsabstractFood picking is trivial for humans but not for robots, as foods are fragile. Presetting foods' physical properties does not help robots much due to the objects' inter- and intra-category diversity. A recent study proved that learning-based fracture anticipation with tactile sensors could overcome this problem; however, the method trains the model for each food to deal with intra-category differences, and tuning robots for each food leads to an undesirable amount of food consumption. This study proposes a novel framework for learning food-picking tasks without consuming foods. The key idea is to leverage the object-breaking experiences of several reusable fragile objects instead of consuming real foods while making the picking ability object-invariant with domain generalization (DG). In real-robot experiments, we trained a model with reusable objects (toy blocks, ping-pong balls, and jellies), selected based on the three common fracture types (crack, rupture, and crush). We then tested the model with four real food objects (tofu, bananas, potato chips, and tomatoes). The results showed that the proposed combination of reusable objects' breaking experiences and DG is effective for the food-picking task. Rinto Yagawa, Reina Ishikawa, Masashi Hamaya, Kazutoshi Tanaka, Atsushi Hashimoto 0001, Hideo Saito 0001 |
ICRA | 3 |
| 2023 | Learning Robotic Powder Weighing from Simulation for Laboratory AutomationabstractThis study focuses on a robotic powder weighing task used in laboratory automation. In this task, a robot weighs a certain amount of powder with a milligram-level target mass using a dispensing spoon. The complex dynamics of the powder, the variations in the materials being weighed, and the need to balance conservative and aggressive actions are significant challenges in the robotics field. Therefore, learning approaches are critical for this task. However, many learning interactions in real-world environments require substantial efforts to clean the spread powder. To overcome this issue, this study employs a sim-to-real transfer learning approach using a domain randomization (DR) technique. This enables the robot to weigh various powders with a small target mass and alleviates the burden of collecting data in a real-world environment. Herein, we formulated weighing manipulation as a reinforcement learning problem. Besides, we developed a powder weighing simulator and carefully selected the dynamics parameters used for DR to adapt to unseen environments. A recurrent neural network-based policy was adopted considering the balance of conservative and aggressive actions. The sim-to-real zero-shot transfer experiments demonstrated that the robot completed the weighing tasks with an average weighing error of 0.1 - 0.2 mg for different powder materials and target masses (5 - 15 mg). Overall, this approach shows promising results and can be useful for automating laboratory tasks that involve weighing powders. Yuki Kadokawa, Masashi Hamaya, Kazutoshi Tanaka |
IROS | 2 |
| 2023 | Robotic Powder Grinding with Audio-Visual Feedback for Laboratory Automation in Materials ScienceabstractThis study focuses on the powder grinding process, which is a necessary step for material synthesis in materials science experiments. In material science, powder grinding is a time-consuming process that is typically executed by hand, as commercial grinding machines are unsuitable for samples of small size. Robotic powder grinding would solve this problem, but it is a challenging task for robots, as it requires observing the powder state and generating appropriate motions. Our previous study proposed a robotic powder grinding system using visual feedback. Although visual feedback is helpful for observing the powder distribution, the particle size during the grinding process remains invisible, leading to suboptimal robot actions. In some cases, the robot chose to gather the powder even though continuing to grind instead would have produced finer powder. In this paper, we present a multi-modal robotic grinding system that utilizes both audio and visual feedback. It makes use of the grinding sound which carries information about the grinding progress, as the particle size strongly affects the audio intensity. The audio feedback enables the robot to grind until the powder is sufficiently fine. In our experiments, the robot ground 80.5% of the powder to a particle size smaller than$250\ \mu\mathrm{m}$with audio and visual feedback and 68% without audio feedback, indicating that multi-modal feedback is an effective tool to produce finer powder. We conclude that the addition of audio feedback provides crucial information to the robot, allowing it to better understand the progress of the grinding process and make more optimal decisions. This robot system can be used to prepare samples in material science experiments and analyze the grinding process. Yusaku Nakajima, Masashi Hamaya, Kazutoshi Tanaka, Takafumi Hawai, Felix von Drigalski, Yasuo Takeichi, Yoshitaka Ushiku, Kanta Ono |
IROS | 2 |
| 2023 | Learning Robotic Assembly by Leveraging Physical Softness and Tactile SensingabstractThis study aims to achieve autonomous robotic assembly under uncertain conditions arising from imprecise goal positioning and variations in the angle of the grasped part. Soft robots are suitable for such uncertain and contact-rich environments and are capable of insertion tasks with imprecise goal positions. However, we may also struggle to handle further uncertainty, such as variations in grasping pose. To address the challenge posed by multiple sources of uncertainty, we equipped the soft robot with a tactile sensor. Our key insight is that tactile signal patterns are closely linked to the subtask transitions in an assembly process, specifically from the search to insertion subtasks. We hypothesize soft robots could complete the task by exploring the transition via tactile signals, even in scenarios with imprecise goal positions and grasp misalignment. To this end, we develop an anomaly detection model using a Variational Autoencoder to identify the timing of these transitions. We then employ learning and heuristic-based controllers to navigate the peg tip to the hole and perform the insertion. Our method was validated through real-robot experiments using a soft wrist and a vision-based tactile sensor. The results demonstrate that our method achieves a 100% success rate in scenarios with less uncertain goal pose ($\sigma=2\text{mm}$) and grasp misalignment (up to 5°) and a 70% success rate in scenarios with uncertain goal pose ($\sigma=10\text{mm}$) and grasp misalignment (up to 20°). Moreover, our anomaly detection model can generalize to different peg diameters without additional training. Joaquín Royo-Miquel, Masashi Hamaya, Cristian C. Beltran-Hernandez, Kazutoshi Tanaka |
IROS | 2 |
| 2023 | Elastic Decision TransformerabstractThis paper introduces Elastic Decision Transformer (EDT), a significant advancement over the existing Decision Transformer (DT) and its variants. Although DT purports to generate an optimal trajectory, empirical evidence suggests it struggles with trajectory stitching, a process involving the generation of an optimal or near-optimal trajectory from the best parts of a set of sub-optimal trajectories. The proposed EDT differentiates itself by facilitating trajectory stitching during action inference at test time, achieved by adjusting the history length maintained in DT. Further, the EDT optimizes the trajectory by retaining a longer history when the previous trajectory is optimal and a shorter one when it is sub-optimal, enabling it to "stitch" with a more optimal trajectory. Extensive experimentation demonstrates EDT's ability to bridge the performance gap between DT-based and Q Learning-based approaches. In particular, the EDT outperforms Q Learning-based methods in a multi-task regime on the D4RL locomotion benchmark and Atari games. Yueh-Hua Wu, Xiaolong Wang 0004, Masashi Hamaya |
NeurIPS | 3 |
| 2022 | Robotic Powder Grinding with a Soft Jig for Laboratory Automation in Material ScienceabstractGrinding materials into a fine powder is a time-consuming task in material science that is generally performed by hand, as current automated grinding machines might not be suitable for preparing small-sized samples. This study presents a robotic powder grinding system for laboratory automation in material science applications that observe the powder's state to improve the grinding outcome. We developed a soft jig consisting of off-the-shelf gel materials and 3D-printed parts, which can be used with any robot arm to perform powder grinding. The jig's physical softness allows for safe grinding without force sensing. In addition, we developed a visual feedback system that observes the powder distribution and decides where to grind and when to gather. The results showed that our system could grind 79 percent of the powder to a particle size smaller than 200 μm by using the soft jig and visual feedback. This ratio was 57% when using only the soft jig without feedback. Our system can be used immediately in laboratories to alleviate the workload of researchers. Yusaku Nakajima, Masashi Hamaya, Takafumi Hawai, Felix von Drigalski, Kazutoshi Tanaka, Yoshitaka Ushiku, Kanta Ono |
IROS | 2 |
| 2021 | Precise Multi-Modal In-Hand Pose Estimation using Low-Precision Sensors for Robotic AssemblyabstractIn industrial assembly tasks, the in-hand pose of grasped objects needs to be known with high precision for subsequent manipulation tasks such as insertion. This problem (in-hand-pose estimation) has traditionally been addressed using visual recognition or tactile sensing. On the one hand, while visual recognition can provide efficient pose estimates, it tends to suffer from low precision due to noise, occlusions and calibration errors. On the other hand, tactile fingertip sensors can provide precise complementary information, but their low durability significantly limits their use in real-world applications. To get the best of both worlds, we propose an efficient method for in-hand pose estimation using off-the-shelf cameras and robot wrist force sensors, which requires no precise camera calibration. The key idea is to utilize visual and contact information adaptively to maximally reduce the uncertainty about the in-hand object pose in a Bayesian state estimation framework. As most of the uncertainty can be resolved from visual observations, our approach reduces the number of physical environment interactions while keeping a high pose estimation accuracy. Our experimental evaluation demonstrates that our approach can estimate object poses with sub-mm precision with an off-the-shelf camera and force-torque sensor. Felix von Drigalski, Kennosuke Hayashi, Yifei Huang 0002, Ryo Yonetani, Masashi Hamaya, Kazutoshi Tanaka, Yoshihisa Ijiri |
ICRA | 5 |
| 2021 | An analytical diabolo model for robotic learning and controlabstractIn this paper, we present a diabolo model that can be used for training agents in simulation to play diabolo, as well as running it on a real dual robot arm system. We first derive an analytical model of the diabolo-string system and compare its accuracy using data recorded via motion capture, which we release as a public dataset of skilled play with diabolos of different dynamics. We show that our model outperforms a deep-learning-based predictor, both in terms of precision and physically consistent behavior. Next, we describe a method based on optimal control to generate robot trajectories that produce the desired diabolo trajectory, as well as a system to transform higher-level actions into robot motions. Finally, we test our method on a real robot system playing the diabolo, and throw it to and catch it from a human player. Felix von Drigalski, Devwrat Joshi, Takayuki Murooka, Kazutoshi Tanaka, Masashi Hamaya, Yoshihisa Ijiri |
ICRA | 5 |
| 2021 | TRANS-AM: Transfer Learning by Aggregating Dynamics Models for Soft Robotic AssemblyabstractPractical industrial assembly scenarios often require robotic agents to adapt their skills to unseen tasks quickly. While transfer reinforcement learning (RL) could enable such quick adaptation, much prior work has to collect many samples from source environments to learn target tasks in a model-free fashion, which still lacks sample efficiency on a practical level. In this work, we develop a novel transfer RL method named TRANSfer learning by Aggregating dynamics Models (TRANS-AM). TRANS-AM is based on model-based RL (MBRL) for its high-level sample efficiency, and only requires dynamics models to be collected from source environments. Specifically, it learns to aggregate source dynamics models adaptively in an MBRL loop to better fit the state-transition dynamics of target environments and execute optimal actions there. As a case study to show the effectiveness of this proposed approach, we address a challenging contact-rich peg-in-hole task with variable hole orientations using a soft robot. Our evaluations with both simulation and real-robot experiments demonstrate that TRANS-AM enables the soft robot to accomplish target tasks with fewer episodes compared when learning the tasks from scratch. Kazutoshi Tanaka, Ryo Yonetani, Masashi Hamaya, Robert Lee, Felix von Drigalski, Yoshihisa Ijiri |
ICRA | 3 |
| 2021 | Learning Robotic Contact JugglingabstractRobotic contact juggling is a challenging task in which robots must control the movement of a ball rapidly and indirectly without holding it while keeping the ball in and sometimes out of contact with the robot’s body. In this work, we address the problem of learning such robotic contact juggling from trial and error via model-based reinforcement learning (MBRL). The key insight is that complex robot-ball interactions of the contact juggling actually consist of a small set of simple dynamics that each corresponds to a distinct interaction "primitive" such as touching and releasing the ball. Accordingly, we develop a tailored MBRL method that incrementally fits a set of simple dynamics models to the movements of a robot and a ball while also learning a switching model that can select a proper dynamics model depending on the current state and action. The learned model can then be used in an MBRL framework to seek optimal juggling control. We demonstrated the effectiveness of our approach on a simulator of contact juggling performed by a robotic arm. Kazutoshi Tanaka, Masashi Hamaya, Devwrat Joshi, Felix von Drigalski, Ryo Yonetani, Takamitsu Matsubara, Yoshihisa Ijiri |
IROS | 2 |
| 2020 | Contact-based in-hand pose estimation using Bayesian state estimation and particle filteringabstractIn industrial assembly tasks, the position of an object grasped by the robot has to be known with high precision in order to insert or place it. In real applications, this problem is commonly solved by jigs that are specially produced for each part. However, they significantly limit flexibility and are prohibitive when the target parts change often, so a flexible method to localize parts with high accuracy after grasping is desired. To solve this problem, we propose a method that can estimate the position of an object in the robot's hand to sub-millimeter precision, and can improve its estimate incrementally, using only minimal calibration and a force sensor. Our method is applicable to any robotic gripper and any rigid object that the gripper can hold, and requires only a force sensor. We demonstrate that the method can determine the position of an object to a precision of under 1 mm without using any part-specific jigs or equipment. Felix von Drigalski, Shohei Taniguchi, Robert Lee, Takamitsu Matsubara, Masashi Hamaya, Kazutoshi Tanaka, Yoshihisa Ijiri |
ICRA | 5 |
| 2020 | Learning Robotic Assembly Tasks with Lower Dimensional Systems by Leveraging Physical Softness and Environmental ConstraintsabstractIn this study, we present a novel control framework for assembly tasks with a soft robot. Typically, existing hard robots require high frequency controllers and precise force/torque sensors for assembly tasks. The resulting robot system is complex, entailing large amounts of engineering and maintenance. Physical softness allows the robot to interact with the environment easily. We expect soft robots to perform assembly tasks without the need for high frequency force/torque controllers and sensors. However, specific data-driven approaches are needed to deal with complex models involving nonlinearity and hysteresis. If we were to apply these approaches directly, we would be required to collect very large amounts of training data. To solve this problem, we argue that by leveraging softness and environmental constraints, a robot can complete tasks in lower dimensional state and action spaces, which could greatly facilitate the exploration of appropriate assembly skills. Then, we apply a highly efficient model-based reinforcement learning method to lower dimensional systems. To verify our method, we perform a simulation for peg-in-hole tasks. The results show that our method learns the appropriate skills faster than an approach that does not consider lower dimensional systems. Moreover, we demonstrate that our method works on a real robot equipped with a compliant module on the wrist. Masashi Hamaya, Robert Lee, Kazutoshi Tanaka, Felix von Drigalski, Chisato Nakashima, Yoshiya Shibata, Yoshihisa Ijiri |
ICRA | 1 |
| 2020 | MULTIPOLAR: Multi-Source Policy Aggregation for Transfer Reinforcement Learning between Diverse Environmental DynamicsabstractTransfer reinforcement learning (RL) aims at improving the learning efficiency of an agent by exploiting knowledge from other source agents trained on relevant tasks. However, it remains challenging to transfer knowledge between different environmental dynamics without having access to the source environments. In this work, we explore a new challenge in transfer RL, where only a set of source policies collected under diverse unknown dynamics is available for learning a target task efficiently. To address this problem, the proposed approach, MULTI-source POLicy AggRegation (MULTIPOLAR), comprises two key techniques. We learn to aggregate the actions provided by the source policies adaptively to maximize the target task performance. Meanwhile, we learn an auxiliary network that predicts residuals around the aggregated actions, which ensures the target policy's expressiveness even when some of the source policies perform poorly. We demonstrated the effectiveness of MULTIPOLAR through an extensive experimental evaluation across six simulated environments ranging from classic control problems to challenging robotics simulations, under both continuous and discrete action spaces. The demo videos and code are available on the project webpage: https://omron-sinicx.github.io/multipolar/. Mohammadamin Barekatain, Ryo Yonetani, Masashi Hamaya |
IJCAI | 3 |
| 2020 | A Compact, Cable-driven, Activatable Soft Wrist with Six Degrees of Freedom for Assembly TasksabstractPhysical softness has been proposed to absorb impacts when establishing contact with a robot or its workpiece, to relax control requirements and improve performance in assembly and insertion tasks. Previous work has focused on special end effector solutions for isolated tasks, such as the peg-in-hole task. However, as many robot tasks require the precision of rigid robots, and their performance would degrade when simply adding compliance, it has been difficult to take advantage of physical softness in real applications. A wrist that could switch between soft and rigid modes could solve this problem, but actuators with sufficient strength for this state transition would increase the size and weight of the module and decrease the payload of the robot. To solve this problem, we propose a novel design of a soft module consisting of a cable-driven mechanism, which allows the robot end effector to change between soft and rigid mode while being very compact and light. The module effectively combines the advantages of soft and rigid robots, and can be retrofitted to existing robots and grippers while preserving the characteristics of the robotic system. We evaluate the effectiveness of our proposed design through experiments modeling assembly tasks, and investigate design parameters quantitatively. Felix von Drigalski, Kazutoshi Tanaka, Masashi Hamaya, Robert Lee, Chisato Nakashima, Yoshiya Shibata, Yoshihisa Ijiri |
IROS | 3 |
| 2020 | Learning Soft Robotic Assembly Strategies from Successful and Failed DemonstrationsabstractPhysically soft robots are promising for robotic assembly tasks as they allow stable contacts with the environment. In this study, we propose a novel learning system for soft robotic assembly strategies. We formulate this problem as a reinforcement learning task and design the reward function from human demonstrations. Our key insight is that the failed demonstrations can be used as constraints to avoid failed behaviors. To this end, we developed a teaching device with which humans can intuitively provide various demonstrations. Moreover, we leverage Physically-Consistent Gaussian Mixture Models to clearly assign Gaussian components to the successful and failed trials. We then create the reference trajectories via Gaussian Mixture Regressions, which fit the successful demonstrations while considering the failed ones. Finally, we apply a sample- efficient deep model-based reinforcement learning method to obtain robust strategies with a few interactions. To validate our method, we developed a real-robot experimental system composed of a rigid collaborative robot arm with a compliant wrist and the teaching device. Our results demonstrated that our method learned the assembly strategies with a higher success rate than when using only successful demonstrations. Masashi Hamaya, Felix von Drigalski, Takamitsu Matsubara, Kazutoshi Tanaka, Robert Lee, Chisato Nakashima, Yoshiya Shibata, Yoshihisa Ijiri |
IROS | 1 |
| 2019 | Exploiting Human and Robot Muscle Synergies for Human-in-the-loop Optimization of EMG-based Assistive StrategiesabstractIn this study, we propose a novel human-in-the-loop optimization approach for exoskeleton robot control. We develop a method to optimize widely-used Electromyography (EMG)-based assistive strategies. If we use multiple EMG channels to control multi-DoF robots, optimization process becomes complex and requires a large amount of data. To make the optimization tractable, we exploit the synergies both of the human muscles and artificial muscles of the exoskeleton robots to reduce the number of parameters of the assistive strategies. We show that we can extract the synergies not only from the user's muscle activities but from pneumatic artificial muscle (PAMs) contractions of the exoskeleton robot. Then, we adopt a Bayesian optimization method to acquire the parameters for assisting human movements by iteratively identifying the user's preferences of the assistive strategies. We conducted experiments to evaluate our proposed method with a PAMs-driven upper-limb exoskeleton robot. Our method successfully learned assistive strategies from the human-in-theloop optimization with a practicable number of interactions. Masashi Hamaya, Takamitsu Matsubara, Jun-ichiro Furukawa, Satoshi Yagi, Tatsuya Teramae, Tomoyuki Noda, Jun Morimoto |
ICRA | 1 |
| 2017 | Learning task-parametrized assistive strategies for exoskeleton robots by multi-task reinforcement learningabstractRecent studies suggest that reinforcement learning has great potential for generating assistive strategies in exoskeletons through physical interactions between a user and a robot. Previous methods focused on a task-specific assistive strategy, where for every single task (situation/context), the user needs to interact with a robot to learn an appropriate assistive strategy. Therefore, the learned strategies cannot be generalized for a new task. Since the sampling cost is expensive for such human-in-the-loop systems as exoskeletons, generalization must be enabled. In this paper, we propose to learn task-parametrized assistive strategies for exoskeleton robots. Our method employs an assistive strategy, which depends on the task parameter and the state variable, that can be learned from multiple sets of human-robot interaction data across different tasks and generalized even for an unseen task, given the task parameter without additional learning. To alleviate the user's burden in the learning process across multiple tasks, we exploit a data-efficient multi-task reinforcement learning framework. To verify the effectiveness of our method, we developed an experimental platform with an exoskeleton robot. We conducted a series of experiments whose experimental results show that our method can learn such a task-parametrized assistive strategy and be generalized for unseen tasks to reduce the user's electromyography signals (EMGs) during tasks. Masashi Hamaya, Takamitsu Matsubara, Tomoyuki Noda, Tatsuya Teramae, Jun Morimoto |
ICRA | 1 |
| 2017 | User-robot collaborative excitation for PAM model identification in exoskeleton robotsabstractPneumatic Artificial Muscle (PAM) actuators have been used as exoskeletons because of their inherited compliance and high power-weight ratio. However, creating accurate models remains difficult mainly due to the compliance issue; the model can be changed by the force applied by the user. Therefore, both user and robot actions need to be considered for sufficient excitation of PAMs that are equipped in exoskeleton robots, unlike typical rigid actuators that can only be sufficiently excited by robot actions. In this paper, we propose a user-robot collaborative excitation approach for PAM model identification as an active learning framework for sequentially collecting data by deriving and executing optimal user and robot actions at each step with Gaussian processes. The optimal actions, which are executed by the robot, are displayed on a monitor that enables the user to execute them. We conducted experiments using a powered elbow exoskeleton with a PAM actuator. Experimental results show that our method can more efficiently identify the PAM model than a standard model identification method that does not use any data acquired through user-robot collaboration. Masashi Hamaya, Takamitsu Matsubara, Tomoyuki Noda, Tatsuya Teramae, Jun Morimoto |
IROS | 1 |
| 2017 | Learning assistive strategies for exoskeleton robots from user-robot physical interactionabstractSocial demand for exoskeleton robots that physically assist humans has been increasing in various situations due to the demographic trends of aging populations. With exoskeleton robots, an assistive strategy is a key ingredient. Since interactions between users and exoskeleton robots are bidirectional, the assistive strategy design problem is complex and challenging. In this paper, we explore a data-driven learning approach for designing assistive strategies for exoskeletons from user-robot physical interaction. We formulate the learning problem of assistive strategies as a policy search problem and exploit a data-efficient model-based reinforcement learning framework. Instead of explicitly providing the desired trajectories in the cost function, our cost function only considers the user’s muscular effort measured by electromyography signals (EMGs) to learn the assistive strategies. The key underlying assumption is that the user is instructed to perform the task by his/her own intended movements. Since the EMGs are observed when the intended movements are achieved by the user’s own muscle efforts rather than the robot’s assistance, EMGs can be interpreted as the “cost” of the current assistance. We applied our method to a 1-DoF exoskeleton robot and conducted a series of experiments with human subjects. Our experimental results demonstrated that our method learned proper assistive strategies that explicitly considered the bidirectional interactions between a user and a robot with only 60 seconds of interaction. We also showed that our proposed method can cope with changes in both the robot dynamics and movement trajectories. Masashi Hamaya, Takamitsu Matsubara, Tomoyuki Noda, Tatsuya Teramae, Jun Morimoto |
Pattern Recognit. Lett. | 1 |
| 2016 | Learning assistive strategies from a few user-robot interactions: Model-based reinforcement learning approachabstractDesigning an assistive strategy for exoskeletons is a key ingredient in movement assistance and rehabilitation. While several approaches have been explored, most studies are based on mechanical models of the human user, i.e., rigid-body dynamics or Center of Mass (CoM)-Zero Moment Point (ZMP) inverted pendulum moECenter of Massdel, or only focus on periodic movements with using oscillator models. On the other hand, the interactions between the user and the robot are often not considered explicitly because of its difficulty in modeling. In this paper, we propose to learn the assistive strategies directly from interactions between the user and the robot. We formulate the learning problem of assistive strategies as a policy search problem. To alleviate heavy burdens to the user for data acquisition, we exploit a data-efficient model-based reinforcement learning framework. To validate the effectiveness of our approach, an experimental platform composed of a real subject, an electromyography (EMG)-measurement system, and a simulated robot arm is developed. Then, a learning experiment with the assistive control task of the robot arm is conducted. As a result, proper assistive strategies that can achieve the robot control task and reduce EMG signals of the user are acquired only by 30 seconds interactions. Masashi Hamaya, Takamitsu Matsubara, Tomoyuki Noda, Tatsuya Teramae, Jun Morimoto |
ICRA | 1 |
| 2016 | Dry-wireless EEG and asynchronous adaptive feature extraction towards a plug-and-play co-adaptive brain robot interfaceabstractThis paper introduces a novel asynchronous adaptive brain machine interface (BMI), based on a dry-wireless headset, to trigger the movement of a lower limb exoskeleton robot by foot motor imagery. Specifically, it addresses two issues that are critical for the development of a plug-and-play brain robot interface (BRI): setup-time and the nonstationarity of the electroencephalogram (EEG). The former is solved by a dry-wireless headset that reduces setup-time compared to gel-based systems, and removes the nuisance of cables. The latter has been extensively studied in the literature, leading to effective adaptive algorithms in synchronous BMI. However, asynchronous BMI has received little attention. We propose an extension of state-of-the-art adaptive methods by defining the forgetting factors according to the time constant of the exponential moving average. In addition, we propose feature adaptation as opposed to the standard bias adaptation of a linear classifier. After calibrating the decoder, the subject with a reliable classification of sensorimotor rhythms was asked to trigger robot squatting. The motion was successfully initialized by foot motor imagery; with an essential contribution of the proposed adaptive BMI, which makes features less prone to nonstationarities and improves classification performance compared to standard adaptive methods. The ultimate goal of our research is to develop a plug-and-play co-adaptive BRI for neuromotor rehabilitation. Giuseppe Lisi, Masashi Hamaya, Tomoyuki Noda, Jun Morimoto |
ICRA | 2 |
| 2015 | Towards balance recovery control for lower body exoskeleton robots with Variable Stiffness Actuators: Spring-loaded flywheel modelabstractThis paper presents a biologically-inspired real-time balance recovery control strategy that is applied to a lower body exoskeleton with variable physical stiffness actuators at its ankle joints. For this purpose, a torsional spring-loaded flywheel model is presented to encapsulate both approximated angular momentum and variable physical stiffness, which are crucial parameters in describing the postural balance. In particular, the incorporation of physical compliance enables us to provide three main contributions: i) A mathematical formulation is developed to express the relation between the dynamic balance criterion ZMP and the physical ankle joint stiffness. Therefore, balancing control can be interpreted in terms of ankle joint stiffness regulation. ii) ‘Variable physical’ stiffness is utilized in the bipedal robot balance control task for the first time in the literature, to the authors' knowledge. iii) The variable physical stiffness strategy is compared with the optimal constant stiffness strategy by conducting experiments on our exoskeleton robot. The results indicate that the proposed method provides a favorable balancing control performance to cope with unperceived perturbations, in terms of center of mass position regulation, ZMP error and mechanical power. Corinne Doppmann, Barkan Ugurlu, Masashi Hamaya, Tatsuya Teramae, Tomoyuki Noda, Jun Morimoto |
ICRA | 3 |