EDBT 2026 Demo / reviewers in the wild / expert
Petar Kormushev
dblp:54/5176
· DBLP profile ↗
34ranked-venue papers
8as first author
8since 2021 · last 2025
0000-0002-6677-3044ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 8 first-author · 6 since 2021Systems, architecture and hardware · 23 · 7 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GraphGarment: Learning Garment Dynamics for Bimanual Cloth Manipulation TasksabstractPhysical manipulation of garments is often crucial when performing fabric-related tasks, such as hanging garments. However, due to the deformable nature of fabrics, these operations remain a significant challenge for robots in household, healthcare, and industrial environments. In this paper, we propose GraphGarment, a novel approach that models garment dynamics based on robot control inputs and applies the learned dynamics model to facilitate garment manipulation tasks such as hanging. Specifically, we use graphs to represent the interactions between the robot end-effector and the garment. GraphGarment uses a graph neural network (GNN) to learn a dynamics model that can predict the next garment state given the current state and input action in simulation. To address the substantial sim-to-real gap, we propose a residual model that compensates for garment state prediction errors, thereby improving real-world performance. The garment dynamics model is then applied to a model-based action sampling strategy, where it is utilized to manipulate the garment to a reference pre-hanging configuration for garment-hanging tasks. We conducted four experiments using six types of garments to validate our approach in both simulation and real-world settings. In simulation experiments, GraphGarment achieves better garment state prediction performance, with a prediction error 0.46 cm lower than the best baseline. Our approach also demonstrates improved performance in the garment-hanging simulation experiment—with enhancements of 12%, 24%, and 10%, respectively. Moreover, real-world robot experiments confirm the robustness of sim-to-real transfer, with an error increase of 0.17 cm compared to simulation results. Supplementary material is available at: https://sites.google.com/view/graphgarment. Kelin Li, Dongmyoung Lee, Xiaoshuai Chen, Rui Zong, Petar Kormushev |
IROS | 6 |
| 2025 | Haptic-ACT: Bridging Human Intuition with Compliant Robotic Manipulation via Immersive VRabstractRobotic manipulation is essential for the widespread adoption of robots in industrial and home settings and has long been a focus within the robotics community. Advances in artificial intelligence have introduced promising learning-based methods to address this challenge, with imitation learning emerging as particularly effective. However, efficiently acquiring high-quality demonstrations remains a challenge. In this work, we introduce an immersive VR-based teleoperation setup designed to collect demonstrations from a remote human user. We also propose an imitation learning framework called Haptic Action Chunking with Transformers (Haptic-ACT). To evaluate the platform, we conducted a pick-and-place task and collected 50 demonstration episodes. Results indicate that the immersive VR platform significantly reduces demonstrator fingertip forces compared to systems without haptic feedback, enabling more delicate manipulation. Additionally, evaluations of the Haptic-ACT framework in both the MuJoCo simulator and on a real robot demonstrate its effectiveness in teaching robots more compliant manipulation compared to the original ACT. Additional materials are available at https://sites.google.com/view/hapticact. Kelin Li, Shubham M. Wagh, Nitish Sharma, Saksham Bhadani, Petar Kormushev |
IROS | 7 |
| 2024 | Task Accuracy Enhancement for a Surgical Macro-Micro Manipulator With Probabilistic Neural Networks and Uncertainty MinimizationabstractAccurate robot kinematic modelling is a major component for autonomous robot control to guarantee safety and precision during task execution. In surgical robotics complex robotic structures and actuation mechanisms are generally employed, therefore machine learning techniques can be adopted to build the model of the robot. Probabilistic neural networks are a class of learning approaches that provide information about the uncertainty of the learnt models. In this work we compare two different probabilistic neural networks (Bayesian and Evidential Neural Networks) to model the kinematics of a surgical robotic instrument and propose a control strategy based on Hierarchical Quadratic Programming (HQP) capable of exploiting the model uncertainty to improve the accuracy and safety of the controller. Simulation and real world experiments on different autonomous path tracking tasks show that the model uncertainty highly affects the control performances and prove the effectiveness of the proposed controller in improving task execution.Note to Practitioners—The push towards reducing invasiveness and patient’s traumas in surgery has lead to the requirement of miniaturized and highly articulated robots. This however comes at the cost of having systems that are hard to model and control, which is one of the major limitations for autonomy in surgical robotics. Machine learning has become very effective in modelling complex systems and probabilistic approaches additionally allow estimating the confidence of the learnt model. In robotics field where high precision is required to perform an autonomous task, like in minimally invasive surgery, the robot model needs to be very accurate and controllers need to guarantee safety in performing the desired task, while satisfying additional motion constraints imposed by the application scenario. This work proposes the use of probabilistic neural networks to model the complexity of a surgical robotic instrument and a control strategy capable of ensuring safety by maximizing model’s confidence and guaranteeing satisfaction of imposed motion constraints. In this work a macro-micro manipulator setup is employed, consisting of an articulated surgical robotic instrument connected to a serial-link manipulator. The proposed modelling and control approaches can be used in any other field where controllers need to highly rely on the robot model due to limitations in using external sensors and where leveraging information about model’s confidence can be beneficial. Currently, the work focuses only on pure kinematic modelling and control, thus neglecting any possible interaction with the environment. Future work will focus on addressing this limitation in order to ensure proper force control and effective autonomy. Francesco Cursi, Weibang Bai, Eric M. Yeatman, Petar Kormushev |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2022 | Optimization of Surgical Robotic Instrument Mounting in a Macro-Micro Manipulator Setup for Improving Task ExecutionabstractIn minimally invasive robotic surgery, the surgical instrument is usually inserted inside the patient’s body through a small incision, which acts as a remote center of motion (RCM). Serial-link manipulators can be used as macro robots on which microsurgical robotic instruments are mounted to increase the number of degrees of freedom of the system and ensure safe task and RCM motion execution. However, the surgical instrument needs to be placed in an appropriate configuration when completing the motion tasks. The contribution of this article is to present a novel framework that preoperatively identifies the best base configuration, in terms of Roll, Pitch, and Yaw angles, of the microsurgical instrument with respect to the macro serial-link manipulator’s end effector in order to achieve the maximum accuracy and dexterity in performing specified tasks. The framework relies on hierarchical quadratic programming for the control, genetic algorithm for the optimization, and on a resilience to error strategy to make sure deviations from the optimum do not affect the system’s performance. Simulation results show that the mounting configuration of the surgical instrument significantly impacts the performance of the whole macro–micro manipulator in executing the desired motion tasks, and both the simulation and experimental results demonstrate that the proposed optimization method improves the overall performance. Francesco Cursi, Weibang Bai, Eric M. Yeatman, Petar Kormushev |
IEEE Trans. Robotics | 4 |
| 2021 | Policy manifold search: exploring the manifold hypothesis for diversity-based neuroevolutionabstractNeuroevolution is an alternative to gradient-based optimisation that has the potential to avoid local minima and allows parallelisation. The main limiting factor is that usually it does not scale well with parameter space dimensionality. Inspired by recent work examining neural network intrinsic dimension and loss landscapes, we hypothesise that there exists a low-dimensional manifold, embedded in the policy network parameter space, around which a high-density of diverse and useful policies are located. This paper proposes a novel method for diversity-based policy search via Neuroevolution, that leverages learned representations of the policy network parameters, by performing policy search in this learned representation space. Our method relies on the Quality-Diversity (QD) framework which provides a principled approach to policy search, and maintains a collection of diverse policies, used as a dataset for learning policy representations. Further, we use the Jacobian of the inverse-mapping function to guide the search in the representation space. This ensures that the generated samples remain in the high-density regions, after mapping back to the original space. Finally, we evaluate our contributions on four continuous-control tasks in simulated environments, and compare to diversity-based baselines. Nemanja Rakicevic, Antoine Cully, Petar Kormushev |
GECCO | 3 |
| 2021 | Learning to Represent Action Values as a Hypergraph on the Action Vertices
Arash Tavakoli, Mehdi Fatemi, Petar Kormushev |
ICLR | 3 |
| 2021 | Kalibrot: A Simple-To-Use Matlab Package for Robot Kinematic CalibrationabstractRobot modelling is an essential part to properly understand how a robotic system moves and how to control it. The kinematic model of a robot is usually obtained by using Denavit-Hartenberg convention, which relies on a set of parameters to describe the end-effector pose in a Cartesian space. These parameters are assigned based on geometrical considerations of the robotic structure, however, the assigned values may be inaccurate. The purpose of robot kinematic calibration is therefore to find optimal parameters which improve the accuracy of the robot model. In this work we present Kalibrot, an open source Matlab package for robot kinematic calibration. Kalibrot has been designed to simplify robot calibration and easily assess the calibration results. Beside computing the optimal parameters, Kalibrot provides a visualization layer showing the values of the calibrated parameters, what parameters can be identified, and the calibrated robotic structure. The capabilities of the package are here shown through simulated and real world experiments. Francesco Cursi, Weibang Bai, Petar Kormushev |
IROS | 3 |
| 2021 | Pre-operative Offline Optimization of Insertion Point Location for Safe and Accurate Surgical Task ExecutionabstractIn robotically assisted surgical procedures the surgical tool is usually inserted in the patient’s body through a small incision, which acts as a constraint for the motion of the robot, known as remote center of Motion (RCM). The location of the insertion point on the patient’s body has huge effects on the performances of the surgical robot. In this work we present an offline pre-operative framework to identify the optimal insertion point location in order to guarantee accurate and safe surgical task execution. The approach is validated using a serial-link manipulator in conjunction with a surgical robotic tool to perform a tumor resection task, while avoiding nearby organs. Results show that the framework is capable of identifying the best insertion point ensuring high dexterity, high tracking accuracy, and safety in avoiding nearby organs. Francesco Cursi, Petar Kormushev |
IROS | 2 |
| 2020 | Scaling All-Goals Updates in Reinforcement Learning Using Convolutional Neural NetworksabstractBeing able to reach any desired location in the environment can be a valuable asset for an agent. Learning a policy to navigate between all pairs of states individually is often not feasible. An all-goals updating algorithm uses each transition to learn Q-values towards all goals simultaneously and off-policy. However the expensive numerous updates in parallel limited the approach to small tabular cases so far. To tackle this problem we propose to use convolutional network architectures to generate Q-values and updates for a large number of goals at once. We demonstrate the accuracy and generalization qualities of the proposed method on randomly generated mazes and Sokoban puzzles. In the case of on-screen goal coordinates the resulting mapping from frames to distance-maps directly informs the agent about which places are reachable and in how many steps. As an example of application we show that replacing the random actions in ε-greedy exploration by several actions towards feasible goals generates better exploratory trajectories on Montezuma's Revenge and Super Mario All-Stars games. Fabio Pardo, Vitaly Levdik, Petar Kormushev |
AAAI | 3 |
| 2020 | Model Predictive Control for a Tendon-Driven Surgical Robot with Safety Constraints in Kinematics and DynamicsabstractIn fields such as minimally invasive surgery, effective control strategies are needed to guarantee safety and accuracy of the surgical task. Mechanical designs and actuation schemes have inevitable limitations such as backlash and joint limits. Moreover, surgical robots need to operate in narrow pathways, which may give rise to additional environmental constraints. Therefore, the control strategies must be capable of satisfying the desired motion trajectories and the imposed constraints. Model Predictive Control (MPC) has proven effective for this purpose, allowing to solve an optimal problem by taking into consideration the evolution of the system states, cost function, and constraints over time. The high nonlinearities in tendon-driven systems, adopted in many surgical robots, are difficult to be modelled analytically. In this work, we use a model learning approach for the dynamics of tendon-driven robots. The dynamic model is then employed to impose constraints on the torques of the robot under consideration and solve an optimal constrained control problem for trajectory tracking by using MPC. To assess the capabilities of the proposed framework, both simulated and real world experiments have been conducted. Francesco Cursi, Valerio Modugno, Petar Kormushev |
IROS | 3 |
| 2020 | Design and Control of SLIDER: An Ultra-lightweight, Knee-less, Low-cost Bipedal Walking RobotabstractMost state-of-the-art bipedal robots are designed to be anthropomorphic and therefore possess legs with knees. Whilst this facilitates more human-like locomotion, there are implementation issues that make walking with straight or near-straight legs difficult. Most bipedal robots have to move with a constant bend in the legs to avoid singularities at the knee joints, and to keep the centre of mass at a constant height for control purposes. Furthermore, having a knee on the leg increases the design complexity as well as the weight of the leg, hindering the robot's performance in agile behaviours such as running and jumping. We present SLIDER, an ultra-lightweight, low-cost bipedal walking robot with a novel knee-less leg design. This non-anthropomorphic straight-legged design reduces the weight of the legs significantly whilst keeping the same functionality as anthropomorphic legs. Simulation results show that SLIDER's low-inertia legs contribute to less vertical motion in the center of mass (CoM) than anthropomorphic robots during walking, indicating that SLIDER's model is closer to the widely used Inverted Pendulum (IP) model. Finally, stable walking on flat terrain is demonstrated both in simulation and in the physical world, and feedback control is implemented to address challenges with the physical robot. Ke Wang 0055, David Marsh, Roni Permana Saputra, Digby Chappell, Zhonghe Jiang, Akshay Raut, Bethany Kon, Petar Kormushev |
IROS | 8 |
| 2019 | Sim-to-Real Learning for Casualty Detection from Ground Projected Point Cloud DataabstractThis paper addresses the problem of human body detection-particularly a human body lying on the ground (a.k.a. casualty)-using point cloud data. This ability to detect a casualty is one of the most important features of mobile rescue robots, in order for them to be able to operate autonomously. We propose a deep-learning-based casualty detection method using a deep convolutional neural network (CNN). This network is trained to be able to detect a casualty using a point-cloud data input. In the method we propose, the point cloud input is pre-processed to generate a depth image-like ground-projected heightmap. This heightmap is generated based on the projected distance of each point onto the detected ground plane within the point cloud data. The generated heightmap-in image form-is then used as an input for the CNN to detect a human body lying on the ground. To train the neural network, we propose a novel sim-to-real approach, in which the network model is trained using synthetic data obtained in simulation and then tested on real sensor data. To make the model transferable to real data implementations, during the training we adopt specific data augmentation strategies with the synthetic training data. The experimental results show that data augmentation introduced during the training process is essential for improving the performance of the trained model on real data. More specifically, the results demonstrate that the data augmentations on raw point-cloud data have contributed to a considerable improvement of the trained model performance. Roni Permana Saputra, Nemanja Rakicevic, Petar Kormushev |
IROS | 3 |
| 2018 | Action Branching Architectures for Deep Reinforcement LearningabstractDiscrete-action algorithms have been central to numerous recent successes of deep reinforcement learning. However, applying these algorithms to high-dimensional action tasks requires tackling the combinatorial increase of the number of possible actions with the number of action dimensions. This problem is further exacerbated for continuous-action tasks that require fine control of actions via discretization. In this paper, we propose a novel neural architecture featuring a shared decision module followed by several network branches, one for each action dimension. This approach achieves a linear increase of the number of network outputs with the number of degrees of freedom by allowing a level of independence for each individual action dimension. To illustrate the approach, we present a novel agent, called Branching Dueling Q-Network (BDQ), as a branching variant of the Dueling Double Deep Q-Network (Dueling DDQN). We evaluate the performance of our agent on a set of challenging continuous control tasks. The empirical results show that the proposed agent scales gracefully to environments with increasing action dimensionality and indicate the significance of the shared decision module in coordination of the distributed action branches. Furthermore, we show that the proposed agent performs competitively against a state-of-the-art continuous control algorithm, Deep Deterministic Policy Gradient (DDPG). Arash Tavakoli, Fabio Pardo, Petar Kormushev |
AAAI | 3 |
| 2018 | Time Limits in Reinforcement LearningabstractIn reinforcement learning, it is common to let an agent interact for a fixed amount of time with its environment before resetting it and repeating the process in a series of episodes. The task that the agent has to learn can either be to maximize its performance over (i) that fixed period, or (ii) an indefinite period where time limits are only used during training to diversify experience. In this paper, we provide a formal account for how time limits could effectively be handled in each of the two cases and explain why not doing so can cause state-aliasing and invalidation of experience replay, leading to suboptimal policies and training instability. In case (i), we argue that the terminations due to time limits are in fact part of the environment, and thus a notion of the remaining time should be included as part of the agent’s input to avoid violation of the Markov property. In case (ii), the time limits are not part of the environment and are only used to facilitate learning. We argue that this insight should be incorporated by bootstrapping from the value of the state at the end of each partial episode. For both cases, we illustrate empirically the significance of our considerations in improving the performance and stability of existing reinforcement learning algorithms, showing state-of-the-art results on several control tasks. Fabio Pardo, Arash Tavakoli, Vitaly Levdik, Petar Kormushev |
ICML | 4 |
| 2017 | Climbing over large obstacles with a humanoid robot via multi-contact motion planningabstractIncremental progress in humanoid robot locomotion over the years has achieved important capabilities such as navigation over flat or uneven terrain, stepping over small obstacles and climbing stairs. However, the locomotion research has mostly been limited to using only bipedal gait and only foot contacts with the environment, using the upper body for balancing without considering additional external contacts. As a result, challenging locomotion tasks like climbing over large obstacles relative to the size of the robot have remained unsolved. In this paper, we address this class of open problems with an approach based on multi-body contact motion planning guided through physical human demonstrations. Our goal is to make the humanoid locomotion problem more tractable by taking advantage of objects in the surrounding environment instead of avoiding them. We propose a multi-contact motion planning algorithm for humanoid robot locomotion which exploits the whole-body motion and multi-body contacts including both the upper and lower body limbs. The proposed motion planning algorithm is applied to a challenging task of climbing over a large obstacle. We demonstrate successful execution of the climbing task in simulation using our multi-contact motion planning algorithm initialized via a transfer from real-world human demonstrations of the task and further optimized. Pavan Kanajar, Darwin G. Caldwell, Petar Kormushev |
RO-MAN | 3 |
| 2015 | Learning symbolic representations of actions from human demonstrationsabstractIn this paper, a robot learning approach is pro- posed which integrates Visuospatial Skill Learning, Imitation Learning, and conventional planning methods. In our approach, the sensorimotor skills (i.e., actions) are learned through a learning from demonstration strategy. The sequence of per- formed actions is learned through demonstrations using Visu- ospatial Skill Learning. A standard action-level planner is used to represent a symbolic description of the skill, which allows the system to represent the skill in a discrete, symbolic form. The Visuospatial Skill Learning module identifies the underlying constraints of the task and extracts symbolic predicates (i.e., action preconditions and effects), thereby updating the planner representation while the skills are being learned. Therefore the planner maintains a generalized representation of each skill as a reusable action, which can be planned and performed inde- pendently during the learning phase. Preliminary experimental results on the iCub robot are presented. Seyed Reza Ahmadzadeh, Ali Paikan, Fulvio Mastrogiovanni, Lorenzo Natale, Petar Kormushev, Darwin G. Caldwell |
ICRA | 5 |
| 2015 | Underwater robot-object contact perception using machine learning on force/torque sensor feedbackabstractAutonomous manipulation of objects requires reliable information on robot-object contact state. Underwater environments can adversely affect sensing modalities such as vision, making them unreliable. In this paper we investigate underwater robot-object contact perception between an autonomous underwater vehicle and a T-bar valve using a force/torque sensor and the robot's proprioceptive information. We present an approach in which machine learning is used to learn a classifier for different contact states, namely, a contact aligned with the central axis of the valve, an edge contact and no contact. To distinguish between different contact states, the robot performs an exploratory behavior that produces distinct patterns in the force/torque sensor. The sensor output forms a multidimensional time-series. A probabilistic clustering algorithm is used to analyze the time-series. The algorithm dissects the multidimensional time-series into clusters, producing a one-dimensional sequence of symbols. The symbols are used to train a hidden Markov model, which is subsequently used to predict novel contact conditions. We show that the learned classifier can successfully distinguish the three contact states with an accuracy of 72% ± 12 %. Nawid Jamali, Petar Kormushev, Arnau Carrera, Marc Carreras, Darwin G. Caldwell |
ICRA | 2 |
| 2015 | Encoderless position control of a two-link robot manipulatorabstractEncoders have been an inseparable part of robots since the very beginning of modern robotics in the 1950s. As a result, the foundations of robot control are built on the concepts of kinematics and dynamics of articulated rigid bodies, which rely on explicitly measuring the robot configuration in terms of joint angles - done by encoders. In this paper, we propose a radically new concept for controlling robots called Encoderless Robot Control (EnRoCo). The concept is based on our hypothesis that it is possible to control a robot without explicitly measuring its joint angles, by measuring instead the effects of the actuation on its end-effector. To prove the feasibility of this unconventional control approach, we propose a proof-of-concept control algorithm for encoderless position control of a robot's end-effector in task space. We demonstrate a prototype implementation of this controller in a dynamics simulation of a two-link robot manipulator. The prototype controller is able to successfully control the robot's end-effector to reach a reference position, as well as to track continuously a desired trajectory. Notably, we demonstrate how this novel controller can cope with something that traditional control approaches fail to do: adapt on-the-fly to changes in the kinematics of the robot, such as changing the lengths of the links. Petar Kormushev, Yiannis Demiris, Darwin G. Caldwell |
ICRA | 1 |
| 2015 | Kinematic-free position control of a 2-DOF planar robot armabstractThis paper challenges the well-established assumption in robotics that in order to control a robot it is necessary to know its kinematic information, that is, the arrangement of links and joints, the link dimensions and the joint positions. We propose a kinematic-free robot control concept that does not require any prior kinematic knowledge. The concept is based on our hypothesis that it is possible to control a robot without explicitly measuring its joint angles, by measuring instead the effects of the actuation on its end-effector. We implement a proof-of-concept encoderless robot controller and apply it for the position control of a physical 2-DOF planar robot arm. The prototype controller is able to successfully control the robot to reach a reference position, as well as to track a continuous reference trajectory. Notably, we demonstrate how this novel controller can cope with something that traditional control approaches fail to do: adapt to drastic kinematic changes such as 100% elongation of a link, 35-degree angular offset of a joint, and even a complete overhaul of the kinematics involving the addition of new joints and links. Petar Kormushev, Yiannis Demiris, Darwin G. Caldwell |
IROS | 1 |
| 2015 | Online regeneration of bipedal walking gait pattern optimizing footstep placement and timingabstractWe propose a new algorithm capable of online regeneration of gait patterns. The algorithm uses a nonlinear optimization technique to find step parameters that will bring the robot from the present state to a desired state. It modifies online not only the footstep positions, but also the step timing in order to maintain dynamic stability during walking. Inclusion of step time modification extends the robustness against rarely addressed disturbances, such as pushes towards the stance foot. The controller is able to recover dynamic stability regardless of the source of the disturbance (e.g. model inaccuracy, reference tracking error or external disturbance). We describe the robot state estimation and center-of-mass feedback controller necessary to realize stable locomotion on our humanoid platform COMAN. We also present a set of experiments performed on the platform that show the performance of the feedback controller and of the gait pattern regenerator. We show how the robot is able to cope with series of pushes, by adjusting step times and positions. Przemyslaw Kryczka, Petar Kormushev, Nikolaos G. Tsagarakis, Darwin G. Caldwell |
IROS | 2 |
| 2015 | Cognitive system for autonomous underwater intervention
Arnau Carrera, Narcís Palomeras, Natàlia Hurtós, Petar Kormushev, Marc Carreras |
Pattern Recognit. Lett. | 4 |
| 2014 | Multi-objective reinforcement learning for AUV thruster failure recoveryabstractThis paper investigates learning approaches for discovering fault-tolerant control policies to overcome thruster failures in Autonomous Underwater Vehicles (AUV). The proposed approach is a model-based direct policy search that learns on an on-board simulated model of the vehicle. When a fault is detected and isolated the model of the AUV is reconfigured according to the new condition. To discover a set of optimal solutions a multi-objective reinforcement learning approach is employed which can deal with multiple conflicting objectives. Each optimal solution can be used to generate a trajectory that is able to navigate the AUV towards a specified target while satisfying multiple objectives. The discovered policies are executed on the robot in a closed-loop using AUV's state feedback. Unlike most existing methods which disregard the faulty thruster, our approach can also deal with partially broken thrusters to increase the persistent autonomy of the AUV. In addition, the proposed approach is applicable when the AUV either becomes under-actuated or remains redundant in the presence of a fault. We validate the proposed approach on the model of the Girona500 AUV. Seyed Reza Ahmadzadeh, Petar Kormushev, Darwin G. Caldwell |
ADPRL | 2 |
| 2014 | Online discovery of AUV control policies to overcome thruster failuresabstractWe investigate methods to improve fault-tolerance of Autonomous Underwater Vehicles (AUVs) to increase their reliability and persistent autonomy. We propose a learning-based approach that is able to discover new control policies to overcome thruster failures as they happen. The proposed approach is a model-based direct policy search that learns on an on-board simulated model of the AUV. The model is adapted to a new condition when a fault is detected and isolated. Since the approach generates an optimal trajectory, the learned fault-tolerant policy is able to navigate the AUV towards a specified target with minimum cost. Finally, the learned policy is executed on the real robot in a closed-loop using the state feedback of the AUV. Unlike most existing methods which rely on the redundancy of thrusters, our approach is also applicable when the AUV becomes under-actuated in the presence of a fault. To validate the feasibility and efficiency of the presented approach, we evaluate it with three learning algorithms and three policy representations with increasing complexity. The proposed method is tested on a real AUV, Girona500. Seyed Reza Ahmadzadeh, Matteo Leonetti, Arnau Carrera, Marc Carreras, Petar Kormushev, Darwin G. Caldwell |
ICRA | 5 |
| 2014 | Robot-object contact perception using symbolic temporal pattern learningabstractThis paper investigates application of machine learning to the problem of contact perception between a robot's gripper and an object. The input data comprises a multidimensional time-series produced by a force/torque sensor at the robot's wrist, the robot's proprioceptive information, namely, the position of the end-effector, as well as the robot's control command. These data are used to train a hidden Markov model (HMM) classifier. The output of the classifier is a prediction of the contact state, which includes no contact, a contact aligned with the central axis of the valve, and an edge contact. To distinguish between contact states, the robot performs exploratory behaviors that produce distinct patterns in the time-series data. The patterns are discovered by first analyzing the data using a probabilistic clustering algorithm that transforms the multidimensional data into a one-dimensional sequence of symbols. The symbols produced by the clustering algorithm are used to train the HMM classifier. We examined two exploratory behaviors: a rotation around the x-axis, and a rotation around the y-axis of the gripper. We show that using these two exploratory behaviors we can successfully predict a contact state with an accuracy of 88 ± 5 % and 81 ± 10 %, respectively. Nawid Jamali, Petar Kormushev, Darwin G. Caldwell |
ICRA | 2 |
| 2014 | Haptic exploration of unknown surfaces with discontinuitiesabstractThis work presents an approach for exploring unknown surfaces with discontinuities using only force/torque information. The motivation is to build an information map of an unknown object or environment by performing a fully-autonomous haptic exploration. Examples of discontinuities considered here are contours with sharp turns (such as wall corners) and abrupt dips (such as cliffs). Compliant motion control using force information has the ability to conform to unknown, smooth surfaces but not to discontinuous surfaces. This paper investigates solutions to address the limitation in compliant motion control over discontinuities while maintaining a desired normal force along the surface. We propose two methods to address the problem: (1) superposition of motion and force control and (2) rotation of axes for force and motion control. The theoretical principles are discussed and experimental results with a KUKA lightweight arm moving in 2D space are presented. Both approaches successfully negotiate objects with sharp 90-degree and 120-degree turns while still maintaining good tracking of the desired force. Rodrigo S. Jamisola, Petar Kormushev, Antonio Bicchi, Darwin G. Caldwell |
IROS | 2 |
| 2013 | Autonomous robotic valve turning: A hierarchical learning approachabstractAutonomous valve turning is an extremely challenging task for an Autonomous Underwater Vehicle (AUV). To resolve this challenge, this paper proposes a set of different computational techniques integrated in a three-layer hierarchical scheme. Each layer realizes specific subtasks to improve the persistent autonomy of the system. In the first layer, the robot acquires the motor skills of approaching and grasping the valve by kinesthetic teaching. A Reactive Fuzzy Decision Maker (RFDM) is devised in the second layer which reacts to the relative movement between the valve and the AUV, and alters the robot's movement accordingly. Apprenticeship learning method, implemented in the third layer, performs tuning of the RFDM based on expert knowledge. Although the long-term goal is to perform the valve turning task on a real AUV, as a first step the proposed approach is tested in a laboratory environment. Seyed Reza Ahmadzadeh, Petar Kormushev, Darwin G. Caldwell |
ICRA | 2 |
| 2013 | Visuospatial skill learning for object reconfiguration tasksabstractWe present a novel robot learning approach based on visual perception that allows a robot to acquire new skills by observing a demonstration from a tutor. Unlike most existing learning from demonstration approaches, where the focus is placed on the trajectories, in our approach the focus is on achieving a desired goal configuration of objects relative to one another. Our approach is based on visual perception which captures the object's context for each demonstrated action. This context is the basis of the visuospatial representation and encodes implicitly the relative positioning of the object with respect to multiple other objects simultaneously. The proposed approach is capable of learning and generalizing multi-operation skills from a single demonstration, while requiring minimum a priori knowledge about the environment. The learned skills comprise a sequence of operations that aim to achieve the desired goal configuration using the given objects. We illustrate the capabilities of our approach using three object reconfiguration tasks with a Barrett WAM robot. Seyed Reza Ahmadzadeh, Petar Kormushev, Darwin G. Caldwell |
IROS | 2 |
| 2013 | On-line identification of autonomous underwater vehicles through global derivative-free optimizationabstractWe describe the design and implementation of an on-line identification scheme for Autonomous Underwater Vehicles (AUVs). The proposed method estimates the dynamic parameters of the vehicle based on a global derivative-free optimization algorithm. It is not sensitive to initial conditions, unlike other on-line identification schemes, and does not depend on the differentiability of the model with respect to the parameters. The identification scheme consists of three distinct modules: a) System Excitation, b) Metric Calculator and c) Optimization Algorithm. The System Excitation module sends excitation inputs to the vehicle. The Optimization Algorithm module calculates a candidate parameter vector, which is fed to the Metric Calculator module. The Metric Calculator module evaluates the candidate parameter vector, using a metric based on the residual of the actual and the predicted commands. The predicted commands are calculated utilizing the candidate parameter vector and the vehicle state vector, which is available via a complete navigation module. Then, the metric is directly fed back to the Optimization Algorithm module, and it is used to correct the estimated parameter vector. The procedure continues iteratively until the convergence properties are met. The proposed method is generic, demonstrates quick convergence and does not require a linear formulation of the model with respect to the parameter vector. The applicability and performance of the proposed algorithm is experimentally verified using the AUV Girona 500. George C. Karras, Charalampos P. Bechlioulis, Matteo Leonetti, Narcís Palomeras, Petar Kormushev, Kostas J. Kyriakopoulos, Darwin G. Caldwell |
IROS | 5 |
| 2013 | Improving the energy efficiency of autonomous underwater vehicles by learning to model disturbancesabstractEnergy efficiency is one of the main challenges for long-term autonomy of AUVs (Autonomous Underwater Vehicles). We propose a novel approach for improving the energy efficiency of AUV controllers based on the ability to learn which external disturbances can safely be ignored. The proposed learning approach uses adaptive oscillators that are able to learn online the frequency, amplitude and phase of zero-mean periodic external disturbances. Such disturbances occur naturally in open water due to waves, currents, and gravity, but also can be caused by the dynamics and hydrodynamics of the AUV itself. We formulate the theoretical basis of the approach, and demonstrate its abilities on a number of input signals. Further experimental evaluation is conducted using a dynamic model of the Girona 500 AUV in simulation on two important underwater scenarios: hovering and trajectory tracking. The proposed approach shows significant energy-saving capabilities while at the same time maintaining high controller gains. The approach is generic and applicable not only for AUV control, but also for other type of control where periodic disturbances exist and could be accounted for by the controller. Petar Kormushev, Darwin G. Caldwell |
IROS | 1 |
| 2012 | Challenges for the policy representation when applying reinforcement learning in roboticsabstractA summary of the state-of-the-art reinforcement learning in robotics is given, in terms of both algorithms and policy representations. Numerous challenges faced by the policy representation in robotics are identified. Two recent examples for application of reinforcement learning to robots are described: pancake flipping task and bipedal walking energy minimization task. In both examples, a state-of-the-art Expectation-Maximization-based reinforcement learning algorithm is used, but different policy representations are proposed and evaluated for each task. The two proposed policy representations offer viable solutions to four rarely-addressed challenges in policy representations: correlations, adaptability, multi-resolution, and globality. Both the successes and the practical difficulties encountered in these examples are discussed. Petar Kormushev, Sylvain Calinon, Darwin G. Caldwell, Barkan Ugurlu |
IJCNN | 1 |
| 2012 | The anatomy of a fall: Automated real-time analysis of raw force sensor data from bipedal walking robots and humansabstractAn automated approach is proposed which can analyze ground reaction force data from bipedal walking robots and humans. The input of the automated analysis is the raw data from force sensors mounted in the feet of a robot. The output is detailed information, such as detected single support, double support, and swing phases, their durations, timings of events like heel strikes, properties of the phase transitions and of the robot itself. The proposed approach is generic, parameter-free, model-free, robust, computationally efficient, and applicable for real-time use during walking. It can detect early indications of instability that could lead to a fall of the robot. Three real-world experiments are presented: with a compliant bipedal robot, with a stiff humanoid robot, and with a human subject. Petar Kormushev, Barkan Ugurlu, Luca Colasanto, Nikolaos G. Tsagarakis, Darwin G. Caldwell |
IROS | 1 |
| 2011 | Upper-body kinesthetic teaching of a free-standing humanoid robotabstractWe present an integrated approach allowing a free-standing humanoid robot to acquire new motor skills by kinesthetic teaching. The proposed method controls simultaneously the upper and lower body of the robot with different control strategies. Imitation learning is used for training the upper body of the humanoid robot via kinesthetic teaching, while at the same time Reaction Null Space method is used for keeping the balance of the robot. During demonstration, a force/torque sensor is used to record the exerted forces, and during reproduction, we use a hybrid position/force controller to apply the learned trajectories in terms of positions and forces to the end effector. The proposed method is tested on a 25-DOF Fujitsu HOAP-2 humanoid robot with a surface cleaning task. Petar Kormushev, Dragomir N. Nenchev, Sylvain Calinon, Darwin G. Caldwell |
ICRA | 1 |
| 2011 | Bipedal walking energy minimization by reinforcement learning with evolving policy parameterizationabstractWe present a learning-based approach for minimizing the electric energy consumption during walking of a passively-compliant bipedal robot. The energy consumption is reduced by learning a varying-height center-of-mass trajectory which uses efficiently the robot's passive compliance. To do this, we propose a reinforcement learning method which evolves the policy parameterization dynamically during the learning process and thus manages to find better policies faster than by using fixed parameterization. The method is first tested on a function approximation task, and then applied to the humanoid robot COMAN where it achieves significant energy reduction. Petar Kormushev, Barkan Ugurlu, Sylvain Calinon, Nikolaos G. Tsagarakis, Darwin G. Caldwell |
IROS | 1 |
| 2010 | Robot motor skill coordination with EM-based Reinforcement LearningabstractWe present an approach allowing a robot to acquire new motor skills by learning the couplings across motor control variables. The demonstrated skill is first encoded in a compact form through a modified version of Dynamic Movement Primitives (DMP) which encapsulates correlation information. Expectation-Maximization based Reinforcement Learning is then used to modulate the mixture of dynamical systems initialized from the user's demonstration. The approach is evaluated on a torque-controlled 7 DOFs Barrett WAM robotic arm. Two skill learning experiments are conducted: a reaching task where the robot needs to adapt the learned movement to avoid an obstacle, and a dynamic pancake-flipping task. Petar Kormushev, Sylvain Calinon, Darwin G. Caldwell |
IROS | 1 |