Dongheui Lee

dblp:96/5115 · DBLP profile ↗
← Back
80ranked-venue papers
13as first author
18since 2021 · last 2025
0000-0003-1897-7664ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 74 · 13 first-author · 15 since 2021Systems, architecture and hardware · 59 · 13 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 since 2021
YearPublicationVenuePosition
2025 Multimodal Anomaly Detection with a Mixture-of-Experts
abstract
With a growing number of robots being deployed across diverse applications, robust multimodal anomaly detection becomes increasingly important. In robotic manipulation, failures typically arise from (1) robot-driven anomalies due to an insufficient task model or hardware limitations, and (2) environment-driven anomalies caused by dynamic environmental changes or external interferences. Conventional anomaly detection methods focus either on the first by low-level statistical modeling of proprioceptive signals or the second by deep learning-based visual environment observation, each with different computational and training data requirements. To effectively capture anomalies from both sources, we propose a mixture-of-experts framework that integrates the complementary detection mechanisms with a visual-language model for environment monitoring and a Gaussian-mixture regression-based detector for tracking deviations in interaction forces and robot motions. We introduce a confidence-based fusion mechanism that dynamically selects the most reliable detector for each situation. We evaluate our approach on both household and industrial tasks using two robotic systems, demonstrating a 60% reduction in detection delay while improving frame-wise anomaly detection performance compared to individual detectors.
Christoph Willibald, Daniel Sliwowski, Dongheui Lee
IROS3
2024 A Unified Masked Autoencoder with Patchified Skeletons for Motion Synthesis
abstract
The synthesis of human motion has traditionally been addressed through task-dependent models that focus on specific challenges, such as predicting future motions or filling in intermediate poses conditioned on known key-poses. In this paper, we present a novel task-independent model called UNIMASK-M, which can effectively address these challenges using a unified architecture. Our model obtains comparable or better performance than the state-of-the-art in each field. Inspired by Vision Transformers (ViTs), our UNIMASK-M model decomposes a human pose into body parts to leverage the spatio-temporal relationships existing in human motion. Moreover, we reformulate various pose-conditioned motion synthesis tasks as a reconstruction problem with different masking patterns given as input. By explicitly informing our model about the masked joints, our UNIMASK-M becomes more robust to occlusions. Experimental results show that our model successfully forecasts human motion on the Human3.6M dataset while achieving state-of-the-art results in motion inbetweening on the LaFAN1 dataset for long transition periods.
Esteve Valls Mascaro, Hyemin Ahn 0001, Dongheui Lee
AAAI3
2024 Shared Autonomy via Variable Impedance Control and Virtual Potential Fields for Encoding Human Demonstrations
abstract
This article introduces a framework for complex human-robot collaboration tasks, such as the co-manufacturing of furniture. For these tasks, it is essential to encode tasks from human demonstration and reproduce these skills in a compliant and safe manner. Therefore, two key components are addressed in this work: motion generation and shared autonomy. We propose a motion generator based on a time-invariant potential field, capable of encoding wrench profiles, complex and closed-loop trajectories, and additionally incorporates obstacle avoidance. Additionally, the paper addresses shared autonomy (SA) which enables synergetic collaboration between human operators and robots by dynamically allocating authority. Variable impedance control (VIC) and force control are employed, where impedance and wrench are adapted based on the human-robot autonomy factor derived from interaction forces. System passivity is ensured by an energy-tank based task passivation strategy. The framework’s efficacy is validated through simulations and an experimental study employing a Franka Emika Research 3 robot.
Shail V. Jadav, Johannes Heidersberger, Christian Ott 0001, Dongheui Lee
ICRA4
2024 Robot Interaction Behavior Generation based on Social Motion Forecasting for Human-Robot Interaction
abstract
Integrating robots into populated environments is a complex challenge that requires an understanding of human social dynamics. In this work, we propose to model social motion forecasting in a shared human-robot representation space, which facilitates us to synthesize robot motions that interact with humans in social scenarios despite not observing any robot in the motion training. We develop a transformer-based architecture called ECHO, which operates in the aforementioned shared space to predict the future motions of the agents encountered in social scenarios. Contrary to prior works, we reformulate the social motion problem as the refinement of the predicted individual motions based on the surrounding agents, which facilitates the training while allowing for single-motion forecasting when only one human is in the scene. We evaluate our model in multi-person and human-robot motion forecasting tasks and obtain state-of-the-art performance by a large margin while being efficient and performing in real-time. Additionally, our qualitative results showcase the effectiveness of our approach in generating human-robot interaction behaviors that can be controlled via text commands.
Esteve Valls Mascaro, Yashuai Yan, Dongheui Lee
ICRA3
2024 Is a Simulation better than Teleoperation for Acquiring Human Manipulation Skill Data?
abstract
This study explores the feasibility of using simulations as a better interface to collect human object manipulation skills for learning from demonstrations (LfD). Recently, numerous researchers have started introducing teleoperation systems to acquire human manipulation skills. However, capturing the subtle, force-involved interaction skills of humans in teleoperation is still challenging due to its inherent dynamic delays and feedback transparency. This research evaluates the effectiveness of demonstration data obtained through simulation versus teleoperation. To evaluate the efficacy of this approach, tasks such as plane cutting, tight peg-in-hole, and deformable pipe plugging were performed to assess the quality of demonstrations acquired. The experimental results highlight the effectiveness of demonstration through simulation in capturing the operator’s force-involved interaction skills. Simulation creates an environment similar to performing tasks with bare hands by minimising dynamic delays due to the exclusion of physical robots and effectively rendering high stiffness. As a result, the demonstration through simulation method has proven effective in extracting interaction data and capturing physical task performance skills.
Kwang-Hyun Lee, Dongheui Lee, Jee-Hwan Ryu
IROS4
2024 Learning From Demonstration of Robot Motions And Stiffness Behaviors For Surgical Blunt Dissection
abstract
In this work, we present a learning from demonstration solution for automating a surgical blunt dissection task. In addition to learning motion trajectories, our goal is to learn variable impedance behaviors that enable the robot to interact safely and compliantly during the task. To that end, we propose a teaching interface using bilateral teleoperation, which allows the natural transfer of human motions and impedance behaviors skills to robots. The demonstrated profiles are captured with Dynamic Movement Primitives and Gaussian Mixture Models, which subsequently provide the robot a reference motion plan, and a stiffness adaptation policy, during physical interaction. Experimental validation on real robot hardware shows the effectiveness of the proposed approach in terms of ensuring successful task execution, as well as safety compared to stiff high-gain control.
Riccardo Arduini, Youssef Michel, Harsimran Singh, Julian Klodmann, Dongheui Lee
RO-MAN5
2024 A Passivity-Based Approach for Variable Stiffness Control With Dynamical Systems
abstract
In this paper, we present a controller that combines motion generation and control in one loop, to endow robots with reactivity and safety. In particular, we propose a control approach that enables to follow the motion plan of a first order Dynamical System (DS) with a variable stiffness profile, in a closed loop configuration where the controller is always aware of the current robot state. This allows the robot to follow a desired path with an interactive behavior dictated by the desired stiffness. We also present two solutions to enable a robot to follow the desired velocity profile, in a manner similar to trajectory tracking controllers, while maintaining the closed-loop configuration. Additionally, we exploit the concept of energy tanks in order to guarantee the passivity during interactions with the environment, as well as the asymptotic stability in free motion, of our closed-loop system. The developed approach is evaluated extensively in simulation, as well as in real robot experiments, in terms of performance and safety both in free motion and during the execution of physical interaction tasks.Note to Practitioners—The approach presented in this work allows for safe and reactive robot motions, as well as the capacity to shape the robot’s physical behavior during interactions. This becomes crucial for performing contact tasks that might require adaptability or for interactions with humans as in shared control or collaborative tasks. Furthermore, the reactive properties of our controller make it adequate for robots that operate in proximity to humans or in dynamic environments where potential collisions are likely to happen.
Youssef Michel, Matteo Saveriano, Dongheui Lee
IEEE Trans Autom. Sci. Eng.3
2023 Can We Use Diffusion Probabilistic Models for 3D Motion Prediction?
abstract
After many researchers observed fruitfulness from the recent diffusion probabilistic model, its effectiveness in image generation is actively studied these days. In this paper, our objective is to evaluate the potential of diffusion probabilistic models for 3D human motion-related tasks. To this end, this pa-per presents a study of employing diffusion probabilistic models to predict future 3D human motion(s) from the previously observed motion. Based on the Human 3.6M and HumanEva-I datasets, our results show that diffusion probabilistic models are competitive for both single (deterministic) and multiple (stochastic) 3D motion prediction tasks, after finishing a single training process. In addition, we find out that diffusion probabilistic models can offer an attractive compromise, since they can strike the right balance between the likelihood and diversity of the predicted future motions. Our code is publicly available on the project website: https://sites.google.com/view/diffusion-motion-prediction.
Hyemin Ahn 0001, Esteve Valls Mascaro, Dongheui Lee
ICRA3
2023 Orientation Control with Variable Stiffness Dynamical Systems
abstract
Recently, several approaches have attempted to combine motion generation and control in one loop to equip robots with reactive behaviors, that cannot be achieved with traditional time-indexed tracking controllers. These approaches however mainly focused on positions, neglecting the orientation part which can be crucial to many tasks e.g. screwing. In this work, we propose a control algorithm that adapts the robot's rotational motion and impedance in a closed-loop manner. Given a first-order Dynamical System representing an orientation motion plan and a desired rotational stiffness profile, our approach enables the robot to follow the reference motion with an interactive behavior specified by the desired stiffness, while always being aware of the current orientation, represented as a Unit Quaternion (UQ). We rely on the Lie algebra to formulate our algorithm, since unlike positions, UQ feature constraints that should be respected in the devised controller. We validate our proposed approach in multiple robot experiments, showcasing the ability of our controller to follow complex orientation profiles, react safely to perturbations, and fulfill physical interaction tasks.
Youssef Michel, Matteo Saveriano, Fares J. Abu-Dakka, Dongheui Lee
IROS4
2023 Fusing Visual Appearance and Geometry for Multi-Modality 6DoF Object Tracking
abstract
In many applications of advanced robotic manipulation, six degrees of freedom (6DoF) object pose estimates are continuously required. In this work, we develop a multi-modality tracker that fuses information from visual appearance and geometry to estimate object poses. The algorithm extends our previous method ICG, which uses geometry, to additionally consider surface appearance. In general, object surfaces contain local characteristics from text, graphics, and patterns, as well as global differences from distinct materials and colors. To incorporate this visual information, two modalities are developed. For local characteristics, keypoint features are used to minimize distances between points from keyframes and the current image. For global differences, a novel region approach is developed that considers multiple regions on the object surface. In addition, it allows the modeling of external geometries. Experiments on the YCB-Video and OPT datasets demonstrate that our approach ICG+ performs best on both datasets, outperforming both conventional and deep learning-based methods. At the same time, the algorithm is highly efficient and runs at more than 300 Hz. The source code of our tracker is publicly available.
Manuel Stoiber, Mariam Elsayed, Anne E. Reichert, Florian Steidle, Dongheui Lee, Rudolph Triebel
IROS5
2023 Intention-Conditioned Long-Term Human Egocentric Action Anticipation
abstract
To anticipate how a person would act in the future, it is essential to understand the human intention since it guides the subject towards a certain action. In this paper, we propose a hierarchical architecture which assumes a sequence of human action (low-level) can be driven from the human intention (high-level). Based on this, we deal with long-term action anticipation task in egocentric videos. Our framework first extracts this low- and high-level human information over the observed human actions in a video through a Hierarchical Multi-task Multi-Layer Perceptrons Mixer (H3M). Then, we constrain the uncertainty of the future through an Intention-Conditioned Variational Auto-Encoder (I-CVAE) that generates multiple stable predictions of the next actions that the observed human might perform. By leveraging human intention as high-level information, we claim that our model is able to anticipate more time-consistent actions in the long-term, thus improving the results over the baseline in Ego4D dataset. This work results in the state-of-the-art for Long-Term Anticipation (LTA) task in Ego4D by providing more plausible anticipated sequences, improving the anticipation scores of nouns and actions. Our work ranked first in both CVPR@2022 and ECCV@2022 Ego4D LTA Challenge.
Esteve Valls Mascaro, Hyemin Ahn 0001, Dongheui Lee
WACV3
2023 Human-object interaction prediction in videos through gaze following
Zhifan Ni, Esteve Valls Mascaro, Hyemin Ahn 0001, Dongheui Lee
Comput. Vis. Image Underst.4
2023 Estimation of 6D Pose of Objects Based on a Variant Adversarial Autoencoder
Hyemin Ahn 0001, Shile Li, Dongheui Lee
Neural Process. Lett.5
2022 Visually Grounding Language Instruction for History-Dependent Manipulation
abstract
This paper emphasizes the importance of a robot's ability to refer to its task history, especially when it exe-cutes a series of pick-and-place manipulations by following language instructions given one by one. The advantage of referring to the manipulation history can be categorized into two folds: (1) the language instructions omitting details but using expressions referring to the past can be interpreted, and (2) the visual information of objects occluded by previous manipulations can be inferred. For this, we introduce a history-dependent manipulation task which objective is to visually ground a series of language instructions for proper pick-and-place manipulations by referring to the past. We also suggest a relevant dataset and model which can be a baseline, and show that our model trained with the proposed dataset can also be applied to the real world based on the CycleGAN. Our dataset and code are publicly available on the project website: https://sites.google.com/view/history-dependent-manipulation.
Hyemin Ahn 0001, Obin Kwon, Kyungdo Kim, Jaeyeon Jeong, Howoong Jun, Hongjung Lee, Dongheui Lee, Songhwai Oh
ICRA7
2022 Robust Human Motion Forecasting using Transformer-based Model
abstract
Comprehending human motion is a fundamental challenge for developing Human-Robot Collaborative applications. Computer vision researchers have addressed this field by only focusing on reducing error in predictions, but not taking into account the requirements to facilitate its implementation in robots. In this paper, we propose a new model based on Transformer that simultaneously deals with the real time 3D human motion forecasting in the short and long term. Our 2-Channel Transformer (2CH-TR) is able to efficiently exploit the spatio-temporal information of a shortly observed sequence (400ms) and generates a competitive accuracy against the current state-of-the-art. 2CH-TR stands out for the efficient performance of the Transformer, being lighter and faster than its competitors. In addition, our model is tested in conditions where the human motion is severely occluded, demonstrating its robustness in reconstructing and predicting 3D human motion in a highly noisy environment. Our experiment results show that the proposed 2CH-TR outperforms the ST-Transformer, which is another state-of-the-art model based on the Transformer, in terms of reconstruction and prediction under the same conditions of input prefix. Our model reduces in 8.89% the mean squared error of ST-Transformer in short-term prediction, and 2.57% in long-term prediction in Human3.6M dataset with 400ms input prefix.
Esteve Valls Mascaro, Hyemin Ahn 0001, Dongheui Lee
IROS4
2022 Multi-Level Task Learning Based on Intention and Constraint Inference for Autonomous Robotic Manipulation
abstract
To perform tasks in unstructured environments, robots need to be able to apply learned skills to different contexts and to autonomously make decisions online. We, therefore, developed a novel data-driven task learning approach that segments a task demonstration into simpler skills and structures them in a high-level task graph. In contrast to other state-of-the-art methods, the presented approach can not only infer the low-level skills and their respective subgoals but also multimodal feature constraints fitted individually to each skill. The inferred feature constraints allow to detect anomalies during autonomous task execution, which can be automatically resolved by a recovery behavior of the task graph. The subgoals encode each skill's intention and thereby enable to flexibly transition between skills and to generalize the behavior to new setups. By separating the subgoal and constraint inference, we achieve a reduced computational complexity and an increased performance compared to state-of-the-art task learning approaches. In a real-world manipulation task, we demonstrate the reusability of skills as well as the autonomous decision-making of our approach.
Christoph Willibald, Dongheui Lee
IROS2
2022 Safety-Aware Hierarchical Passivity-Based Variable Compliance Control for Redundant Manipulators
abstract
This article presents a hierarchical passivity-based compliance controller that exploits robot redundancy and aims at achieving an impedance behavior with a time-varying stiffness on all the priority levels. Unfortunately, this gives rise to certain control actions that lead to the loss of the safety-critical passivity feature. To deal with this problem, we employ the concept of virtualenergy tanksthat keep track of the passivity violating energy in the system, ensuring that it remains bounded. This restores the passivity in the system, which guarantees the stable interaction with any passive environment. Furthermore, we augment our controller with an additional safety layer, which ensures that the energy injected through the tank into the system remains below a safe limit, defined based on the maximum kinetic energy allowed in the system. Finally, our approach is validated in terms of performance during task execution and safety both in simulations and on real-robot hardware.
Youssef Michel, Christian Ott 0001, Dongheui Lee
IEEE Trans. Robotics3
2021 Refining Action Segmentation with Hierarchical Video Representations
abstract
In this paper, we propose Hierarchical Action Segmentation Refiner (HASR), which can refine temporal action segmentation results from various models by understanding the overall context of a given video in a hierarchical way. When a backbone model for action segmentation estimates how the given video can be segmented, our model extracts segment-level representations based on frame-level features, and extracts a video-level representation based on the segment-level representations. Based on these hierarchical representations, our model can refer to the overall context of the entire video, and predict how the segment labels that are out of context should be corrected. Our HASR can be plugged into various action segmentation models (MS-TCN, SSTDA, ASRF), and improve the performance of state-of-the-art models based on three challenging datasets (GTEA, 50Salads, and Breakfast). For example, in 50Salads dataset, the segmental edit score improves from 67.9% to 77.4% (MS-TCN), from 75.8% to 77.3% (SSTDA), from 79.3% to 81.0% (ASRF). In addition, our model can refine the segmentation result from the unseen backbone model, which was not referred to when training HASR. This generalization performance would make HASR be an effective tool for boosting up the existing approaches for temporal action segmentation. Our code is available at https://github.com/cotton-ahn/HASR_iccv2021.
Hyemin Ahn 0001, Dongheui Lee
ICCV2
2020 Measuring Generalisation to Unseen Viewpoints, Articulations, Shapes and Objects for 3D Hand Pose Estimation Under Hand-Object Interaction
Anil Armagan, Guillermo Garcia-Hernando, Seungryul Baek, Shreyas Hampali, Mahdi Rad, Shipeng Xie, Mingxiu Chen, Boshen Zhang, Fu Xiong, Yang Xiao 0007, Zhiguo Cao 0001, Junsong Yuan 0001, Pengfei Ren 0001, Weiting Huang, Haifeng Sun 0001, Marek Hrúz, Jakub Kanis, Zdenek Krnoul, Qingfu Wan, Shile Li, Linlin Yang 0001, Dongheui Lee, Angela Yao, Weiguo Zhou, Sijia Mei, Adrian Spurr, Umar Iqbal 0001, Pavlo Molchanov 0001, Philippe Weinzaepfel, Romain Brégier, Grégory Rogez, Vincent Lepetit, Tae-Kyun Kim 0001
ECCV (23)23
2020 Mini-Batched Online Incremental Learning Through Supervisory Teleoperation with Kinesthetic Coupling
abstract
We propose an online incremental learning approach through teleoperation which allows an operator to partially modify a learned model, whenever it is necessary, during task execution. Compared to conventional incremental learning approaches, the proposed approach is applicable for teleoperation-based teaching and it needs only partial demonstration without any need to obstruct the task execution. Dynamic authority distribution and kinesthetic coupling between the operator and the agent helps the operator to correctly perceive the exact instance where modification needs to be asserted in the agent's behaviour online using partial trajectory. For this, we propose a variation of the Expectation-Maximization algorithm for updating original model through mini batches of the modified partial trajectory. The proposed approach reduces human workload and latency for a rhythmic peg-in-hole teleoperation task where online partial modification is required during the task operation.
Hiba Ovais Latifee, Affan Pervez, Jee-Hwan Ryu, Dongheui Lee
ICRA4
2020 Hand Pose Estimation for Hand-Object Interaction Cases using Augmented Autoencoder
abstract
Hand pose estimation with objects is challenging due to object occlusion and the lack of large annotated datasets. To tackle these issues, we propose an Augmented Autoencoder based deep learning method using augmented clean hand data. Our method takes 3D point cloud of a hand with an augmented object as input and encodes the input to latent representation of the hand. From the latent representation, our method decodes 3D hand pose and we propose to use an auxiliary point cloud decoder to assist the formation of the latent space. Through quantitative and qualitative evaluation on both synthetic dataset and real captured data containing objects, we demonstrate state-of-the-art performance for hand pose estimation with objects, even using only a small number of annotated hand-object samples.
Shile Li, Dongheui Lee
ICRA3
2020 Collaborative Programming of Conditional Robot Tasks
abstract
Conventional robot programming methods are not suited for non-experts to intuitively teach robots new tasks. For this reason, the potential of collaborative robots for production cannot yet be fully exploited. In this work, we propose an active learning framework, in which the robot and the user collaborate to incrementally program a complex task. Starting with a basic model, the robot's task knowledge can be extended over time if new situations require additional skills. An on-line anomaly detection algorithm therefore automatically identifies new situations during task execution by monitoring the deviation between measured- and commanded sensor values. The robot then triggers a teaching phase, in which the user decides to either refine an existing skill or demonstrate a new skill. The different skills of a task are encoded in separate probabilistic models and structured in a high-level graph, guaranteeing robust execution and successful transition between skills. In the experiments, our approach is compared to two state-of-the-art Programming by Demonstration frameworks on a real system. Increased intuitiveness and task performance of the method can be shown, allowing shop-floor workers to program industrial tasks with our framework.
Christoph Willibald, Thomas Eiband, Dongheui Lee
IROS3
2020 Hand Pose-based Task Learning from Visual Observations with Semantic Skill Extraction
abstract
Learning from Demonstrations is a promising technique to transfer task knowledge from a user to a robot. We propose a framework for task programming by observing the human hand pose and object locations solely with a depth camera. By extracting skills from the demonstrations, we are able to represent what the robot has learned, generalize to unseen object locations and optimize the robotic execution instead of replaying a non-optimal behavior. A two-staged segmentation algorithm that employs skill template matching via Hidden Markov Models has been developed to extract motion primitives from the demonstration and gives them semantic meanings. In this way, the transfer of task knowledge has been improved from a simple replay of the demonstration towards a semantically annotated, optimized and generalized execution. We evaluated the extraction of a set of skills in simulation and prove that the task execution can be optimized by such means.
Zeju Qiu, Thomas Eiband, Shile Li, Dongheui Lee
RO-MAN4
2019 Point-To-Pose Voting Based Hand Pose Estimation Using Residual Permutation Equivariant Layer
abstract
Recently, 3D input data based hand pose estimation methods have shown state-of-the-art performance, because 3D data capture more spatial information than the depth image. Whereas 3D voxel-based methods need a large amount of memory, PointNet based methods need tedious preprocessing steps such as K-nearest neighbour search for each point. In this paper, we present a novel deep learning hand pose estimation method for an unordered point cloud. Our method takes 1024 3D points as input and does not require additional information. We use Permutation Equivariant Layer (PEL) as the basic element, where a residual network version of PEL is proposed for the hand pose estimation task. Furthermore, we propose a voting-based scheme to merge information from individual points to the final pose output. In addition to the pose estimation task, the voting-based scheme can also provide point cloud segmentation result without ground-truth for segmentation. We evaluate our method on both NYU dataset and the Hands2017Challenge dataset, where our method outperforms recent state-of-the-art methods.
Shile Li, Dongheui Lee
CVPR2
2019 Aligning Latent Spaces for 3D Hand Pose Estimation
abstract
Hand pose estimation from monocular RGB inputs is a highly challenging task. Many previous works for monocular settings only used RGB information for training despite the availability of corresponding data in other modalities such as depth maps. In this work, we propose to learn a joint latent representation that leverages other modalities as weak labels to boost the RGB-based hand pose estimator. By design, our architecture is highly flexible in embedding various diverse modalities such as heat maps, depth maps and point clouds. In particular, we find that encoding and decoding the point cloud of the hand surface can improve the quality of the joint latent representation. Experiments show that with the aid of other modalities during training, our proposed method boosts the accuracy of RGB-based hand pose estimation systems and significantly outperforms state-of-the-art on two public benchmarks.
Linlin Yang 0001, Shile Li, Dongheui Lee, Angela Yao
ICCV3
2019 Learning Haptic Exploration Schemes for Adaptive Task Execution
abstract
The recent generation of compliant robots enables kinesthetic teaching of novel skills by human demonstration. This enables strategies to transfer tasks to the robot in a more intuitive way than conventional programming interfaces. Programming physical interactions can be achieved by manually guiding the robot to learn the behavior from the motion and force data. To let the robot react to changes in the environment, force sensing can be used to identify constraints and act accordingly. While autonomous exploration strategies in the whole workspace are time consuming, we propose a way to learn these schemes from human demonstrations in an object targeted manner. The presented teaching strategy and the learning framework allow to generate adaptive robot behaviors relying on the robot's sense of touch in a systematically changing environment. A generated behavior consists of a hierarchical representation of skills, where haptic exploration skills are used to touch the environment with the end effector, and relative manipulation skills, which are parameterized according to previous exploration events. The effectiveness of the approach has been proven in a manipulation task, where the adaptive task structure is able to generalize to unseen object locations. The robot autonomously manipulates objects without relying on visual feedback.
Thomas Eiband, Matteo Saveriano, Dongheui Lee
ICRA3
2019 Merging Position and orientation Motion Primitives
abstract
In this paper, we focus on generating complex robotic trajectories by merging sequential motion primitives. A robotic trajectory is a time series of positions and orientations ending at a desired target. Hence, we first discuss the generation of converging pose trajectories via dynamical systems, providing a rigorous stability analysis. Then, we present approaches to merge motion primitives which represent both the position and the orientation part of the motion. Developed approaches preserve the shape of each learned movement and allow for continuous transitions among succeeding motion primitives. Presented methodologies are theoretically described and experimentally evaluated, showing that it is possible to generate a smooth pose trajectory out of multiple motion primitives.
Matteo Saveriano, Felix Franzel, Dongheui Lee
ICRA3
2019 Learning Barrier Functions for Constrained Motion Planning with Dynamical Systems
abstract
Stable dynamical systems are a flexible tool to plan robotic motions in real-time. In the robotic literature, dynamical system motions are typically planned without considering possible limitations in the robot's workspace. This work presents a novel approach to learn workspace constraints from human demonstrations and to generate motion trajectories for the robot that lie in the constrained workspace. Training data are incrementally clustered into different linear subspaces and used to fit a low dimensional representation of each subspace. By considering the learned constraint subspaces as zeroing barrier functions, we are able to design a control input that keeps the system trajectory within the learned bounds. This control input is effectively combined with the original system dynamics preserving eventual asymptotic properties of the unconstrained system. Simulations and experiments on a real robot show the effectiveness of the proposed approach.
Matteo Saveriano, Dongheui Lee
IROS2
2019 A Transfer Learning Approach to Cross-Modal Object Recognition: From Visual Observation to Robotic Haptic Exploration
abstract
In this paper, we introduce the problem of cross-modal visuo-tactile object recognition with robotic active exploration. With this term, we mean that the robot observes a set of objects with visual perception, and later on, it is able to recognize such objects only with tactile exploration, without having touched any object before. Using a machine learning terminology, in our application, we have a visual training set and a tactile test set, or vice versa. To tackle this problem, we propose an approach constituted by four steps: finding a visuo-tactile common representation, defining a suitable set of features, transferring the features across the domains, and classifying the objects. We show the results of our approach using a set of 15 objects, collecting 40 visual examples and five tactile examples for each object. The proposed approach achieves an accuracy of 94.7%, which is comparable with the accuracy of the monomodal case, i.e., when using visual data both as training set and test set. Moreover, it performs well compared to the human ability, which we have roughly estimated carrying out an experiment with ten participants.
Pietro Falco, Shuang Lu, Ciro Natale, Salvatore Pirozzi, Dongheui Lee
IEEE Trans. Robotics5
2018 Depth-Based 3D Hand Pose Estimation: From Current Achievements to Future Goals
abstract
In this paper, we strive to answer two questions: What is the current state of 3D hand pose estimation from depth images? And, what are the next challenges that need to be tackled? Following the successful Hands In the Million Challenge (HIM2017), we investigate the top 10 state-of-the-art methods on three tasks: single frame 3D pose estimation, 3D hand tracking, and hand pose estimation during object interaction. We analyze the performance of different CNN structures with regard to hand shape, joint visibility, view point and articulation distributions. Our findings include: (1) isolated 3D hand pose estimation achieves low mean errors (10 mm) in the view point range of [70, 120] degrees, but it is far from being solved for extreme view points; (2) 3D volumetric representations outperform 2D CNNs, better capturing the spatial structure of the depth data; (3) Discriminative methods still generalize poorly to unseen hand shapes; (4) While joint occlusions pose a challenge for most methods, explicit modeling of structure constraints can significantly narrow the gap between errors on visible and occluded joints.
Shanxin Yuan, Guillermo Garcia-Hernando, Björn Stenger, Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee, Pavlo Molchanov 0001, Jan Kautz, Sina Honari, Liuhao Ge, Junsong Yuan 0001, Xinghao Chen 0001, Guijin Wang, Fan Yang 0032, Kai Akiyama, Yang Wu 0001, Qingfu Wan, Meysam Madadi, Sergio Escalera, Shile Li, Dongheui Lee, Iasonas Oikonomidis, Antonis A. Argyros, Tae-Kyun Kim 0001
CVPR21
2018 Incremental Skill Learning of Stable Dynamical Systems
abstract
Efficient skill acquisition, representation, and online adaptation to different scenarios has become of fundamental importance for assistive robotic applications. In the past decade, dynamical systems (DS) have arisen as a flexible and robust tool to represent learned skills and to generate motion trajectories. This work presents a novel approach to incrementally modify the dynamics of a generic autonomous DS when new demonstrations of a task are provided. A control input is learned from demonstrations to modify the trajectory of the system while preserving the stability properties of the reshaped DS. Learning is performed incrementally through Gaussian process regression, increasing the robot's knowledge of the skill every time a new demonstration is provided. The effectiveness of the proposed approach is demonstrated with experiments on a publicly available dataset of complex motions.
Matteo Saveriano, Dongheui Lee
IROS2
2017 Cross-modal visuo-tactile object recognition using robotic active exploration
abstract
In this work, we propose a framework to deal with cross-modal visuo-tactile object recognition. By cross-modal visuo-tactile object recognition, we mean that the object recognition algorithm is trained only with visual data and is able to recognize objects leveraging only tactile perception. The proposed cross-modal framework is constituted by three main elements. The first is a unified representation of visual and tactile data, which is suitable for cross-modal perception. The second is a set of features able to encode the chosen representation for classification applications. The third is a supervised learning algorithm, which takes advantage of the chosen descriptor. In order to show the results of our approach, we performed experiments with 15 objects common in domestic and industrial environments. Moreover, we compare the performance of the proposed framework with the performance of 10 humans in a simple cross-modal recognition task.
Pietro Falco, Shuang Lu, Andrea Cirillo, Ciro Natale, Salvatore Pirozzi, Dongheui Lee
ICRA6
2017 Data-efficient control policy search using residual dynamics learning
abstract
In this work, we propose a model-based and data efficient approach for reinforcement learning. The main idea of our algorithm is to combine simulated and real rollouts to efficiently find an optimal control policy. While performing rollouts on the robot, we exploit sensory data to learn a probabilistic model of the residual difference between the measured state and the state predicted by a simplified model. The simplified model can be any dynamical system, from a very accurate system to a simple, linear one. The residual difference is learned with Gaussian processes. Hence, we assume that the difference between real and simplified model is Gaussian distributed, which is less strict than assuming that the real system is Gaussian distributed. The combination of the partial model and the learned residuals is exploited to predict the real system behavior and to search for an optimal policy. Simulations and experiments show that our approach significantly reduces the number of rollouts needed to find an optimal control policy for the real system.
Matteo Saveriano, Yuchao Yin, Pietro Falco, Dongheui Lee
IROS4
2017 Special issue on user profiling and behavior adaptation for human-robot interaction
Silvia Rossi 0002, Dongheui Lee
Pattern Recognit. Lett.2
2017 Human-aware motion reshaping using dynamical systems
Matteo Saveriano, Fabian Hirt, Dongheui Lee
Pattern Recognit. Lett.3
2016 Encoding human actions with a frequency domain approach
abstract
In this work, we propose a Frequency-based Action Descriptor (FADE) to represent human actions. In robotics, with the development of Programming by Demonstration (PbD) methods, representing and recognizing large sets of actions has become crucial to build autonomous systems that learn from humans. The FADE descriptor leverages Fast Fourier Transform (FFT) for action representation and is combined with the Manhattan distance for measuring similarities between actions. It is characterized by a low time and space complexity and is particularly suitable for classification of human actions. For clustering problems, we propose a modified version of FADE, called Uncompressed-FADE (U-FADE), which performs well in combination with Spectral Clustering algorithms at the price of a reduced compression. We compare FADE with action descriptors based on Singular Value Decomposition (SVD) and Hidden Markov Models (HMM) on the entire HDM05 motion capture database. Despite the high dimensionality of the problem, we obtained on the entire database a promising recognition rate of 78% combining FADE with a simple 1-NN classification algorithm. Furthermore, we achieved a rate of 98% on a small action set and 88% on a medium action set.
Dharmil Shah, Pietro Falco, Matteo Saveriano, Dongheui Lee
IROS4
2016 Learning and Generalization of Compensative Zero-Moment Point Trajectory for Biped Walking
abstract
This paper presents an online learning framework for improving the robustness of zero-moment point (ZMP)-based biped walking controllers. The key idea is to learn a feedforward compensative ZMP (CZMP) trajectory from measured ZMP errors during repetitive walking motions by applying iterative learning control theory. The learned CZMP trajectory adjusts the reference ZMP and reduces the effect of unmodeled dynamics at the pattern-generation stage. From individual learned CZMP trajectories of typical walking parameters, we can build up a CZMP database. This database can be used for generating an initial CZMP whenever a new walking pattern is executed. A prediction from the database is done by k-nearest neighbor regression based on the Mahalonobis distance. Compared with state-of-the-art model-based methods, the proposed learning approach is model free and allows online adaptation to constant unknown disturbances. Enhanced walking robustness can be observed from reduced average ZMP error and more robust reaction against external disturbances on the DLR humanoid robot TORO.
Kai Hu 0009, Christian Ott 0001, Dongheui Lee
IEEE Trans. Robotics3
2015 Prioritized Inverse Kinematics with Multiple Task Definitions
abstract
We are proposing a general framework that incorporates multiple task definitions in the prioritized inverse kinematics problem. First, a mathematical description of multiple task definitions is constructed that provides an efficient way to show unprioritized or prioritized accumulations of tasks. Then, smooth transitions between all task definitions are studied, so a method, called task transition control, is developed that interpolates joint trajectories using barycentric coordinates and linear dynamical systems to overcome difficulties of interpolating task trajectories in the conventional methods. Consequently, smooth, arbitrary, and consecutive task transitions are achieved in a simple, direct, and general manner and also boundedness of joint trajectories is assured regardless of singularity. Lastly, the idea is tested by two kinematic simulations: obstacle avoidance with the KUKA LWR and task scheduling of a humanoid robot.
Sang-ik An, Dongheui Lee
ICRA2
2015 Online iterative learning control of zero-moment point for biped walking stabilization
abstract
Biped walking control based on simplified models relies much on online feedback stabilizers to compensate the zero-moment point (ZMP) error which partially comes from the model inconsistency of pattern generation. Inspired by the fact that human improves the performance by practicing a task for multiple times, this paper presents an online learning control framework for improving the robustness during the dominant repetitive phases of walking. The key idea is to learn a compensative feedforward ZMP term from previous ZMP error trajectories in order to achieve better ZMP tracking. Based on the iterative learning control theory, the learning process is conducted online continuously with minimal iteration of two footsteps, which can practically run in parallel with state-of-the-art walking controllers. A varying forgetting factor is designed to reduce the influence of the landing impact. Convergence of the learning control algorithm and improved ZMP tracking performance is verified both in dynamics simulation and experiment on the DLR humanoid robot TORO.
Kai Hu 0009, Christian Ott 0001, Dongheui Lee
ICRA3
2015 Incremental kinesthetic teaching of end-effector and null-space motion primitives
abstract
In this paper, we propose a unified approach to teach and iteratively refine both end-effector and null-space movements. Hence, the robot can be taught to make use of all its degrees-of-freedom (DoF) to adapt its behavior to new dynamic scenarios. In order to achieve this goal we propose an incremental learning approach in a framework of kinesthetic teaching based on a multi-priority kinematic controller, the so-called Task Transition Control (TTC). The learning algorithm is responsible for skill acquisition and their incremental update. On the real-time level, end-effector and null-space motion primitives, as well as the physical guidance are considered as prioritized tasks. The transitions among these tasks and their insertion and removal are managed by the TTC according to the specified transition parameters. This allows to introduce a customized task which guarantees a proper and smooth response to the applied external forces during the kinesthetic teaching. Experimental results on a 7 DoF KUKA lightweight manipulator show the effectiveness of the proposed approach.
Matteo Saveriano, Sang-ik An, Dongheui Lee
ICRA3
2015 A bidirectional invariant representation of motion for gesture recognition and reproduction
abstract
Human action representation, recognition and learning is of importance to guarantee a fruitful human-robot cooperation. In this paper, we propose a novel coordinate-free, scale invariant representation of 6D (position and orientation) motion trajectories. The advantages of the proposed invariant representation are twofold. First the performance of gesture recognition can be improved thanks to its invariance to different viewpoints and different body sizes of the actors. Secondly, the proposed representation is bi-directional. Not only the original Cartesian trajectory can be converted into the 6 invariant values, but also the motion in the original space can be retrieved back from the invariants. While the former aspect handles robust human gesture recognition, the latter allows the execution of robot motions without the need to store the Cartesian data. Experimental results illustrate the effectiveness of the proposed invariant representation for gesture recognition and accurate trajectory reconstruction.
Raffaele Soloperto, Matteo Saveriano, Dongheui Lee
ICRA3
2015 Real-time and model-free object tracking using particle filter with Joint Color-Spatial Descriptor
abstract
This paper presents a novel point-cloud descriptor for robust and real-time tracking of multiple objects without any object knowledge. Following with the framework of incremental model-free multiple object tracking from our previous work [5][7][6], 6 DoF pose of each object is firstly estimated with input point-cloud data which is then segmented according to the estimated objects, and incremental model of each object is updated from the segmented point-clouds. Here, we propose Joint Color-Spatial Descriptor (JCSD) to enhance the robustness of the pose hypothesis evaluation to the point-cloud scene in the particle filtering framework. The method outperforms widely used point-to-point comparison methods, especially in the partially occluded scene, which is frequently happened in the dynamic object manipulation cases. By means of the robust descriptor, we achieved unsupervised multiple object segmentation accuracy higher than 99%. The model-free multiple object tracking was implemented by using a particle filtering with JCSD as a likelihood function. The robust likelihood function is implemented with GPU, thus facilitating real-time tracking of multiple objects.
Shile Li, Seong-Yong Koo, Dongheui Lee
IROS3
2015 Generalization of optimal motion trajectories for bipedal walking
abstract
Control of robot locomotion profits from the use of pre-planned trajectories. This paper presents a way to generalize globally optimal and dynamically consistent trajectories for cyclic bipedal walking. A small task-space consisting of stride-length and step time is mapped to spline parameters which fully define the optimal joint space motion. The paper presents the impact of different machine learning algorithms for velocity and torque optimal trajectories with respect to optimality and feasibility. To demonstrate the usefulness of the trajectories, a control approach is presented that allows general walking including transitions between points in the task-space.
Alexander Werner, Dietrich Trautmann, Dongheui Lee, Roberto Lampariello
IROS3
2014 Prioritized inverse kinematics using QR and cholesky decompositions
abstract
This paper proposes new methods for the prioritized inverse kinematics (PIK) by using the QR decomposition (QRD) and the Cholesky decomposition (CLD) on the purpose of separation between orthogonalization and inversion processes that are essential parts of the PIK. The distinctive approach eliminates the interference between two processes which usually induces inaccuracy and sometimes instability on the prioritized inverse solutions. Two degenerate properties of using the QRD are explained and the remedies are provided as the modified damped least-squares pseudoinverse and the numerical reconditioning. The effectiveness are examined by the kinematic simulations with the n-link manipulators in the two-dimensional case and the KUKA LWR in the three-dimensional case.
Sang-ik An, Dongheui Lee
ICRA2
2014 Online human walking imitation in task and joint space based on quadratic programming
abstract
This paper presents an online methodology for imitating human walking motion of a humanoid robot in task and joint space simultaneously. Two aspects are essential for a successful walking imitation: stable footprints represented in task space and motion similarity represented in joint space. The human footprints are recognized from the captured motion data and imitated by the robot through conventional zero-moment point (ZMP) control scheme. Additionally we focus on similar knee joint trajectories for the motion similarity, which are related to knee stretching and swing leg motion. The inverse kinematics suffers from three problems: knee singularity, strongly conflicting tasks and underactuation. We formulate this problem as a quadratic programming (QP) with dynamic equality and inequality constraints. The discontinuity of dynamic task switching is solved by introducing an activation buffer, resulting in a cascaded QP form. Finally we evaluate the effectiveness of the proposed approach on the DLR humanoid robot TORO.
Kai Hu 0009, Christian Ott 0001, Dongheui Lee
ICRA3
2014 Distance based dynamical system modulation for reactive avoidance of moving obstacles
abstract
An algorithm which allows the robot to avoid moving obstacles and to reach the assigned goal is proposed. For this purpose, a dynamical system (DS) modulation matrix is calculated using the distance from the obstacles and their velocity, without the need of an analytical representation of the obstacles. This matrix modulates a generic first order DS, used to generate the desired path, saving the equilibrium points of the modulated system. The effectiveness of the proposed approach is validated with numerical simulations and experiments on a 7 DOF KUKA light weight arm.
Matteo Saveriano, Dongheui Lee
ICRA2
2014 Unsupervised object individuation from RGB-D image sequences
abstract
In this paper, we propose a novel unified framework for unsupervised object individuation from RGB-D image sequences. The proposed framework integrates existing location-based and feature-based object segmentation methods to achieve both computational efficiency and robustness in unstructured and dynamic situations. Based on the infant's object indexing theory, the newly proposed ambiguity graph plays as a key component of the framework to detect falsely segmented objects and rectify them by using both location and feature information. In order to evaluate the proposed method, three table-top multiple object manipulation scenarios were performed: stacking, unstacking, and occluding tasks. The results showed that the proposed method is more robust than the location-only method and more efficient than the feature-only method.
Seong-Yong Koo, Dongheui Lee, Dong-Soo Kwon
IROS2
2014 A Bayesian approach for task recognition and future human activity prediction
abstract
Task recognition and future human activity prediction are of importance for a safe and profitable human-robot cooperation. In real scenarios, the robot has to extract this information merging the knowledge of the task with contextual information from the sensors, minimizing possible misunderstandings. In this paper, we focus on tasks that can be represented as a sequence of manipulated objects and performed actions. The task is modelled with a Dynamic Bayesian Network (DBN), which takes as input manipulated objects and performed actions. Objects and actions are separately classified starting from RGB-D raw data. The DBN is responsible for estimating the current task, predicting the most probable future pairs of action-object and correcting possible misclassification. The effectiveness of the proposed approach is validated on a case of study, consisting of three typical tasks of a kitchen scenario.
Vito Magnanimo, Matteo Saveriano, Silvia Rossi 0002, Dongheui Lee
RO-MAN4
2014 Incremental object learning and robust tracking of multiple objects from RGB-D point set data
Seong-Yong Koo, Dongheui Lee, Dong-Soo Kwon
J. Vis. Commun. Image Represent.2
2013 GMM-based 3D object representation and robust tracking in unconstructed dynamic environments
abstract
Operating in unstructured dynamic human environments, it is desirable for a robot to identify dynamic objects and robustly track them without prior knowledge. This paper proposes a novel model-free approach for probabilistic representation and tracking of moving objects from 3D point set data based on Gaussian Mixture Model (GMM). GMM is inherently flexible such that represents any shape of objects as 3D probability distribution of the true positions. In order to achieve the robustness of the model, the proposed tracking method consists of GMM-based 3D registration, Gaussian Sum Filtering, and GMM simplification processes. The tracking performance of the proposed method was evaluated in the moving two human hands with one object, and it performed over 87% tracking accuracy together with processing 5 frames per second.
Seong-Yong Koo, Dongheui Lee, Dong-Soo Kwon
ICRA2
2013 Multiple object tracking using an RGB-D camera by hierarchical spatiotemporal data association
abstract
In this paper, we propose a novel multiple object tracking method from RGB-D point set data by introducing the hierarchical spatiotemporal data association method (HSTA) in order to robustly track multiple objects without prior knowledge. HSTA is able to construct not only temporal associations between multiple objects, but also component-level spatiotemporal associations that allow the correction of falsely detected objects in the presence of various types of interaction among multiple objects. The proposed method was evaluated using the four representative interaction cases such as split, complete occlusion, partial occlusion, and multiple contacts. As a result, HSTA showed significantly more robust performance than did other temporal data association methods in the experiments.
Seong-Yong Koo, Dongheui Lee, Dong-Soo Kwon
IROS2
2013 Kinesthetic teaching of humanoid motion based on whole-body compliance control with interaction-aware balancing
abstract
In this work we present a framework for kinesthetic teaching and iterative refinement of whole body motions. For detection of external forces we apply a momentum based disturbance observer known from manipulator control to the floating-base model of a humanoid robot. These external forces are used as a trigger for implementing a compliant behavior at the interaction point and are integrated into a predictive balancing algorithm. For representation of the motion data, a hidden Markov model is used, which allows for an iterative update of the discrete motion states as well as a smooth generation of continuous motion data. Finally, we present an application of these algorithms on the humanoid robot TORO.
Christian Ott 0001, Bernd Henze, Dongheui Lee
IROS3
2013 Point cloud based dynamical system modulation for reactive avoidance of convex and concave obstacles
abstract
The ability of the robot to avoid undesired collisions with humans and objects in its workspace is of importance in the field of human-robot interaction. In this paper, we propose an algorithm which allows the robot to avoid obstacles and to reach the assigned goal as long as the goal does not lie within obstacles. For this purpose, dynamical system modulation approach is adopted which ensures the avoidance of convex and concave obstacles. A modulation matrix can be calculated directly from the point cloud data of obstacles in the scene, without the need of analytical representation of the obstacles. This matrix modulates a generic first order dynamical system, used to generate the goal. In this way we guarantee the obstacles avoidance and the reaching of the goal. The effectiveness of the proposed approach is validated with numerical simulations and experiments on a 7 DOF KUKA light weight arm.
Matteo Saveriano, Dongheui Lee
IROS2
2013 Invariant representation for user independent motion recognition
abstract
Human gesture recognition is of importance for smooth and efficient human robot interaction. One of difficulties in gesture recognition is that different actors have different styles in performing even same gestures. In order to move towards more realistic scenarios, a robot is required to handle not only different users, but also different view points and noisy incomplete data from onboard sensors on the robot. Facing these challenges, we propose a new invariant representation of rigid body motions, which is invariant to translation, rotation and scaling factors. For classification, Hidden Markov Models based approach and Dynamic Time Warping based approach are modified by weighting the importances of body parts. The proposed method is tested with two Kinect datasets and it is compared with another invariant representation and a typical non-invariant representation. The experimental results show good recognition performance of our proposed approach.
Matteo Saveriano, Dongheui Lee
RO-MAN2
2012 Risk-Sensitive Optimal Feedback Control for Haptic Assistance
abstract
While human behavior prediction can increase the capability of a robotic partner to generate anticipatory behavior during physical human robot interaction (pHRI), predictions in uncertain situations can lead to large disturbances for the human if they do not match the human intentions. In this paper we present a novel control concept in which the assistive control parameters are adapted to the uncertainty in the sense that a the robot takes a more or less active role depending on its confidence in the human behavior prediction. The approach is based on risk-sensitive optimal feedback control. The human behavior is modeled using probabilistic learning methods and any unexpected disturbance is considered as a source of noise. The proposed approach is validated in situations with different uncertainties, process noise and risk-sensitivities in a tow- Degree-of-Freedom virtual reality experiment.
Jose Ramon Medina, Dongheui Lee, Sandra Hirche
ICRA2
2012 Learning and generalizing force control policies for sculpting
abstract
Humans exhibit exceptional skills in using tools and manipulating objects of their environment by skillfully controlling exerted force and arm impedance. One of the basic components of this mechanism is the generation of internal models which associate kinematic variables with applied force. On the other hand, making robots capable of skillfully using tools and adapting their motor behavior to new environmental conditions is rather complex. In the present paper, we investigate learning of force control policies for robotic sculpting given multiple task demonstrations. These policies express the relationship between constrained motions and exerted force and are learned in Cartesian space where the coupling of dynamics between different directions of motion is also taken into account. In addition, a novel algorithm is proposed to generalize these policies to new motion tasks, executed in a sufficiently homogeneous environment, same with that in demonstrations, but in presence of new motion-dependent external forces. To this aim, a differential calculus approach is proposed where not only the mapping from motion to force but also from difference in motion to difference in force is learned to generalize the policies to new contexts. This is achieved by learning apart from a set of policy parameters, some newly introduced quantities, so called weight differentials, which express the rate of change of the policy parameters. The proposed approach is validated in simple real-world sculpting experiments by using a two degrees-of-freedom haptic device.
Vasiliki Koropouli, Sandra Hirche, Dongheui Lee
IROS3
2012 Feedback motion planning and learning from demonstration in physical robotic assistance: differences and synergies
abstract
Goal-directed physical assistance to the human is one of the most challenging problems in the area of human-robot interaction. Planning and learning from demonstration represent two conceptually different approaches to achieve goal-directed behavior. Here we examine the properties of a planning-based and a learning-based approach in the context of physical robotic assistance for the prototypical task of cooperative object maneuvering. In order to exploit the complementary strengths of planning and learning-based approaches we derive three novel synergy strategies. The algorithms are experimentally evaluated in a human user study in a planar virtual-reality scenario and in a proof-of-concept study with a human-sized mobile robot with two 7DoF arms. The results show that combinations of planning and learning algorithms are superior over the individual approaches.
Martin Lawitzky, Jose Ramon Medina, Dongheui Lee, Sandra Hirche
IROS3
2012 Disagreement-aware physical assistance through risk-sensitive optimal feedback control
abstract
Proactive physical robotic assistance in the presence of human prediction uncertainty is a very challenging control problem. In this paper we propose a risk-sensitive optimal feedback controller for physical assistance that autonomously adapts the robot's behavior even during unknown situations. Using a probabilistic model to represent the cooperative task execution behavior and modeling the human as a source of process noise in the system, the proposed assistive controller proactively contributes to the task anticipating the human motion. Estimating online the current level of disagreement and prediction uncertainty, the assistive controller consequently calculates the optimal task contribution providing higher adaptability. A psychological evaluation compares different assistive control strategies in a virtual scenario using a two-Degree-of-Freedom haptic experimental setup. Results show that considering the current level of disagreement enhances the performance of the controller in terms of helpfulness and human effort minimization.
Jose Ramon Medina, Tamara Lorenz, Dongheui Lee, Sandra Hirche
IROS3
2012 Tire mounting on a car using the real-time control architecture ARCADE
abstract
In comparison to industrial settings with structured environments, the operation of autonomous robots in unstructured and uncertain environments is more challenging. This video presents a generic control and system architecture ARCADE, applicable for real-time robot control in complex task situations. Several methods to cope with uncertainties are demonstrated with the example task of changing tires on a car. Approaches of object detection (applied to car, tires, and humans), robust real-time control of robot arms under perception uncertainty, and human-friendly haptic interaction are detailed. The video shows two robots jointly performing the task of mounting a mock-up tire to a real car using the proposed methods, realizing robust performance in an uncertain environment.
Thomas Nierhoff, Lei Lou, Vasiliki Koropouli, Martin Eggers, Timo Fritzsch, Omiros Kourakos, Kolja Kühnlenz, Dongheui Lee, Bernd Radig, Martin Buss, Sandra Hirche
IROS8
2012 Real-time human motion tracking using multiple depth cameras
abstract
In this paper, we consider the problem of tracking human motion with a 22-DOF kinematic model from depth images. In contrast to existing approaches, our system naturally scales to multiple sensors. The motivation behind our approach, termed Multiple Depth Camera Approach (MDCA), is that by using several cameras, we can significantly improve the tracking quality and reduce ambiguities as for example caused by occlusions. By fusing the depth images of all available cameras into one joint point cloud, we can seamlessly incorporate the available information from multiple sensors into the pose estimation. To track the high-dimensional human pose, we employ state-of-the-art annealed particle filtering and partition sampling. We compute the particle likelihood based on the truncated signed distance of each observed point to a parameterized human shape model. We apply a coarse-to-fine scheme to recognize a wide range of poses to initialize the tracker. In our experiments, we demonstrate that our approach can accurately track human motion in real-time (15Hz) on a GPGPU. In direct comparison to two existing trackers (OpenNI, Microsoft Kinect SDK), we found that our approach is significantly more robust for unconstrained motions and under (partial) occlusions.
Licong Zhang, Jürgen Sturm, Daniel Cremers, Dongheui Lee
IROS4
2012 Towards interactive physical robotic assistance: Parameterizing motion primitives through natural language
abstract
Natural language interaction between humans and robots is a very challenging topic, especially when it refers to motion descriptions in a certain environment. This problem is particularly relevant during physical human-robot interaction, e.g. in cooperative transportation tasks, where the partners' physical coupling requires an agreement on the way to follow. Understanding in depth the link between sentences, words, environmental properties and motions can deeply enhance the interaction between humans and robots. In this work, we propose a novel approach for learning relations and dependencies between motion, natural language and environmental properties using parameterized left-to-right time-based Hidden Markov Models. A natural language model represents the link between language and motion symbols while the HMMs parameterization corresponds to the explicit influence on motions of both words and environmental features. The proposed PHMM approach parameterizes the output and the transition probabilities using a non-linear dependency estimation. The method is validated by learning and generating navigation primitives in a 2 Degrees-Of-Freedom (DoF) virtual scenario.
Jose Ramon Medina, Michael Shelley, Dongheui Lee, Wataru Takano, Sandra Hirche
RO-MAN3
2011 Physical human robot interaction in imitation learning
abstract
This video presents our recent research on the integration of physical human-robot interaction (pHRI) into imitation learning. First, a marker control approach for real time human motion imitation is shown. Secondly, physical coaching in addition to observational learning is applied for the incremental learning of motion primitives. Last, we extend imitation learning to learning pHRI which includes the establishment of intended physical contacts. The proposed methods were implemented and tested using the IRT humanoid robot and DLR's humanoid upper-body robot Justin.
Dongheui Lee, Christian Ott 0001, Yoshihiko Nakamura, Gerd Hirzinger
ICRA1
2011 Learning interaction control policies by demonstration
abstract
This paper explores learning of interaction force skills by human demonstration in dynamic interaction tasks. Skillful force regulation is required in many cases to achieve the goal of a task and at the same time, not to cause undesired stress on the manipulator or the object under manipulation which could result in physical failure. For example, manipulation of compliant objects with varying physical properties or artistic tasks such as engraving require skillful force modulation. Humans gracefully manipulate objects by using their sense of touch and skillfully regulating exerted forces. To learn the demonstrated force for a task by demonstration, an interaction force control policy, in terms of a goal-directed dynamical system, is proposed which stems from the parallel force/position control. The control policy is parameterized and its parameters are learned by Locally Weighted Regression from human demonstrated data to learn a force trajectory. Scaling of learned force is possible by modifying the goal of the system. The proposed method is evaluated in virtual manipulation tasks using a two degrees-of-freedom haptic device.
Vasiliki Koropouli, Dongheui Lee, Sandra Hirche
IROS2
2011 Particle filter based monocular human tracking with a 3D cardbox model and a novel deterministic resampling strategy
abstract
The challenge of markerless human motion tracking is the high dimensionality of the search space. Thus, efficient exploration in the search space is of great significance. In this paper, a motion capturing algorithm is proposed for upper body motion tracking. The proposed system tracks human motion based on monocular silhouette-matching, and it is built on the top of a hierarchical particle filter, within which a novel deterministic resampling strategy (DRS) is applied. The proposed system is evaluated quantitatively with the ground truth data measured by an inertial sensor system. In addition, we compare the DRS with the stratified resampling strategy (SRS). It is shown in experiments that DRS outperforms SRS with the same amount of particles. Moreover, a new 3D articulated human upper body model with the name 3D cardbox model is created and is proven to work successfully for motion tracking. Experiments show that the proposed system can robustly track upper body motion without self-occlusion. Motions towards the camera can also be well tracked.
Dongheui Lee, Wolfgang Sepp
IROS2
2011 An experience-driven robotic assistant acquiring human knowledge to improve haptic cooperation
abstract
Physical cooperation with humans greatly enhances the capabilities of robotic systems when leaving standardized industrial settings. Our novel cognition-enabled control framework presented in this paper enables a robotic assistant to enrich its own experience by acquisition of human task knowledge during joint manipulation. Our robot incrementally learns semantic task structures during joint task execution using hierarchically clustered Hidden Markov Models. A semantic labeling of recognized task segments is acquired from the human partner through speech. After a small number of repetitions, the robot uses an anticipated task progress to generate a feed-forward set point for an admittance feedback control scheme. This paper describes the framework and its implementation on a mobile bi-manual platform. The evolution of the robot's task knowledge is presented and discussed. Finally, the cooperation quality is measured in terms of the robot's task contribution.
Jose Ramon Medina, Martin Lawitzky, Alexander Mortl, Dongheui Lee, Sandra Hirche
IROS4
2011 Imitation learning of human grasping skills from motion and force data
abstract
Imitation learning, also known as Programming by Demonstration, allows a non-expert user to teach complex skills to a robot. While so far researchers focused on abstracting kinematic relations, only little attention has been paid to force information. In this work we study imitation learning of human grasping skills from motion and force data. For this purpose a teleoperation system is realized that allows a human to control a simulated robotic hand and to grasp objects in a virtual environment. Haptic rendering algorithms are implemented to calculate interaction forces that occur when touching the virtual object. While learning of fingertip interaction forces is shown to result in physical inconsistency compared to the demonstrations, we show that learning of internal tensions leads to stable reproductions of the demonstrated grasping skill. Obtained results further indicate an enlarged generalisation capability of grasping skills learnt on the basis of motion and force data compared to grasping skills that encode kinematic relations only.
Alexander M. Schmidts, Dongheui Lee, Angelika Peer
IROS2
2010 Incremental motion primitive learning by physical coaching using impedance control
abstract
We present an approach for kinesthetic teaching of motion primitives for a humanoid robot. The proposed teaching method allows for iterative execution and motion refinement using a forgetting factor. During the iterative motion refinement, a confidence value specifies an area of allowed refinement around the nominal trajectory. A novel method for continuous generation of motions from a hidden Markov model (HMM) representation of motion primitives is proposed, which incorporates relative time information for each state. On the real-time control level, the kinesthetic teaching is handled by a customized impedance controller, which combines tracking performance with soft physical interaction and allows to implement soft boundaries for the motion refinement. The proposed methods were implemented and tested using DLR's humanoid upper-body robot Justin.
Dongheui Lee, Christian Ott 0001
IROS1
2009 Whole body motion primitive segmentation from monocular video
abstract
This paper proposes a novel approach for motion primitive segmentation from continuous full body human motion captured on monocular video. The proposed approach does not require a kinematic model of the person, nor any markers on the body. Instead, optical flow computed directly in the image plane is used to estimate the location of segment points. The approach is based on detecting tracking features in the image based on the Shi and Thomasi algorithm [1]. The optical flow at each feature point is then estimated using the Lucas Kanade Pyramidal Optical Flow estimation algorithm [2]. The feature points are clustered and tracked on-line to find regions of the image with coherent movement. The appearance and disappearance of these coherent clusters indicates the start and end points of motion primitive segments. The algorithm performance is validated on full body motion video sequences, and compared to a joint-angle, motion capture based approach. The results show that the segmentation performance is comparable to the motion capture based approach, while using much simpler hardware and at a lower computational effort.
Dana Kulic, Dongheui Lee, Yoshihiko Nakamura
ICRA2
2009 Mimetic communication with impedance control for physical human-robot interaction
abstract
In this paper, mimetic communication is extended to human-robot interaction tasks, in which physical contact transitions must be handled. The mimetic communication consists of imitation learning for learning low level motion primitives and a higher level interaction learning stage in which also the information about the human-robot contacts is included. For the imitation learning, Cartesian marker data from a motion capture system is used. A modification of the low level marker trajectory following algorithm is presented, which allows to reshape the trajectory of the motion primitive in accordance with the human hand motion in real-time. Moreover, for performing safe contact motion, an appropriate impedance controller is integrated into the setting. All the presented concepts are evaluated in experiments with a humanoid robot.
Dongheui Lee, Christian Ott 0001, Yoshihiko Nakamura
ICRA1
2009 Associating and reshaping of whole body motions for object manipulation
abstract
Since humanoid robots have similar body structures to humans, a humanoid robot is expected to perform various dynamic tasks including object manipulation. This research focuses on issues related to learning and performing object manipulation. Basic motion primitives for tasks are learned from observation of human's behaviors. An object manipulation task is divided into two types of motion primitives, which are represented as hidden Markov models (HMMs): one for a body motion primitive and the other for the relation between the object and body parts, which manipulate the object. When performing a task, a natural whole body motion is associated from an object motion by using learned motion primitives. Furthermore, the associated body motion is reshaped in both spatial and temporal space, in a more precise way. The reshaping in spatial space is realized in two stages by a feedback control policy learned with reinforcement learning and by constrained inverse kinematics. Key features like end-effectors for manipulation and timing for a task are extracted and used for the feedback control policy learning. The reshaping in temporal space is realized by comparing a predicted and observed object motion speed.
Hirotoshi Kunori, Dongheui Lee, Yoshihiko Nakamura
IROS2
2008 Missing motion data recovery using factorial hidden Markov models
abstract
This paper proposes a method to recover missing data during observation by factorial hidden Markov models (FHMMs). The fundamental idea of the proposed method originates from the mimesis model, inspired by the mirror neuron system. By combining the motion recognition from partial observation algorithm and the proto-symbol based duplication of observed motion algorithm, whole body motion imitation from partial observation can be achieved. The algorithm for missing data recovery uses the same basic strategy as the whole body motion imitation from partial observation, but requires more accurate spatial representability. FHMMs allow for more efficient representation of a continuous data sequence by distributed state representation compared to hidden Markov models (HMMs). The proposed algorithm is tested with human motion data and the experimental results show improved representability compared to the conventional HMMs.
Dongheui Lee, Dana Kulic, Yoshihiko Nakamura
ICRA1
2008 Association of whole body motion from tool knowledge for humanoid robots
abstract
Since humanoid robots have similar body structures to humans, they are expected to perform various tasks including tool-use manipulation tasks instead of humans. This research studies on learning and performing tool-use manipulation tasks. For tool-use manipulations, understanding the relation between tool motion and whole body motion is crucial. In this paper, a tool-use motion model is designed with tool knowledge and body motion knowledge. The authors propose a method which enables a humanoid robot to associate whole body motion from tool knowledge by adopting the mimesis method from partial observations [1]. When a specific tool trajectory of a tool-use motion is given, appropriate hand motion is associated. From the calculated hand motion, appropriate whole body motion is associated successively. The proposed algorithm is implemented on a humanoid robot.
Dongheui Lee, Hirotoshi Kunori, Yoshihiko Nakamura
IROS1
2007 Mimesis Scheme using a Monocular Vision System on a Humanoid Robot
abstract
Optical motion capturing systems are widely used to acquire human beings' motion patterns in humanoid imitation learning research. However, optical motion capturing systems have a restricted movable area. This paper proposes the HMM based mimesis scheme using a monocular camera mounted on a humanoid. This scheme releases the restriction of movable area and enables imitation in daily life environments. Also, natural human-robot-interaction is expected during imitation. From two-dimensional image sequences of the demonstrator's motion, the demonstrator's pose and motion is estimated and recognized through the mimesis model and the humanoid generates its joint motor commands for imitation in 3D space. The feasibility of the proposed scheme is demonstrated by simulation.
Dongheui Lee, Yoshihiko Nakamura
ICRA1
2007 Motion capturing from monocular vision by statistical inference based on motion database: Vector field approach
abstract
This paper proposes a 3D motion recovery method from monocular images by statistical inference. The fundamental idea of the paper originates from the mimesis model, inspired by the mirror neuron system. The mimesis model is extended to include motion understanding from monocular image sequences and to imitate whole-body motion patterns in 3D space. In order to achieve this goal, (1) conversion of 3D motion database, represented in probabilistic form, into various spaces is adopted. (2) A vector field approach is developed for natural motion understanding. (3) With the particle filter, a demonstrator’s pose is estimated.
Dongheui Lee, Yoshihiko Nakamura
IROS1
2006 Stochastic Model of Imitating a New Observed Motion Based on the Acquired Motion Primitives
abstract
Generally, imitation of a motion means generation of a close motion to the observation. Moreover, it means that conversion into its own motion, which is adoptable to its body structure, by integrating with its prior knowledge. From this perspective, a new imitation scheme is proposed. The scheme is based on hidden Markov models by employing Viterbi algorithm. The proposed scheme enables to imitate a new observed motion without learning the motion by applying its prior knowledge. Online motion primitive acquisition method is considered. Evaluation factors, such as inheritance coordinate and matching error, are introduced to evaluate imitation performance. The feasibility of the proposed scheme is demonstrated by simulation on a 20 degrees of freedom humanoid robot configuration with the evaluation factors
Dongheui Lee, Yoshihiko Nakamura
IROS1
2005 Dependable localization strategy in dynamic real environments
abstract
Due to dynamic changes of an environment and various kinds of uncertainties in a real world, mobile robot localization is difficult to be solved by a single continuous algorithm. In order to achieve a practical localization solution generally, this paper proposes a strategy to deal with various uncertainties using explicit discretization of robot's status. Discrete status of localization is designed with three criteria as follows: (i) polygonal environment and non-polygonal environment; (ii) static environment and dynamic environment; and (iii) global positioning problem and local tracking problem are defined. An appropriate strategy is adopted according to the robot's status. The feasibility of the proposed method is demonstrated by simulation results.
Dongheui Lee, Woojin Chung
IROS1
2005 Mimesis from partial observations
abstract
In this paper, a new mimesis scheme is proposed. This scheme enables for a humanoid to imitate human's motion even though the humanoid cannot see human's whole-body motion and the humanoid has not seen the exactly same motion so far. Mimesis framework is based on continuous hidden Markov model. Viterbi algorithm is applied in order to generate more various motion patterns than the number of existing hidden Markov models. In order to imitate other's motion in a smooth way, a smoothing technique in generation problem is realized. The feasibility of this method is demonstrated by simulation on 20 degrees of freedom humanoid robot configuration.
Dongheui Lee, Yoshihiko Nakamura
IROS1
2004 Integrated Localization of the Service Robot PSR
abstract
Although a great deal of localization methods have been proposed, when it comes to human coexisting real world there are still many unsolved problems. It is because real world contains various kinds of uncertainties. For reliable navigation in such a world, this paper proposes a new localization synthesis integrated localization. The integrated localization is a dependable active localization approach, which is the structural synthesis of navigation modules. Due to a dynamic change of an environment, robot navigation cannot be solved by a single algorithm. In this paper, various situations are classified into different status, which is modeled as discrete events. Then, developed algorithms are synthesized in a structured way. The discrete event control structure enables efficient combination of position estimation algorithms and synthesis of navigation modules. Furthermore, the scheme provides structural framework for dead lock avoidance. The proposed technique is applied to KIST public service robots and shown to be useful in real experiments.
Dongheui Lee, Woojin Chung
ICRA1
2003 Reliable position estimation method of the service robot by map matching
abstract
In this paper, a reliable position estimation method of the indoor service robot is proposed. The service robot PSR1 is a wheeled mobile manipulator which navigates in office buildings. Our localization method is a map-matching scheme using scanned range data, without using any artificial landmark. The proposed algorithm can provide solutions for both a global localization problem and a local position tracking. A probabilistic position estimation scheme is designed based on MCL (Monte Carlo localization). Two measure functions are developed for computing positional probabilities. The robot automatically decides whether it uses geometric pattern matching (i.e. walls, pillars) by Hough transform. The proposed scheme shows reliable performance in both polygonal environments and non-polygonal environments even there exist many obstacles. Experimental results demonstrate the validity and feasibility of the proposed localization algorithm for the service robot to navigate in an office building, using the natural environmental characteristics.
Dongheui Lee, Woojin Chung
ICRA1
2003 Autonomous map building and smart localization of the service robot PSR
abstract
In this paper, an autonomous map building method and an intelligent position estimation method for an indoor service robot are presented. Map building is composed of three processes: (1) environmental information gathering, (2) scan registration, and (3) grid map building. A grid map of the environment can be successfully generated using the proposed strategy. Previously [Dongheui Lee et al., 2003] we proposed a localization method which is a map-matching scheme using scanned range data, without using any artificial landmarks. In this paper, an extended localization method called smart localization is presented. Smart localization includes the Petri net based discrete event control concept as well as the position estimation scheme using map matching. A mobile robot is able to act intelligently even if various real world problems arise when using discrete event control. For example, when the robot is unable to compute its position, discrete event based error handling logics are activated according to the predetermined behavioral configuration. Experimental results demonstrate the validity and feasibility of the proposed algorithm for a service robot to navigate in an office building.
Dongheui Lee, Woojin Chung
IROS1