Danica Kragic

dblp:82/1211 · also Danica Kragic Jensfelt · DBLP profile ↗
← Back
229ranked-venue papers
16as first author
60since 2021 · last 2026
0000-0003-2965-2953ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 199 · 12 first-author · 50 since 2021Systems, architecture and hardware · 143 · 9 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 31 · 1 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 22 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AI That Moves With You: A Review of Interactive Technologies Powered by Large Foundation Models for Mobility Impairment
abstract
Large foundation models (FMs) – including large language model (LLM), large vision model (LVM), vision language model (VLM), and related variants – are rapidly reshaping interactive assistive technologies during past years. We present a review of FM-enabled interactive systems for people with mobility impairments, covering work published from January 2020 to May 2025. Searching five databases, we screened 6,249 records and included 26 full papers. We first summerize descriptive results including study design and evaluation approaches of the reviewed studies. We then synthesize FM techniques, model integration patterns, interaction paradigms, and mobility impairment contexts. Our analysis surfaces and distills both technical and ethical challenges existed, lighting up future research topics. We contribute: (i) a conceptualization of FM-enabled interactions for mobility impairment functioning as a design space; (ii) a tabulated corpus with a reproducible codebook; and (iii) a forward agenda to guide and inspire the design of future mobility-assistance interactive systems within human-computer interaction (HCI) and CHI community.
Duosi Dai, Yuchong Zhang 0001, Yong Ma 0003, Danica Kragic
CHI4
2026 Data Augmentation with LLMs for Cold Start Recommendation in E-Commerce
Natalija Glisovic, Martin Tegner, Danica Kragic
ECIR (4)3
2025 HyperSteiner: Computing Heuristic Hyperbolic Steiner Minimal Trees
abstract
We propose HyperSteiner – an efficient heuristic algorithm for computing Steiner minimal trees in the hyperbolic space. HyperSteiner extends the Euclidean Smith-Lee-Liebman algorithm, which is grounded in a divide-and-conquer approach involving the Delaunay triangulation. The central idea is rephrasing Steiner tree problems with three terminals as a system of equations in the Klein-Beltrami model. Motivated by the fact that hyperbolic geometry is well-suited for representing hierarchies, we explore applications to hierarchy discovery in data. Results show that HyperSteiner infers more realistic hierarchies than the Minimum Spanning Tree and is more scalable to large datasets than Neighbor Joining.
Alejandro García-Castellanos, Aniss Aiman Medbouhi, Giovanni Luca Marchetti, Erik J. Bekkers, Danica Kragic
ALENEX5
2025 A Non-Adversarial Approach to Idempotent Generative Modelling
abstract
Idempotent Generative Networks (IGNs) are deep generative models that also function as local data manifold projectors, mapping arbitrary inputs back onto the manifold. They are trained to act as identity operators on the data and as idempotent operators off the data manifold. However, IGNs suffer from mode collapse, mode dropping, and training instability due to their objectives, which contain adversarial components and can cause the model to cover the data manifold only partially – an issue shared with generative adversarial networks. We introduce Non-Adversarial Idempotent Generative Networks (NAIGNs) to address these issues. Our loss function combines reconstruction with the non-adversarial generative objective of Implicit Maximum Likelihood Estimation (IMLE). This improves on IGN’s ability to restore corrupted data and generate new samples that closely match the data distribution. We moreover demonstrate that NAIGNs implicitly learn the distance field to the data manifold, as well as an energy-based model.
Mohammed Al-Jaff, Giovanni Luca Marchetti, Michael C. Welle, Jens Lundell, Mats G. Gustafsson, Gustav Eje Henter, Hossein Azizpour, Danica Kragic
ECAI8
2025 Deep Learning Amplified Early Stopping Bias: Overestimating Performance on Small Datasets
abstract
Cross-validation is commonly used to estimate machine learning model performance on new samples. However, using it for both hyperparameter selection and error estimation can lead to overestimating model performance, especially with extensive hyperparameter searches that overly tailor models to validation data. We demonstrate that deep learning further amplifies this bias, with even minor model adjustments causing significant overestimation. Our extensive experiments on simulated and real data focus on the bias from early stopping during cross-validation. We find that overestimation intensifies with network depth and is especially severe in small datasets, which are common in physiological signal processing applications. Selecting the early stopping point during cross-validation can result in ROC-AUC estimates exceeding 90% on random data, and this effect persists across various sample sizes, architectures, and network sizes1.
Nona Rajabi, Antônio H. Ribeiro, Miguel Vasco, Danica Kragic
ICASSP4
2025 A Riemannian Framework for Learning Reduced-order Lagrangian Dynamics
abstract
By incorporating physical consistency as inductive bias, deep neural networks display increased generalization capabilities and data efficiency in learning nonlinear dynamic models. However, the complexity of these models generally increases with the system dimensionality, requiring larger datasets, more complex deep networks, and significant computational effort. We propose a novel geometric network architecture to learn physically-consistent reduced-order dynamic parameters that accurately describe the original high-dimensional system behavior. This is achieved by building on recent advances in model-order reduction and by adopting a Riemannian perspective to jointly learn a non-linear structure-preserving latent space and the associated low-dimensional dynamics. Our approach enables accurate long-term predictions of the high-dimensional dynamics of rigid and deformable systems with increased data efficiency by inferring interpretable and physically-plausible reduced Lagrangian models.
Katharina Friedl, Noémie Jaquier, Jens Lundell, Tamim Asfour, Danica Kragic
ICLR5
2025 Human-Aligned Image Models Improve Visual Decoding from the Brain
abstract
Decoding visual images from brain activity has significant potential for advancing brain-computer interaction and enhancing the understanding of human perception. Recent approaches align the representation spaces of images and brain activity to enable visual decoding. In this paper, we introduce the use of human-aligned image encoders to map brain signals to images. We hypothesize that these models more effectively capture perceptual attributes associated with the rapid visual stimuli presentations commonly used in visual brain data recording experiments. Our empirical results support this hypothesis, demonstrating that this simple modification improves image retrieval accuracy by up to 21\% compared to state-of-the-art methods. Comprehensive experiments confirm consistent performance improvements across diverse EEG architectures, image encoders, alignment methods, participants, and brain imaging modalities.
Nona Rajabi, Antônio H. Ribeiro, Miguel Vasco, Farzaneh Taleb, Mårten Björkman, Danica Kragic
ICML6
2025 Flora: Sample-Efficient Preference-Based Rl Via Low-Rank Style Adaptation of Reward Functions
abstract
Preference-based reinforcement learning (PbRL) is a suitable approach for style adaptation of pre-trained robotic behavior: adapting the robot's policy to follow human user preferences while still being able to perform the original task. However, collecting preferences for the adaptation process in robotics is often challenging and time-consuming. In this work we explore the adaptation of pre-trained robots in the low-preference-data regime. We show that, in this regime, recent adaptation approaches suffer from catastrophic reward forgetting (CRF), where the updated reward model overfits to the new preferences, leading the agent to become unable to perform the original task. To mitigate CRF, we propose to enhance the original reward model with a small number of parameters (low-rank matrices) responsible for modeling the preference adaptation. Our evaluation shows that our method can efficiently and effectively adjust robotic behavior to human preferences across simulation benchmark tasks and multiple real-world robotic tasks. We provide videos of our results and source code at https://sites.google.com/view/preflora/.
Daniel Marta, Simon Holk, Miguel Vasco, Jens Lundell, Timon Homberger, Finn Lukas Busch, Olov Andersson, Danica Kragic, Iolanda Leite
ICRA8
2025 Feature Extractor or Decision Maker: Rethinking the Role of Visual Encoders in Visuomotor Policies
abstract
An end-to-end (E2E) visuomotor policy is typically treated as a unified whole, but recent approaches using out-of-domain (OOD) data to pretrain the visual encoder have cleanly separated the visual encoder from the network, with the remainder referred to as the policy. We propose Visual Alignment Testing, an experimental framework designed to evaluate the validity of this functional separation. Our results indicate that in E2E-trained models, visual encoders actively contribute to decision-making resulting from motor data supervision, contradicting the assumed functional separation. In contrast, OOD-pretrained models, where encoders lack this capability, experience an average performance drop of 42% in our benchmark results, compared to the state-of-the-art performance achieved by E2E policies. We believe this initial exploration of visual encoders' role can provide a first step towards guiding future pretraining methods to address their decision-making ability, such as developing task-conditioned or context-aware encoders.
Zheyu Zhuang, Shutong Jin, Nils Ingelhag, Danica Kragic, Florian T. Pokorny
ICRA5
2025 FLAME: A Federated Learning Benchmark for Robotic Manipulation
abstract
Recent progress in robotic manipulation has been fueled by large-scale datasets collected across diverse environments. Training robotic manipulation policies on these datasets is traditionally performed in a centralized manner, raising concerns regarding scalability, adaptability, and data privacy. While federated learning enables decentralized, privacy-preserving training, its application to robotic manipulation remains largely unexplored. We introduce FLAME (Federated Learning Across Manipulation Environments), the first benchmark designed for federated learning in robotic manipulation. FLAME consists of: (i) a set of large-scale datasets of over 160,000 expert demonstrations of multiple manipulation tasks, collected across a wide range of simulated environments; (ii) a training and evaluation framework for robotic policy learning in a federated setting. We evaluate standard federated learning algorithms in FLAME, showing their potential for distributed policy learning and highlighting key challenges. Our benchmark establishes a foundation for scalable, adaptive, and privacy-aware robotic learning. The code is publicly available at https://github.com/KTH-RPL/ELSA-Robotics-Challenge.
Santiago Bou Betran, Alberta Longhini, Miguel Vasco, Yuchong Zhang 0001, Danica Kragic
IROS5
2025 Real-time Iteration Scheme for Diffusion Policy
abstract
Diffusion Policies have demonstrated impressive performance in robotic manipulation tasks. However, their long inference time, resulting from an extensive iterative denoising process, and the need to execute an action chunk before the next prediction to maintain consistent actions limit their applicability to latency-critical tasks or simple tasks with a short cycle time. While recent methods explored distillation or alternative policy structures to accelerate inference, these often demand additional training, which can be resource-intensive for large robotic models. In this paper, we introduce a novel approach inspired by the Real-Time Iteration (RTI) Scheme, a method from optimal control that accelerates optimization by leveraging solutions from previous time steps as initial guesses for subsequent iterations. We explore the application of this scheme in diffusion inference and propose a scaling-based method to effectively handle discrete actions, such as grasping, in robotic manipulation. The proposed scheme significantly reduces runtime computational costs without the need for distillation or policy redesign. This enables a seamless integration into many pre-trained diffusion-based models, in particular, to resource-demanding large models. We also provide theoretical conditions for the contractivity which could be useful for estimating the initial denoising step. Quantitative results from extensive simulation experiments show a substantial reduction in inference time, with comparable overall performance compared with Diffusion Policy using full-step denoising. Our project page with additional resources is available at: https://rti-dp.github.io/
Yufei Duan, Danica Kragic
IROS3
2025 Towards Safe Reinforcement Learning with Reduced Conservativeness: A Case Study on Drone Flight Control
abstract
Incorporating formal methods into reinforcement learning (RL) has the potential to result in the best of both worlds, combining the robustness of formal guarantees with the adaptability and learning capabilities of RL, though careful design is needed to balance safety and exploration. In this work, we propose a framework to mitigate this loss of exploration while still allowing for the safety of the system to be ensured. Specifically, we introduce a less restrictive method that can reduce the conservativeness of formal methods by refining a disturbance model using online collected data and it evaluates the safety of a learning-based controller, using computationally efficient zonotopic reachability analysis for the safety analysis to facilitate a real-time implementation. We validate the framework in a real-world drone flight through a canyon, where the drone is subjected to unknown external disturbances and the framework is tasked with learning those disturbances online and adjusting the safety guarantees accordingly. The results show that the framework enables a less restrictive online training of learning-based controllers without compromising the safety of the system.
Loizos Hadjiloizou, Michael C. Welle, Hang Yin 0001, Danica Kragic
IROS4
2025 Learning Dexterous In-Hand Manipulation with Multifingered Hands via Visuomotor Diffusion
abstract
We present a framework for learning dexterous in-hand manipulation with multifingered hands using visuo-motor diffusion policies. Our system enables complex in-hand manipulation tasks, such as unscrewing a bottle lid with one hand, by leveraging a fast and responsive teleoperation setup for the four-fingered Allegro Hand. We collect high-quality expert demonstrations using an augmented reality (AR) interface that tracks hand movements and applies inverse kinematics and motion retargeting for precise control. The AR headset provides real-time visualization, while gesture controls streamline teleoperation. To enhance policy learning, we introduce a novel demonstration outlier removal approach based on HDBSCAN clustering and the Global-Local Outlier Score from Hierarchies (GLOSH) algorithm, effectively filtering out low-quality demonstrations that could degrade performance. We evaluate our approach extensively in real-world settings and provide all experimental videos on the project website.1.
Piotr Koczy, Michael C. Welle, Danica Kragic
IROS3
2025 Mind Meets Robots: A Review of EEG-Based Brain-Robot Interaction Systems
abstract
Brain-robot interaction (BRI) empowers individuals to control (semi-)automated machines through brain activity, either passively or actively. In the past decade, BRI systems have advanced significantly, primarily leveraging electroencephalogram (EEG) signals. This article presents an up-to-date review of 87 curated studies published between 2018 and 2023, identifying the research landscape of EEG-based BRI systems. The review consolidates methodologies, interaction modes, application contexts, system evaluation, existing challenges, and future directions in this domain. Based on our analysis, we propose a BRI system model comprising three entities: Brain, Robot, and Interaction, depicting their internal relationships. We especially examine interaction modes between human brains and robots, an aspect not yet fully explored. Within this model, we scrutinize and classify current research, extract insights, highlight challenges, and offer recommendations for future studies. Our findings provide a structured design space for human-robot interaction (HRI), informing the development of more efficient BRI frameworks.
Yuchong Zhang 0001, Nona Rajabi, Farzaneh Taleb, Andrii Matviienko, Yong Ma 0003, Mårten Björkman, Danica Kragic
Int. J. Hum. Comput. Interact.7
2024 Scalable Motion Style Transfer with Constrained Diffusion Generation
abstract
Current training of motion style transfer systems relies on consistency losses across style domains to preserve contents, hindering its scalable application to a large number of domains and private data. Recent image transfer works show the potential of independent training on each domain by leveraging implicit bridging between diffusion models, with the content preservation, however, limited to simple data patterns. We address this by imposing biased sampling in backward diffusion while maintaining the domain independence in the training stage. We construct the bias from the source domain keyframes and apply them as the gradient of content constraints, yielding a framework with keyframe manifold constraint gradients (KMCGs). Our validation demonstrates the success of training separate models to transfer between as many as ten dance motion styles. Comprehensive experiments find a significant improvement in preserving motion contents in comparison to baseline and ablative diffusion-based style transfer models. In addition, we perform a human study for a subjective assessment of the quality of generated dance motions. The results validate the competitiveness of KMCGs.
Yi Yu 0001, Hang Yin 0001, Danica Kragic, Mårten Björkman
AAAI4
2024 Scalable Unsupervised Feature Selection with Reconstruction Error Guarantees via QMR Decomposition
abstract
Unsupervised feature selection (UFS) methods have garnered significant attention for their capability to eliminate redundant features without relying on class label information. However, their scalability to large datasets remains a challenge, rendering common UFS methods impractical for such applications. To address this issue, we introduce QMR-FS, a greedy forward filtering approach that selects linearly independent features up to a specified relative tolerance, ensuring that any excluded features can be reconstructed from the retained set within this tolerance. This is achieved through the QMR matrix decomposition, which builds upon the well-known QR decomposition. QMR-FS benefits from linear complexity relative to the number of instances and boasts exceptional performance due to its ability to leverage parallelized computation on both CPU and GPU. Despite its greedy nature, QMR-FS achieves comparable classification and clustering accuracies across multiple datasets when compared to other UFS methods, while achieving runtimes approximately 10 times faster than recently proposed scalable UFS methods for datasets ranging from 100 million to 1 billion elements.
Ciwan Ceylan, Kambiz Ghoorchian, Danica Kragic
CIKM3
2024 Harmonics of Learning: Universal Fourier Features Emerge in Invariant Networks
abstract
In this work, we formally prove that, under certain conditions, if a neural network is invariant to a finite group then its weights recover the Fourier transform on that group. This provides a mathematical explanation for the emergence of Fourier features – a ubiquitous phenomenon in both biological and artificial learning systems. The results hold even for non-commutative groups, in which case the Fourier transform encodes all the irreducible unitary group representations. Our findings have consequences for the problem of symmetry discovery. Specifically, we demonstrate that the algebraic structure of an unknown group can be recovered from the weights of a network that is at least approximately invariant within certain bounds. Overall, this work contributes to a foundation for an algebraic learning theory of invariant neural network representations.
Giovanni Luca Marchetti, Christopher Hillar, Danica Kragic, Sophia Sanborn
COLT3
2024 Imitation or Innovation? Translating Features of Expressive Motion from Humans to Robots
abstract
Expressive robot motion can help establish acceptance of this technology in everyday life, but understanding what makes movement expressive is a complex and multifaceted task. This paper presents the results of an online study with 46 participants, it aims to explore how people perceive and interpret the expressive qualities of human movement and how they envision the translation of their description into an imagined non-humanoid, quadrupedal robot. Through a qualitative analysis of responses, we conceptualize three themes: their understanding of intent, their interpretations of movement qualities, and finally, their translation from human to robot movement. Respondents’ descriptions of their initial understanding of the performer’s intent fall into two modes, bio-mechanical and narrative. We illustrate their interpretations of movement qualities through four strategies: movement features as kinematic indicators, intent indicators, attributed context, and perceived internal states. Lastly, we observe their translation from human to robot movement, with a particular focus on respondents’ use of kinaesthetic empathy and anthropomorphism. Our findings aim to support a bottom-up approach, using users’ general knowledge for designing expressive robot motion.
Benedikte Wallace, Marieke van Otterdijk, Yuchong Zhang 0001, Nona Rajabi, Diego Marin-Bucio, Danica Kragic, Jim Tørresen
HAI6
2024 Standardization of Cloth Objects and its Relevance in Robotic Manipulation
abstract
The field of robotics faces inherent challenges in manipulating deformable objects, particularly in understanding and standardising fabric properties like elasticity, stiffness, and friction. While the significance of these properties is evident in the realm of cloth manipulation, accurately categorising and comprehending them in real-world applications remains elusive. This study sets out to address two primary objectives: (1) to provide a framework suitable for robotics applications to characterise cloth objects, and (2) to study how these properties influence robotic manipulation tasks. Our preliminary results validate the framework’s ability to characterise cloth properties and compare cloth sets, and reveal the influence that different properties have on the outcome of five manipulation primitives. We believe that, in general, results on the manipulation of clothes should be reported along with a better description of the garments used in the evaluation. This paper proposes a set of these measures.
Irene Garcia-Camacho, Alberta Longhini, Michael C. Welle, Guillem Alenyà, Danica Kragic, Júlia Borràs Sol
ICRA5
2024 Ensemble Latent Space Roadmap for Improved Robustness in Visual Action Planning
abstract
Planning in learned latent spaces helps to decrease the dimensionality of raw observations. In this work, we propose to leverage the ensemble paradigm to enhance the robustness of latent planning systems. We rely on our Latent Space Roadmap (LSR) framework, which builds a graph in a learned structured latent space to perform planning. Given multiple LSR framework instances, that differ either on their latent spaces or on the parameters for constructing the graph, we use the action information as well as the embedded nodes of the produced plans to define similarity measures. These are then utilized to select the most promising plans. We validate the performance of our Ensemble LSR (ENS-LSR) on simulated box stacking and grape harvesting tasks as well as on a real-world robotic T-shirt folding experiment.
Martina Lippi, Michael C. Welle, Andrea Gasparri, Danica Kragic
ICRA4
2024 Raising Body Ownership in End-to-End Visuomotor Policy Learning via Robot-Centric Pooling
abstract
We present Robot-centric Pooling (RcP), a novel pooling method designed to enhance end-to-end visuomo-tor policies by enabling differentiation between the robots and similar entities or their surroundings. Given an image-proprioception pair, RcP guides the aggregation of image features by highlighting image regions correlating with the robot’s proprioceptive states, thereby extracting robot-centric image representations for policy learning. Leveraging contrastive learning techniques, RcP integrates seamlessly with existing visuomotor policy learning frameworks and is trained jointly with the policy using the same dataset, requiring no extra data collection involving self-distractors. We evaluate the proposed method with reaching tasks in both simulated and real-world settings. The results demonstrate that RcP significantly enhances the policies’ robustness against various unseen distractors, including self-distractors, positioned at different locations. Additionally, the inherent robot-centric characteristic of RcP enables the learnt policy to be far more resilient to aggressive pixel shifts compared to the baselines. Code available at: https://github.com/Zheyu-Zhuang/RcP
Zheyu Zhuang, Ville Kyrki, Danica Kragic
IROS3
2024 Can Transformers Smell Like Humans?
abstract
The human brain encodes stimuli from the environment into representations that form a sensory perception of the world. Despite recent advances in understanding visual and auditory perception, olfactory perception remains an under-explored topic in the machine learning community due to the lack of large-scale datasets annotated with labels of human olfactory perception. In this work, we ask the question of whether pre-trained transformer models of chemical structures encode representations that are aligned with human olfactory perception, i.e., can transformers smell like humans? We demonstrate that representations encoded from transformers pre-trained on general chemical structures are highly aligned with human olfactory perception. We use multiple datasets and different types of perceptual representations to show that the representations encoded by transformer models are able to predict: (i) labels associated with odorants‌‌ provided by experts; (ii) continuous ratings provided by human participants with respect to pre-defined descriptors; and (iii) similarity ratings between odorants provided by human participants. Finally, we evaluate the extent to which this alignment is associated with physicochemical features of odorants known to be relevant for olfactory decoding.
Farzaneh Taleb, Miguel Vasco, Antônio H. Ribeiro, Mårten Björkman, Danica Kragic
NeurIPS5
2024 Hyperbolic Delaunay Geometric Alignment
Aniss Aiman Medbouhi, Giovanni Luca Marchetti, Vladislav Polianskii, Alexander Kravberg, Petra Poklukar, Anastasia Varava, Danica Kragic
ECML/PKDD (3)7
2024 A Robotic Skill Learning System Built Upon Diffusion Policies and Foundation Models
abstract
In this paper, we build upon two major recent developments in the field, Diffusion Policies for visuomotor manipulation and large pre-trained multimodal foundational models to obtain a robotic skill learning system. The system can obtain new skills via the behavioral cloning approach of visuomotor diffusion policies given teleoperated demonstrations. Foundational models are being used to perform skill selection given the user’s prompt in natural language. Before executing a skill the foundational model performs a precondition check given an observation of the workspace. We compare the performance of different foundational models to this end and give a detailed experimental evaluation of the skills taught by the user in simulation and the real world. Finally, we showcase the combined system on a challenging food serving scenario in the real world. Videos of all experimental executions, as well as the process of teaching new skills in simulation and the real world, are available on the project’s website1.
Nils Ingelhag, Jesper Munkeby, Jonne van Haastregt, Anastasia Varava, Michael C. Welle, Danica Kragic
RO-MAN6
2024 Visual Action Planning with Multiple Heterogeneous Agents
abstract
Visual planning methods are promising to handle complex settings where extracting the system state is challenging. However, none of the existing works tackles the case of multiple heterogeneous agents which are characterized by different capabilities and/or embodiment. In this work, we propose a method to realize visual action planning in multi-agent settings by exploiting a roadmap built in a low-dimensional structured latent space and used for planning. To enable multi-agent settings, we infer possible parallel actions from a dataset composed of tuples associated with individual actions. Next, we evaluate feasibility and cost of them based on the capabilities of the multi-agent system and endow the roadmap with this information, building a capability latent space roadmap (C-LSR). Additionally, a capability suggestion strategy is designed to inform the human operator about possible missing capabilities when no paths are found. The approach is validated in a simulated burger cooking task and a real-world box packing task.
Martina Lippi, Michael C. Welle, Marco Moletta, Alessandro Marino, Andrea Gasparri, Danica Kragic
RO-MAN6
2024 Low-Cost Teleoperation with Haptic Feedback through Vision-based Tactile Sensors for Rigid and Soft Object Manipulation
abstract
Haptic feedback is essential for humans to successfully perform complex and delicate manipulation tasks. A recent rise in tactile sensors has enabled robots to leverage the sense of touch and expand their capability drastically. However, many tasks still need human intervention/guidance. For this reason, we present a teleoperation framework designed to provide haptic feedback to human operators based on the data from camera-based tactile sensors mounted on the robot gripper. Partial autonomy is introduced to prevent slippage of grasped objects during task execution. Notably, we rely exclusively on low-cost off-the-shelf hardware to realize an affordable solution. We demonstrate the versatility of the framework on nine different objects ranging from rigid to soft and fragile ones, using three different operators on real hardware.
Martina Lippi, Michael C. Welle, Maciej Wozniak 0001, Andrea Gasparri, Danica Kragic
RO-MAN5
2023 An Efficient and Continuous Voronoi Density Estimator
abstract
We introduce a non-parametric density estimator deemed Radial Voronoi Density Estimator (RVDE). RVDE is grounded in the geometry of Voronoi tessellations and as such benefits from local geometric adaptiveness and broad convergence properties. Due to its radial definition RVDE is continuous and computable in linear time with respect to the dataset size. This amends for the main shortcomings of previously studied VDEs, which are highly discontinuous and computationally expensive. We provide a theoretical study of the modes of RVDE as well as an empirical investigation of its performance on high-dimensional data. Results show that RVDE outperforms other non-parametric density estimators, including recently introduced VDEs.
Giovanni Luca Marchetti, Vladislav Polianskii, Anastasiia Varava, Florian T. Pokorny, Danica Kragic
AISTATS5
2023 Equivariant Representation Learning via Class-Pose Decomposition
abstract
We introduce a general method for learning representations that are equivariant to symmetries of data. Our central idea is to decompose the latent space into an invariant factor and the symmetry group itself. The components semantically correspond to intrinsic data classes and poses respectively. The learner is trained on a loss encouraging equivariance based on supervision from relative symmetry information. The approach is motivated by theoretical results from group theory and guarantees representations that are lossless, interpretable and disentangled. We provide an empirical investigation via experiments involving datasets with a variety of symmetries. Results show that our representations capture the geometry of data and outperform other equivariant representation learning frameworks.
Giovanni Luca Marchetti, Gustaf Tegnér, Anastasiia Varava, Danica Kragic
AISTATS4
2023 TD-GEM: Text-Driven Garment Editing Mapper
Reza Dadfar, Sanaz Sabzevari, Mårten Björkman, Danica Kragic
BMVC4
2023 EDO-Net: Learning Elastic Properties of Deformable Objects from Graph Dynamics
abstract
We study the problem of learning graph dynamics of deformable objects that generalizes to unknown physical properties. Our key insight is to leverage a latent representation of elastic physical properties of cloth-like deformable objects that can be extracted, for example, from a pulling interaction. In this paper we propose EDO-Net (Elastic Deformable Object - Net), a model of graph dynamics trained on a large variety of samples with different elastic properties that does not rely on ground-truth labels of the properties. EDO-Net jointly learns an adaptation module, and a forward-dynamics module. The former is responsible for extracting a latent representation of the physical properties of the object, while the latter leverages the latent representation to predict future states of cloth-like objects represented as graphs. We evaluate EDO-Net both in simulation and real world, assessing its capabilities of: 1) generalizing to unknown physical properties, 2) transferring the learned representation to new downstream tasks.
Alberta Longhini, Marco Moletta, Alfredo Reichlin, Michael C. Welle, David Held, Zackory Erickson, Danica Kragic
ICRA7
2023 Elastic Context: Encoding Elasticity for Data-driven Models of Textiles Elastic Context: Encoding Elasticity for Data-driven Models of Textiles
abstract
Physical interaction with textiles, such as assistive dressing or household tasks, requires advanced dexterous skills. The complexity of textile behavior during stretching and pulling is influenced by the material properties of the yarn and by the textile's construction technique, which are often unknown in real-world settings. Moreover, identification of physical properties of textiles through sensing commonly available on robotic platforms remains an open problem. To address this, we introduce Elastic Context (EC), a method to encode the elasticity of textiles using stress-strain curves adapted from textile engineering for robotic applications. We employ EC to learn generalized elastic behaviors of textiles and examine the effect of EC dimension on accurate force modeling of real-world non-linear elastic behaviors.
Alberta Longhini, Marco Moletta, Alfredo Reichlin, Michael C. Welle, Alexander Kravberg, Yufei Wang 0007, David Held, Zackory Erickson, Danica Kragic
ICRA9
2023 Generating Scenarios from High-Level Specifications for Object Rearrangement Tasks
abstract
Rearranging objects is an essential skill for robots. To quickly teach robots new rearrangements tasks, we would like to generate training scenarios from high-level specifications that define the relative placement of objects for the task at hand. Ideally, to guide the robot's learning we also want to be able to rank these scenarios according to their difficulty. Prior work has shown how generating diverse scenario from specifications and providing the robot with easy-to-difficult samples can improve the learning. Yet, existing scenario generation methods typically cannot generate diverse scenarios while controlling their difficulty. We address this challenge by conditioning generative models on spatial logic specifications to generate spatially-structured scenarios that meet the specification and desired difficulty level. Our experiments showed that generative models are more effective and data-efficient than rejection sam-pling and that the spatially-structured scenarios can drastically improve training of downstream tasks by orders of magnitude.
Sanne van Waveren, Christian Pek, Iolanda Leite, Jana Tumova, Danica Kragic
IROS5
2023 Learning Geometric Representations of Objects via Interaction
Alfredo Reichlin, Giovanni Luca Marchetti, Hang Yin 0001, Anastasiia Varava, Danica Kragic
ECML/PKDD (4)5
2023 Equivariant Representation Learning in the Presence of Stabilizers
Luis A. Pérez Rey, Giovanni Luca Marchetti, Danica Kragic, Dmitri Jarnikov, Mike Holenderski
ECML/PKDD (4)3
2023 Detecting the Intention of Object Handover in Human-Robot Collaborations: An EEG Study
abstract
Human-robot collaboration (HRC) relies on smooth and safe interactions. In this paper, we focus on the human-to-robot handover scenario, where the robot acts as a taker. We investigate the feasibility of detecting the intention of a human-to-robot handover action through the analysis of electroencephalogram (EEG) signals. Our study confirms that temporal patterns in EEG signals provide information about motor planning and can be leveraged to predict the likelihood of an individual executing a motor task with an average accuracy of 94.7%. We also suggest the effectiveness of the time-frequency features of EEG signals in the final second prior to the movement for distinguishing between handover action and other actions. Furthermore, we classify human intentions for different tasks based on time-frequency representations of pre-movement EEG signals and achieve an average accuracy of 63.5% for contrasting every two tasks against each other. The result encourages the possibility of using EEG signals to detect human handover intention in HRC tasks.
Nona Rajabi, Parag Khanna, Sumeyra Demir Kanik, Elmira Yadollahi, Miguel Vasco, Mårten Björkman, Christian Smith, Danica Kragic
RO-MAN8
2023 Controllable Motion Synthesis and Reconstruction with Autoregressive Diffusion Models
abstract
Data-driven and controllable human motion synthesis and prediction are active research areas with various applications in interactive media and social robotics. Challenges remain in these fields for generating diverse motions given past observations and dealing with imperfect poses. This paper introduces MoDiff, an autoregressive probabilistic diffusion model over motion sequences conditioned on control contexts of other modalities. Our model integrates a cross-modal Transformer encoder and a Transformer-based decoder, which are found effective in capturing temporal correlations in motion and control modalities. We also introduce a new data dropout method based on the diffusion forward process to provide richer data representations and robust generation. We demonstrate the superior performance of MoDiff in controllable motion synthesis for locomotion with respect to two baselines and show the benefits of diffusion data dropout for robust synthesis and reconstruction of high-fidelity motion close to recorded data.
Ruibo Tu, Hang Yin 0001, Danica Kragic, Hedvig Kjellström, Mårten Björkman
RO-MAN4
2023 Dance Style Transfer with Cross-modal Transformer
abstract
We present CycleDance, a dance style transfer system to transform an existing motion clip in one dance style to a motion clip in another dance style while attempting to preserve motion context of the dance. Our method extends an existing CycleGAN architecture for modeling audio sequences and integrates multimodal transformer encoders to account for music context. We adopt sequence length-based curriculum learning to stabilize training. Our approach captures rich and long-term intra-relations between motion frames, which is a common challenge in motion transfer and synthesis work. We further introduce new metrics for gauging transfer strength and content preservation in the context of dance movements. We perform an extensive ablation study as well as a human study including 30 participants with 5 or more years of dance experience. The results demonstrate that CycleDance generates realistic movements with the target style, significantly outperforming the baseline CycleGAN on naturalness, transfer strength, and content preservation.1
Hang Yin 0001, Kim Baraka, Danica Kragic, Mårten Björkman
WACV4
2023 Multimodal dance style transfer
abstract
Abstract This paper first presents CycleDance, a novel dance style transfer system that transforms an existing motion clip in one dance style into a motion clip in another dance style while attempting to preserve the motion context of the dance. CycleDance extends existing CycleGAN architectures with multimodal transformer encoders to account for the music context. We adopt a sequence length-based curriculum learning strategy to stabilize training. Our approach captures rich and long-term intra-relations between motion frames, which is a common challenge in motion transfer and synthesis work. Building upon CycleDance, we further propose StarDance, which enables many-to-many mappings across different styles using a single generator network. Additionally, we introduce new metrics for gauging transfer strength and content preservation in the context of dance movements. To evaluate the performance of our approach, we perform an extensive ablation study and a human study with 30 participants, each with 5 or more years of dance experience. Our experimental results show that our approach can generate realistic movements with the target style, outperforming the baseline CycleGAN and its variants on naturalness, transfer strength, and content preservation. Our proposed approach has potential applications in choreography, gaming, animation, and tool development for artistic and scientific innovations in the field of dance.
Hang Yin 0001, Kim Baraka, Danica Kragic, Mårten Björkman
Mach. Vis. Appl.4
2023 Enabling Visual Action Planning for Object Manipulation Through Latent Space Roadmap
abstract
In this article, we present a framework for visual action planning of complex manipulation tasks with high-dimensional state spaces, focusing on manipulation of deformable objects. We propose a latent space roadmap (LSR) for task planning, which is a graph-based structure globally capturing the system dynamics in a low-dimensional latent space. Our framework consists of the following three parts. First, a mapping module (MM) that maps observations is given in the form of images into a structured latent space extracting the respective states as well as generates observations from the latent states. Second, the LSR, which builds and connects clusters containing similar states in order to find the latent plans between start and goal states, extracted by MM. Third, the action proposal module that complements the latent plan found by the LSR with the corresponding actions. We present a thorough investigation of our framework on simulated box stacking and rope/box manipulation tasks, and a folding task executed on a real robot.
Martina Lippi, Petra Poklukar, Michael C. Welle, Anastasia Varava, Hang Yin 0001, Alessandro Marino, Danica Kragic
IEEE Trans. Robotics7
2023 Safe Data-Driven Model Predictive Control of Systems With Complex Dynamics
abstract
In this article, we address the task and safety performance of data-driven model predictive controllers (DD-MPC) for systems with complex dynamics, i.e., temporally or spatially varying dynamics that may also be discontinuous. The three challenges we focus on are the accuracy of learned models, the receding horizon-induced myopic predictions of DD-MPC, and the active encouragement of safety. To learn accurate models for DD-MPC, we cautiously, yet effectively, explore the dynamical system with rapidly exploring random trees (RRT) to collect a uniform distribution of samples in the state-input space and overcome the common distribution shift in model learning. The learned model is further used to construct an RRT tree that estimates how close the model's predictions are to the desired target. This information is used in the cost function of the DD-MPC to minimize the short-sighted effect of its receding horizon nature. To promote safety, we approximate sets of safe states using demonstrations of exclusively safe trajectories, i.e., without unsafe examples, and encourage the controller to generate trajectories close to the sets. As a running example, we use abrokenversion of an inverted pendulum where the friction abruptly changes in certain regions. Furthermore, we showcase the adaptation of our method to a real-world robotic application with complex dynamics: robotic food-cutting. Our results show that our proposed control framework effectively avoids unsafe states with higher success rates than baseline controllers that employ models from controlled demonstrations and even random actions.
Ioanna Mitsioni, Pouria Tajvar, Danica Kragic, Jana Tumova, Christian Pek
IEEE Trans. Robotics3
2023 Deep Learning Approaches to Grasp Synthesis: A Review
abstract
Grasping is the process of picking up an object by applying forces and torques at a set of contacts. Recent advances in deep learning methods have allowed rapid progress in robotic object grasping. In this systematic review, we surveyed the publications over the last decade, with a particular interest in grasping an object using all six degrees of freedom of the end-effector pose. Our review found four common methodologies for robotic grasping: sampling-based approaches, direct regression, reinforcement learning, and exemplar approaches In addition, we found two “supporting methods” around grasping that use deep learning to support the grasping process, shape approximation, and affordances. We have distilled the publications found in this systematic review (85 papers) into ten key takeaways we consider crucial for future robotic grasping and manipulation research.
Rhys Newbury, Morris Gu, Lachlan Chumbley, Arsalan Mousavian, Clemens Eppner, Jürgen Leitner, Jeannette Bohg, Antonio Morales, Tamim Asfour, Danica Kragic, Dieter Fox, Akansel Cosgun
IEEE Trans. Robotics10
2022 Delaunay Component Analysis for Evaluation of Data Representations
Petra Poklukar, Vladislav Polianskii, Anastasiia Varava, Florian T. Pokorny, Danica Kragic
ICLR5
2022 Active Nearest Neighbor Regression Through Delaunay Refinement
abstract
We introduce an algorithm for active function approximation based on nearest neighbor regression. Our Active Nearest Neighbor Regressor (ANNR) relies on the Voronoi-Delaunay framework from computational geometry to subdivide the space into cells with constant estimated function value and select novel query points in a way that takes the geometry of the function graph into account. We consider the recent state-of-the-art active function approximator called DEFER, which is based on incremental rectangular partitioning of the space, as the main baseline. The ANNR addresses a number of limitations that arise from the space subdivision strategy used in DEFER. We provide a computationally efficient implementation of our method, as well as theoretical halting guarantees. Empirical results show that ANNR outperforms the baseline for both closed-form functions and real-world examples, such as gravitational wave parameter inference and exploration of the latent space of a generative model.
Alexander Kravberg, Giovanni Luca Marchetti, Vladislav Polianskii, Anastasiia Varava, Florian T. Pokorny, Danica Kragic
ICML6
2022 Geometric Multimodal Contrastive Representation Learning
abstract
Learning representations of multimodal data that are both informative and robust to missing modalities at test time remains a challenging problem due to the inherent heterogeneity of data obtained from different channels. To address it, we present a novel Geometric Multimodal Contrastive (GMC) representation learning method consisting of two main components: i) a two-level architecture consisting of modality-specific base encoders, allowing to process an arbitrary number of modalities to an intermediate representation of fixed dimensionality, and a shared projection head, mapping the intermediate representations to a latent representation space; ii) a multimodal contrastive loss function that encourages the geometric alignment of the learned representations. We experimentally demonstrate that GMC representations are semantically rich and achieve state-of-the-art performance with missing modality information on three different learning problems including prediction and reinforcement learning tasks.
Petra Poklukar, Miguel Vasco, Hang Yin 0001, Francisco S. Melo, Ana Paiva 0001, Danica Kragic
ICML6
2022 Comparing Reconstruction- and Contrastive-based Models for Visual Task Planning
abstract
Learning state representations enables robotic planning directly from raw observations such as images. Several methods learn state representations by utilizing losses based on the reconstruction of the raw observations from a lower-dimensional latent space. The similarity between observations in the space of images is often assumed and used as a proxy for estimating similarity between the underlying states of the system. However, observations commonly contain task-irrelevant factors of variation which are nonetheless important for reconstruction, such as varying lighting and different camera viewpoints. In this work, we define relevant evaluation metrics and perform a thorough study of different loss functions for state representation learning. We show that models exploiting task priors, such as Siamese networks with a simple contrastive loss, outperform reconstruction-based representations in visual task planning in case of task-irrelevant factors of variations.
Constantinos Chamzas, Martina Lippi, Michael C. Welle, Anastasia Varava, Lydia E. Kavraki, Danica Kragic
IROS6
2022 Augment-Connect-Explore: a Paradigm for Visual Action Planning with Data Scarcity
abstract
Visual action planning particularly excels in applications where the state of the system cannot be computed explicitly, such as manipulation of deformable objects, as it enables planning directly from raw images. Even though the field has been significantly accelerated by deep learning techniques, a crucial requirement for their success is the availability of a large amount of data. In this work, we propose the Augment-Connect-Explore (ACE) paradigm to enable visual action planning in cases of data scarcity. We build upon the Latent Space Roadmap (LSR) framework which performs planning with a graph built in a low dimensional latent space. In particular, ACE is used to i) Augment the available training dataset by autonomously creating new pairs of datapoints, ii) create new unobserved Connections among representations of states in the latent graph, and iii) Explore new regions of the latent space in a targeted manner. We validate the proposed approach on both simulated box stacking and real-world folding task showing the applicability for rigid and deformable object manipulation tasks, respectively.
Martina Lippi, Michael C. Welle, Petra Poklukar, Alessandro Marino, Danica Kragic
IROS5
2022 Back to the Manifold: Recovering from Out-of-Distribution States
abstract
Learning from previously collected datasets of expert data offers the promise of acquiring robotic policies without unsafe and costly online explorations. However, a major challenge is a distributional shift between the states in the training dataset and the ones visited by the learned policy at the test time. While prior works mainly studied the distribution shift caused by the policy during the offline training, the problem of recovering from out-of-distribution states at the deployment time is not very well studied yet. We alleviate the distributional shift at the deployment time by introducing a recovery policy that brings the agent back to the training manifold whenever it steps out of the in-distribution states, e.g., due to an external perturbation. The recovery policy relies on an approximation of the training data density and a learned equivariant mapping that maps visual observations into a latent space in which translations correspond to the robot actions. We demonstrate the effectiveness of the proposed method through several manipulation experiments on a real robotic platform. Our results show that the recovery policy enables the agent to complete tasks while the behavioral cloning alone fails because of the distributional shift problem.
Alfredo Reichlin, Giovanni Luca Marchetti, Hang Yin 0001, Ali Ghadirzadeh, Danica Kragic
IROS5
2022 Consensus-based Normalizing-Flow Control: A Case Study in Learning Dual-Arm Coordination
abstract
We develop two consensus-based learning algorithms for multi-robot systems applied on complex tasks involving collision constraints and force interactions, such as the cooperative peg-in-hole placement. The proposed algorithms integrate multi-robot distributed consensus and normalizing-flow-based reinforcement learning. The algorithms guarantee the stability and the consensus of the multi-robot system's generalized variables in a transformed space. This transformed space is obtained via a diffeomorphic transformation parameterized by normalizing-flow models that the algorithms use to train the underlying task, learning hence skillful, dexterous trajectories required for the task accomplishment. We validate the proposed algorithms by parameterizing reinforcement learning policies, demonstrating efficient cooperative learning, and strong generalization of dual-arm assembly skills in a dynamics-engine simulator.
Hang Yin 0001, Christos K. Verginis, Danica Kragic
IROS3
2022 Embedding Koopman Optimal Control in Robot Policy Learning
abstract
Embedding an optimization process has been explored for imposing efficient and flexible policy structures. Existing work often build upon nonlinear optimization with explicitly iteration steps, making policy inference prohibitively expensive for online learning and real-time control. Our approach embeds a linear-quadratic-regulator (LQR) formulation with a Koopman representation, thus exhibiting the tractability from a closed-form solution and richness from a non-convex neural network. We use a few auxiliary objectives and reparameterization to enforce optimality conditions of the policy that can be easily integrated to standard gradient-based learning. Our approach is shown to be effective for learning policies rendering an optimality structure and efficient reinforcement learning, including simulated pendulum control, 2D and 3D walking, and manipulation for both rigid and deformable objects. We also demonstrate real world application in a robot pivoting task.
Hang Yin 0001, Michael C. Welle, Danica Kragic
IROS3
2022 Voronoi density estimator for high-dimensional data: Computation, compactification and convergence
abstract
The Voronoi Density Estimator (VDE) is an established density estimation technique that adapts to the local geometry of data. However, its applicability has been so far limited to problems in two and three dimensions. This is because Voronoi cells rapidly increase in complexity as dimensions grow, making the necessary explicit computations infeasible. We define a variant of the VDE deemed Compactified Voronoi Density Estimator (CVDE), suitable for higher dimensions. We propose computationally efficient algorithms for numerical approximation of the CVDE and formally prove convergence of the estimated density to the original one. We implement and empirically validate the CVDE through a comparison with the Kernel Density Estimator (KDE). Our results indicate that the CVDE outperforms the KDE on sound and image data.
Vladislav Polianskii, Giovanni Luca Marchetti, Alexander Kravberg, Anastasiia Varava, Florian T. Pokorny, Danica Kragic
UAI6
2022 Training and Evaluation of Deep Policies Using Reinforcement Learning and Generative Models
abstract
We present a data-efficient framework for solving sequential decision-making problems which exploits the combination of reinforcement learning (RL) and latent variable generative models. The framework, called GenRL, trains deep policies by introducing an action latent variable such that the feed-forward policy search can be divided into two parts: (i) training a sub-policy that outputs a distribution over the action latent variable given a state of the system, and (ii) unsupervised training of a generative model that outputs a sequence of motor actions conditioned on the latent action variable. GenRL enables safe exploration and alleviates the data-inefficiency problem as it exploits prior knowledge about valid sequences of motor actions. Moreover, we provide a set of measures for evaluation of generative models such that we are able to predict the performance of the RL policy training prior to the actual training on a physical robot. We experimentally determine the characteristics of generative models that have most influence on the performance of the final policy training on two robotics tasks: shooting a hockey puck and throwing a basketball. Furthermore, we empirically demonstrate that GenRL is the only method which can safely and efficiently solve the robotics tasks compared to two state-of-the-art RL methods.
Ali Ghadirzadeh, Petra Poklukar, Karol Arndt, Chelsea Finn, Ville Kyrki, Danica Kragic, Mårten Björkman
J. Mach. Learn. Res.6
2022 In Memoriam: Jan-Olof Eklundh
Atsuto Maki, Danica Kragic, Hedvig Kjellström, Hossein Azizpour, Josephine Sullivan, Mårten Björkman, Patric Jensfelt, Stefan Carlsson, Tony Lindeberg, Yngve Sundblad
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 GeomCA: Geometric Evaluation of Data Representations
abstract
Evaluating the quality of learned representations without relying on a downstream task remains one of the challenges in representation learning. In this work, we present Geometric Component Analysis (GeomCA) algorithm that evaluates representation spaces based on their geometric and topological properties. GeomCA can be applied to representations of any dimension, independently of the model that generated them. We demonstrate its applicability by analyzing representations obtained from a variety of scenarios, such as contrastive learning models, generative models and supervised learning models.
Petra Poklukar, Anastasiia Varava, Danica Kragic
ICML3
2021 Learning Stable Normalizing-Flow Control for Robotic Manipulation
abstract
Reinforcement Learning (RL) of robotic manipulation skills, despite its impressive successes, stands to benefit from incorporating domain knowledge from control theory. One of the most important properties that is of interest is control stability. Ideally, one would like to achieve stability guarantees while staying within the framework of state-of-the-art deep RL algorithms. Such a solution does not exist in general, especially one that scales to complex manipulation tasks. We contribute towards closing this gap by introducing normalizing-flow control structure, that can be deployed in any latest deep RL algorithms. While stable exploration is not guaranteed, our method is designed to ultimately produce deterministic controllers with provable stability. In addition to demonstrating our method on challenging contact-rich manipulation tasks, we also show that it is possible to achieve considerable exploration efficiency–reduced state space coverage and actuation efforts– without losing learning efficiency.
Shahbaz Abdul Khader, Hang Yin 0001, Pietro Falco, Danica Kragic
ICRA4
2021 Interpretability in Contact-Rich Manipulation via Kinodynamic Images
abstract
Deep Neural Networks (NNs) have been widely utilized in contact-rich manipulation tasks to model the complicated contact dynamics. However, NN-based models are often difficult to decipher which can lead to seemingly inexplicable behaviors and unidentifiable failure cases. In this work, we address the interpretability of NN-based models by introducing the kinodynamic images. We propose a methodology that creates images from kinematic and dynamic data of contact-rich manipulation tasks. By using images as the state representation, we enable the application of interpretability modules that were previously limited to vision-based tasks. We use this representation to train a Convolutional Neural Network (CNN) and we extract interpretations with Grad-CAM to produce visual explanations. Our method is versatile and can be applied to any classification problem in manipulation tasks to visually interpret which parts of the input drive the model’s decisions and distinguish its failure modes, regardless of the features used. Our experiments demonstrate that our method enables detailed visual inspections of sequences in a task, and high-level evaluations of a model’s behavior. Code for this work is available at [1].
Ioanna Mitsioni, Joonatan Mänttäri, Yiannis Karayiannidis, John Folkesson, Danica Kragic
ICRA5
2021 Bayesian Meta-Learning for Few-Shot Policy Adaptation Across Robotic Platforms
abstract
Reinforcement learning methods can achieve significant performance but require a large amount of training data collected on the same robotic platform. A policy trained with expensive data is rendered useless after making even a minor change to the robot hardware. In this paper, we address the challenging problem of adapting a policy, trained to perform a task, to a novel robotic hardware platform given only few demonstrations of robot motion trajectories on the target robot. We formulate it as a few-shot meta-learning problem where the goal is to find a meta-model that captures the common structure shared across different robotic platforms such that data-efficient adaptation can be performed. We achieve such adaptation by introducing a learning framework consisting of a probabilistic gradient-based meta-learning algorithm that models the uncertainty arising from the few-shot setting with a low-dimensional latent variable. We experimentally evaluate our framework on a simulated reaching and a real-robot picking task using 400 simulated robots generated by varying the physical parameters of an existing set of robotic platforms. Our results show that the proposed method can successfully adapt a trained policy to different robotic platforms with novel physical parameters and the superiority of our meta-learning algorithm compared to state-of-the-art methods for the introduced few-shot policy adaptation problem.
Ali Ghadirzadeh, Xi Chen 0051, Petra Poklukar, Chelsea Finn, Mårten Björkman, Danica Kragic
IROS6
2021 Textile Taxonomy and Classification Using Pulling and Twisting
abstract
Identification of textile properties is an important milestone toward advanced robotic manipulation tasks that consider interaction with clothing items such as assisted dressing, laundry folding, automated sewing, textile recycling and reusing. Despite the abundance of work considering this class of deformable objects, many open problems remain. These relate to the choice and modelling of the sensory feedback as well as the control and planning of the interaction and manipulation strategies. Most importantly, there is no structured approach for studying and assessing different approaches that may bridge the gap between the robotics community and textile production industry. To this end, we outline a textile taxonomy considering fiber types and production methods, commonly used in textile industry. We devise datasets according to the taxonomy, and study how robotic actions, such as pulling and twisting of the textile samples, can be used for the classification. We also provide important insights from the perspective of visualization and interpretability of the gathered data.
Alberta Longhini, Michael C. Welle, Ioanna Mitsioni, Danica Kragic
IROS4
2021 Graph-based Task-specific Prediction Models for Interactions between Deformable and Rigid Objects
abstract
Capturing scene dynamics and predicting the future scene state is challenging but essential for robotic manipulation tasks, especially when the scene contains both rigid and deformable objects. In this work, we contribute a simulation environment and generate a novel dataset for task-specific manipulation, involving interactions between rigid objects and a deformable bag. The dataset incorporates a rich variety of scenarios including different object sizes, object numbers and manipulation actions. We approach dynamics learning by proposing an object-centric graph representation and two modules which are Active Prediction Module (APM) and Position Prediction Module (PPM) based on graph neural networks with an encode-process-decode architecture. At the inference stage, we build a two-stage model based on the learned modules for single time step prediction. We combine modules with different prediction horizons into a mixed-horizon model which addresses long-term prediction. In an ablation study, we show the benefits of the two-stage model for single time step prediction and the effectiveness of the mixed-horizon model for long-term prediction tasks. Supplementary material is available at https://github.com/wengzehang/deformable_rigid_interaction_prediction
Zehang Weng, Fabian Paus, Anastasiia Varava, Hang Yin 0001, Tamim Asfour, Danica Kragic
IROS6
2021 Learning Task Constraints in Visual-Action Planning from Demonstrations
abstract
Visual planning approaches have shown great success for decision making tasks with no explicit model of the state space. Learning a suitable representation and constructing a latent space where planning can be performed allows non-experts to setup and plan motions by just providing images. However, learned latent spaces are usually not semantically-interpretable, and thus it is difficult to integrate task constraints. We propose a novel framework to determine whether plans satisfy constraints given demonstrations of policies that satisfy or violate the constraints. The demonstrations are realizations of Linear Temporal Logic formulas which are employed to train Long Short-Term Memory (LSTM) networks directly in the latent space representation. We demonstrate that our architecture enables designers to easily specify, compose and integrate task constraints and achieves high performance in terms of accuracy. Furthermore, this visual planning framework enables human interaction, coping the environment changes that a human worker may involve. We show the flexibility of the method on a box pushing task in a simulated warehouse setting with different task constraints.
Francesco Esposito, Christian Pek, Michael C. Welle, Danica Kragic
RO-MAN4
2021 Graph-based Normalizing Flow for Human Motion Generation and Reconstruction
abstract
Data-driven approaches for modeling human skeletal motion have found various applications in interactive media and social robotics. Challenges remain in these fields for generating high-fidelity samples and robustly reconstructing motion from imperfect input data, due to e.g. missed marker detection. In this paper, we propose a probabilistic generative model to synthesize and reconstruct long horizon motion sequences conditioned on past information and control signals, such as the path along which an individual is moving. Our method adapts the existing work MoGlow by introducing a new graph-based model. The model leverages the spatial-temporal graph convolutional network (ST-GCN) to effectively capture the spatial structure and temporal correlation of skeletal motion data at multiple scales. We evaluate the models on a mixture of motion capture datasets of human locomotion with foot-step and bone-length analysis. The results demonstrate the advantages of our model in reconstructing missing markers and achieving comparable results on generating realistic future poses. When the inputs are imperfect, our model shows improvements on robustness of generation.
Hang Yin 0001, Danica Kragic, Mårten Björkman
RO-MAN3
2020 Discrete Bimanual Manipulation for Wrench Balancing
abstract
Dual-arm robots can overcome grasping force and payload limitations of a single arm by jointly grasping an object. However, if the distribution of mass of the grasped object is not even, each arm will experience different wrenches that can exceed its payload limits. In this work, we consider the problem of balancing the wrenches experienced by a dual-arm robot grasping a rigid tray. The distribution of wrenches among the robot arms changes due to objects being placed on the tray. We present an approach to reduce the wrench imbalance among arms through discrete bimanual manipulation. Our approach is based on sequential sliding motions of the grasp points on the surface of the object, to attain a more balanced configuration. We validate our modeling approach and system design through a set of robot experiments.
Silvia Cruciani, Diogo Almeida, Danica Kragic, Yiannis Karayiannidis
ICRA3
2020 In-Hand Manipulation of Objects with Unknown Shapes
abstract
This work addresses the problem of changing grasp configurations on objects with an unknown shape through in-hand manipulation. Our approach leverages shape priors, learned as deep generative models, to infer novel object shapes from partial visual sensing. The Dexterous Manipulation Graph method is extended to build incrementally and account for object shape uncertainty when planning a sequence of manipulation actions. We show that our approach successfully solves in-hand manipulation tasks with unknown objects, and demonstrate the validity of these solutions with robot experiments.
Silvia Cruciani, Hang Yin 0001, Danica Kragic
ICRA3
2020 Variational Auto-Regularized Alignment for Sim-to-Real Control
abstract
General-purpose simulators can be a valuable data source for flexible learning and control approaches. However, training models or control policies in simulation and then directly applying to hardware can yield brittle control. Instead, we propose a novel way to use simulators as regularizers. Our approach regularizes a decoder of a variational autoencoder to a black-box simulation, with the latent space bound to a subset of simulator parameters. This enables successful encoder training from a small number of real-world trajectories (10 in our experiments), yielding a latent space with simulation parameter distribution that matches the real-world setting. We use a learnable mixture for the latent prior/posterior, which implies a highly flexible class of densities for the posterior fit. Our approach is scalable and does not require restrictive distributional assumptions. We demonstrate ability to recover matching parameter distributions on a range of benchmarks, challenging custom simulation environments and several real-world scenarios. Our experiments using ABB YuMi robot hardware show ability to help reinforcement learning approaches overcome cases of severe sim-to-real mismatch.
Martin Hwasser, Danica Kragic, Rika Antonova
ICRA2
2020 Latent Space Roadmap for Visual Action Planning of Deformable and Rigid Object Manipulation
abstract
We present a framework for visual action planning of complex manipulation tasks with high-dimensional state spaces such as manipulation of deformable objects. Planning is performed in a low-dimensional latent state space that embeds images. We define and implement a Latent Space Roadmap (LSR) which is a graph-based structure that globally captures the latent system dynamics. Our framework consists of two main components: a Visual Foresight Module (VFM) that generates a visual plan as a sequence of images, and an Action Proposal Network (APN) that predicts the actions between them. We show the effectiveness of the method on a simulated box stacking task as well as a T-shirt folding task performed with a real robot.
Martina Lippi, Petra Poklukar, Michael C. Welle, Anastasiia Varava, Hang Yin 0001, Alessandro Marino, Danica Kragic
IROS7
2020 Multi-Object Rearrangement with Monte Carlo Tree Search: A Case Study on Planar Nonprehensile Sorting
abstract
In this work, we address a planar non-prehensile sorting task. Here, a robot needs to push many densely packed objects belonging to different classes into a configuration where these classes are clearly separated from each other. To achieve this, we propose to employ Monte Carlo tree search equipped with a task-specific heuristic function. We evaluate the algorithm on various simulated and real-world sorting tasks. We observe that the algorithm is capable of reliably sorting large numbers of convex and non-convex objects, as well as convex objects in the presence of immovable obstacles.
Haoran Song, Joshua A. Haustein, Weihao Yuan 0001, Kaiyu Hang, Michael Yu Wang, Danica Kragic, Johannes A. Stork
IROS6
2019 VPE: Variational Policy Embedding for Transfer Reinforcement Learning
abstract
Reinforcement Learning methods are capable of solving complex problems, but resulting policies might perform poorly in environments that are even slightly different. In robotics especially, training and deployment conditions often vary and data collection is expensive, making retraining undesirable. Simulation training allows for feasible training times, but on the other hand suffer from a reality-gap when applied in real-world settings. This raises the need of efficient adaptation of policies acting in new environments.We consider the problem of transferring knowledge within a family of similar Markov decision processes. We assume that Q-functions are generated by some low-dimensional latent variable. Given such a Q-function, we can find a master policy that can adapt given different values of this latent variable. Our method learns both the generative mapping and an approximate posterior of the latent variables, enabling identification of policies for new tasks by searching only in the latent space, rather than the space of all policies. The low-dimensional space, and master policy found by our method enables policies to quickly adapt to new environments. We demonstrate the method on both a pendulum swing-up task in simulation, and for simulation-to-real transfer on a pushing task.
Isac Arnekvist, Danica Kragic, Johannes A. Stork
ICRA2
2019 Reinforcement Learning in Topology-based Representation for Human Body Movement with Whole Arm Manipulation
abstract
Moving a human body or a large and bulky object may require the strength of whole arm manipulation (WAM). This type of manipulation places the load on the robot's arms and relies on global properties of the interaction to succeed- rather than local contacts such as grasping or non-prehensile pushing. In this paper, we learn to generate motions that enable WAM for holding and transporting of humans in certain rescue or patient care scenarios. We model the task as a reinforcement learning problem in order to provide a robot behavior that can directly respond to external perturbation and human motion. For this, we represent global properties of the robot-human interaction with topology-based coordinates that are computed from arm and torso positions. These coordinates also allow transferring the learned policy to other body shapes and sizes. For training and evaluation, we simulate a dynamic sea rescue scenario and show in quantitative experiments that the policy can solve unseen scenarios with differently-shaped humans, floating humans, or with perception noise. Our qualitative experiments show the subsequent transporting after holding is achieved and we demonstrate that the policy can be directly transferred to a real world setting.
Weihao Yuan 0001, Kaiyu Hang, Haoran Song, Danica Kragic, Michael Yu Wang, Johannes A. Stork
ICRA4
2019 Long-term Prediction of Motion Trajectories Using Path Homology Clusters
abstract
In order for robots to share their workspace with people, they need to reason about human motion efficiently. In this work we leverage large datasets of paths in order to infer local models that are able to perform long-term predictions of human motion. Further, since our method is based on simple dynamics, it is conceptually simple to understand and allows one to interpret the predictions produced, as well as to extract a cost function that can be used for planning. The main difference between our method and similar systems, is that we employ a map of the space and translate the motion of groups of paths into vector fields on that map. We test our method on synthetic data and show its performance on the Edinburgh forum pedestrian long-term tracking dataset [1] where we were able to outperform a Gaussian Mixture Model tasked with extracting dynamics from the paths.
J. Frederico Carvalho 0001, Mikael Vejdemo-Johansson, Florian T. Pokorny, Danica Kragic
IROS4
2019 Fast Adaptation with Meta-Reinforcement Learning for Trust Modelling in Human-Robot Interaction
abstract
In socially assistive robotics, an important research area is the development of adaptation techniques and their effect on human-robot interaction. We present a meta-learning based policy gradient method for addressing the problem of adaptation in human-robot interaction and also investigate its role as a mechanism for trust modelling. By building an escape room scenario in mixed reality with a robot, we test our hypothesis that bi-directional trust can be influenced by different adaptation algorithms. We found that our proposed model increased the perceived trustworthiness of the robot and influenced the dynamics of gaining human's trust. Additionally, participants evaluated that the robot perceived them as more trustworthy during the interactions with the meta-learning based adaptation compared to the previously studied statistical adaptation model.
Alex Yuan Gao, Elena Sibirtseva, Ginevra Castellano, Danica Kragic
IROS4
2019 Object Placement Planning and optimization for Robot Manipulators
abstract
We address the problem of planning the placement of a rigid object with a dual-arm robot in a cluttered environment. In this task, we need to locate a collision-free pose for the object that a) facilitates the stable placement of the object, b) is reachable by the robot and c) optimizes a user-given placement objective. In addition, we need to select which robot arm to perform the placement with. To solve this task, we propose an anytime algorithm that integrates sampling-based motion planning with a novel hierarchical search for suitable placement poses. Our algorithm incrementally produces approach motions to stable placement poses, reaching placements with better objective as runtime progresses. We evaluate our approach for two different placement objectives, and observe its effectiveness even in challenging scenarios.
Joshua A. Haustein, Kaiyu Hang, Johannes A. Stork, Danica Kragic
IROS4
2019 Learning to Estimate Pose and Shape of Hand-Held Objects from RGB Images
abstract
We develop a system for modeling hand-object interactions in 3D from RGB images that show a hand which is holding a novel object from a known category. We design a Convolutional Neural Network (CNN) for Hand-held Object Pose and Shape estimation called HOPS-Net and utilize prior work to estimate the hand pose and configuration. We leverage the insight that information about the hand facilitates object pose and shape estimation by incorporating the hand into both training and inference of the object pose and shape as well as the refinement of the estimated pose. The network is trained on a large synthetic dataset of objects in interaction with a human hand. To bridge the gap between real and synthetic images, we employ an image-to-image translation model (Augmented CycleGAN) that generates realistically textured objects given a synthetic rendering. This provides a scalable way of generating annotated data for training HOPS-Net. Our quantitative experiments show that even noisy hand parameters significantly help object pose and shape estimation. The qualitative experiments show results of pose and shape estimation of objects held by a hand “in the wild”.
Mia Kokic, Danica Kragic, Jeannette Bohg
IROS2
2019 Partial Caging: A Clearance-Based Definition and Deep Learning
abstract
Caging grasps limit the mobility of an object to a bounded component of configuration space. We introduce a notion of partial cage quality based on maximal clearance of an escaping path. As this is a computationally demanding task even in a two-dimensional scenario, we propose a deep learning approach. We design two convolutional neural networks and construct a pipeline for real-time partial cage quality estimation directly from 2D images of object models and planar caging tools. One neural network, CageMaskNN, is used to identify caging tool locations that can support partial cages, while a second network that we call CageClearanceNN is trained to predict the quality of those configurations. A dataset of 3811 images of objects and more than 19 million caging tool configurations is used to train and evaluate these networks on previously unseen objects and caging tool configurations. Furthermore, the networks are trained jointly on configurations for both 3 and 4 caging tool configurations whose shape varies along a 1-parameter family of increasing elongation. In experiments, we study how the networks' performance depends on the size of the training dataset, as well as how to efficiently deal with unevenly distributed training data. In further analysis, we show that the evaluation pipeline can approximately identify connected regions of successful caging tool placements and we evaluate the continuity of the cage quality score evaluation along caging tool trajectories. Experiments show that evaluation of a given configuration on a GeForce GTX 1080 GPU takes less than 6 ms.
Anastasiia Varava, Michael C. Welle, Jeffrey Mahler, Kenneth Y. Goldberg, Danica Kragic, Florian T. Pokomy
IROS5
2019 Robust Motion Planning for Non-holonomic Robots with Planar Geometric Constraints
Pouria Tajvar, Anastasiia Varava, Danica Kragic, Jana Tumova
ISRR3
2018 Evaluating the Quality of Non-Prehensile Balancing Grasps
abstract
Assessing grasp quality and, subsequently, predicting grasp success is useful for avoiding failures in many autonomous robotic applications. In addition, interest in nonprehensile grasping and manipulation has been growing as it offers the potential for a large increase in dexterity. However, while force-closure grasping has been the subject of intense study for many years, few existing works have considered quality metrics for non-prehensile grasps. Furthermore, no studies exist to validate them in practice. In this work we use a real-world data set of non-prehensile balancing grasps and use it to experimentally validate a wrench-based quality metric by means of its grasp success prediction capability. The overall accuracy of up to 84 % is encouraging and in line with existing results for force-closure grasps.
Robert Krug 0002, Yasemin Bekiroglu, Danica Kragic, Máximo A. Roa
ICRA3
2018 Anticipating Many Futures: Online Human Motion Prediction and Generation for Human-Robot Interaction
abstract
Fluent and safe interactions of humans and robots require both partners to anticipate the others' actions. The bottleneck of most methods is the lack of an accurate model of natural human motion. In this work, we present a conditional variational autoencoder that is trained to predict a window of future human motion given a window of past frames. Using skeletal data obtained from RGB depth images, we show how this unsupervised approach can be used for online motion prediction for up to 1660 ms. Additionally, we demonstrate online target prediction within the first 300-500 ms after motion onset without the use of target specific training data. The advantage of our probabilistic approach is the possibility to draw samples of possible future motion patterns. Finally, we investigate how movements and kinematic cues are represented on the learned low dimensional manifold.
Judith Bütepage, Hedvig Kjellström, Danica Kragic
ICRA3
2018 Path Clustering with Homology Area
abstract
Path clustering has found many applications in recent years. Common approaches to this problem use aggregates of the distances between points to provide a measure of dissimilarity between paths which do not satisfy the triangle inequality. Furthermore, they do not take into account the topology of the space where the paths are embedded. To tackle this, we extend previous work in path clustering with relative homology, by employing minimum homology area as a measure of distance between homologous paths in a triangulated mesh. Further, we show that the resulting distance satisfies the triangle inequality, and how we can exploit the properties of homology to reduce the amount of pairwise distance calculations necessary to cluster a set of paths. We further compare the output of our algorithm with that of DTW on a toy dataset of paths, as well as on a dataset of real-world paths.
J. Frederico Carvalho 0001, Mikael Vejdemo-Johansson, Danica Kragic, Florian T. Pokorny
ICRA3
2018 Rearrangement with Nonprehensile Manipulation Using Deep Reinforcement Learning
abstract
Rearranging objects on a tabletop surface by means of nonprehensile manipulation is a task which requires skillful interaction with the physical world. Usually, this is achieved by precisely modeling physical properties of the objects, robot, and the environment for explicit planning. In contrast, as explicitly modeling the physical environment is not always feasible and involves various uncertainties, we learn a nonprehensile rearrangement strategy with deep reinforcement learning based on only visual feedback. For this, we model the task with rewards and train a deep Q-network. Our potential field-based heuristic exploration strategy reduces the amount of collisions which lead to suboptimal outcomes and we actively balance the training set to avoid bias towards poor examples. Our training process leads to quicker learning and better performance on the task as compared to uniform exploration and standard experience replay. We demonstrate empirical evidence from simulation that our method leads to a success rate of 85%, show that our system can cope with sudden changes of the environment, and compare our performance with human level performance.
Weihao Yuan 0001, Johannes A. Stork, Danica Kragic, Michael Yu Wang, Kaiyu Hang
ICRA3
2018 Interactive, Collaborative Robots: Challenges and Opportunities
abstract
Robotic technology has transformed manufacturing industry ever since the first industrial robot was put in use in the beginning of the 60s. The challenge of developing flexible solutions where production lines can be quickly re-planned, adapted and structured for new or slightly changed products is still an important open problem. Industrial robots today are still largely preprogrammed for their tasks, not able to detect errors in their own performance or to robustly interact with a complex environment and a human worker. The challenges are even more serious when it comes to various types of service robots. Full robot autonomy, including natural interaction, learning from and with human, safe and flexible performance for challenging tasks in unstructured environments will remain out of reach for the foreseeable future. In the envisioned future factory setups, home and office environments, humans and robots will share the same workspace and perform different object manipulation tasks in a collaborative manner. We discuss some of the major challenges of developing such systems and provide examples of the current state of the art.
Danica Kragic, Joakim Gustafson, Hakan Karaoguz, Patric Jensfelt, Robert Krug 0002
IJCAI1
2018 Dexterous Manipulation Graphs
abstract
We propose the Dexterous Manipulation Graph as a tool to address in-hand manipulation and reposition an object inside a robot's end-effector. This graph is used to plan a sequence of manipulation primitives so to bring the object to the desired end pose. This sequence of primitives is translated into motions of the robot to move the object held by the end-effector. We use a dual arm robot with parallel grippers to test our method on a real system and show successful planning and execution of in-hand manipulation.
Silvia Cruciani, Christian Smith, Danica Kragic, Kaiyu Hang
IROS3
2018 A Comparison of Visualisation Methods for Disambiguating Verbal Requests in Human-Robot Interaction
abstract
Picking up objects requested by a human user is a common task in human-robot interaction. When multiple objects match the user's verbal description, the robot needs to clarify which object the user is referring to before executing the action. Previous research has focused on perceiving user's multimodal behaviour to complement verbal commands or minimising the number of follow up questions to reduce task time. In this paper, we propose a system for reference disambiguation based on visualisation and compare three methods to disambiguate natural language instructions. In a controlled experiment with a YuMi robot, we investigated realtime augmentations of the workspace in three conditions - head-mounted display, projector, and a monitor as the baseline - using objective measures such as time and accuracy, and subjective measures like engagement, immersion, and display interference. Significant differences were found in accuracy and engagement between the conditions, but no differences were found in task time. Despite the higher error rates in the head-mounted display condition, participants found that modality more engaging than the other two, but overall showed preference for the projector condition over the monitor and head-mounted display conditions.
Elena Sibirtseva, Dimosthenis Kontogiorgos, Olov Nykvist, Hakan Karaoguz, Iolanda Leite, Joakim Gustafson, Danica Kragic
RO-MAN7
2018 Free Space of Rigid Objects: Caging, Path Non-existence, and Narrow Passage Detection
Anastasiia Varava, J. Frederico Carvalho 0001, Florian T. Pokorny, Danica Kragic
WAFR4
2017 Deep Representation Learning for Human Motion Prediction and Classification
abstract
Generative models of 3D human motion are often restricted to a small number of activities and can therefore not generalize well to novel movements or applications. In this work we propose a deep learning framework for human motion capture data that learns a generic representation from a large corpus of motion capture data and generalizes well to new, unseen, motions. Using an encoding-decoding network that learns to predict future 3D poses from the most recent past, we extract a feature representation of human motion. Most work on deep learning for sequence prediction focuses on video and speech. Since skeletal data has a different structure, we present and evaluate different network architectures that make different assumptions about time dependencies and limb correlations. To quantify the learned features, we use the output of different layers for action classification and visualize the receptive fields of the network units. Our method outperforms the recent state of the art in skeletal motion prediction even though these use action specific training data. Our results show that deep feedforward networks, trained from a generic mocap database, can successfully be used for feature extraction from human motion data and that this representation can be used as a foundation for classification and prediction.
Judith Bütepage, Michael J. Black, Danica Kragic, Hedvig Kjellström
CVPR3
2017 Acting, Interacting, Collaborative Robots
abstract
The current trend in computer vision is development of data-driven approaches where the use of large amounts of data tries to compensate for the complexity of the world captured by cameras. Are these approaches also viable solutions in robotics? Apart from 'seeing', a robot is capable of acting, thus purposively change what and how it sees the world around it. There is a need for an interplay between processes such as attention, segmentation, object detection, recognition and categorization in order to interact with the environment. In addition, the parameterization of these is inevitably guided by the task or the goal a robot is supposed to achieve. In this talk, I will present the current state of the art in the area of robot vision and discuss open problems in the area. I will also show how visual input can be integrated with proprioception, tactile and force-torque feedback in order to plan, guide and assess robot's action and interaction with the environment. Interaction between two agents builds on the ability to engage in mutual prediction and signaling. Thus, human-robot interaction requires a system that can interpret and make use of human signaling strategies in a social context. Our work in this area focuses on developing a framework for human motion prediction in the context of joint action in HRI. We base this framework on the idea that social interaction is highly influences by sensorimotor contingencies (SMCs). Instead of constructing explicit cognitive models, we rely on the interaction between actions the perceptual change that they induce in both the human and the robot. This approach allows us to employ a single model for motion prediction and goal inference and to seamlessly integrate the human actions into the environment and task context.
Danica Kragic
HRI1
2017 Collaborative robots: from action and interaction to collaboration (keynote)
abstract
The integral ability of any robot is to act in the environment, interact and collaborate with people and other robots. Interaction between two agents builds on the ability to engage in mutual prediction and signaling. Thus, human-robot interaction requires a system that can interpret and make use of human signaling strategies in a social context. Our work in this area focuses on developing a framework for human motion prediction in the context of joint action in HRI. We base this framework on the idea that social interaction is highly influences by sensorimotor contingencies (SMCs). Instead of constructing explicit cognitive models, we rely on the interaction between actions the perceptual change that they induce in both the human and the robot. This approach allows us to employ a single model for motion prediction and goal inference and to seamlessly integrate the human actions into the environment and task context. The current trend in computer vision is development of data-driven approaches where the use of large amounts of data tries to compensate for the complexity of the world captured by cameras. Are these approaches also viable solutions in robotics? Apart from 'seeing', a robot is capable of acting, thus purposively change what and how it sees the world around it. There is a need for an interplay between processes such as attention, segmentation, object detection, recognition and categorization in order to interact with the environment. In addition, the parameterization of these is inevitably guided by the task or the goal a robot is supposed to achieve. In this talk, I will present the current state of the art in the area of robot vision and discuss open problems in the area. I will also show how visual input can be integrated with proprioception, tactile and force-torque feedback in order to plan, guide and assess robot's action and interaction with the environment. We employ a deep generative model that makes inferences over future human motion trajectories given the intention of the human and the history as well as the task setting of the interaction. With help predictions drawn from the model, we can determine the most likely future motion trajectory and make inferences over intentions and objects of interest.
Danica Kragic
ICMI1
2017 Integrating motion and hierarchical fingertip grasp planning
abstract
In this work, we present an algorithm that simultaneously searches for a high quality fingertip grasp and a collision-free path for a robot hand-arm system to achieve it. The algorithm combines a bidirectional sampling-based motion planning approach with a hierarchical contact optimization process. Rather than tackling these problems in a decoupled manner, the grasp optimization is guided by the proximity to collision-free configurations explored by the motion planner. We implemented the algorithm for a 13-DoF manipulator and show that it is capable of efficiently planning reachable high quality grasps in cluttered environments. Further, we show that our algorithm outperforms a decoupled integration in terms of planning runtime.
Joshua A. Haustein, Kaiyu Hang, Danica Kragic
ICRA3
2017 Deep predictive policy training using reinforcement learning
abstract
Skilled robot task learning is best implemented by predictive action policies due to the inherent latency of sensorimotor processes. However, training such predictive policies is challenging as it involves finding a trajectory of motor activations for the full duration of the action. We propose a data-efficient deep predictive policy training (DPPT) framework with a deep neural network policy architecture which maps an image observation to a sequence of motor activations. The architecture consists of three sub-networks referred to as the perception, policy and behavior super-layers. The perception and behavior super-layers force an abstraction of visual and motor data trained with synthetic and simulated training samples, respectively. The policy super-layer is a small subnetwork with fewer parameters that maps data in-between the abstracted manifolds. It is trained for each task using methods for policy search reinforcement learning. We demonstrate the suitability of the proposed architecture and learning framework by training predictive policies for skilled object grasping and ball throwing on a PR2 robot. The effectiveness of the method is illustrated by the fact that these tasks are trained using only about 180 real robot attempts with qualitative terminal rewards.
Ali Ghadirzadeh, Atsuto Maki, Danica Kragic, Mårten Björkman
IROS3
2017 Estimating deformability of objects using meshless shape matching
abstract
Humans interact with deformable objects on a daily basis but this still represents a challenge for robots. To enable manipulation of and interaction with deformable objects, robots need to be able to extract and learn the deformability of objects both prior to and during the interaction. Physics-based models are commonly used to predict the physical properties of deformable objects and simulate their deformation accurately. The most popular simulation techniques are force-based models that need force measurements. In this paper, we explore the applicability of a geometry-based simulation method called meshless shape matching (MSM) for estimating the deformability of objects. The main advantages of MSM are its controllability and computational efficiency that make it popular in computer graphics to simulate complex interactions of multiple objects at the same time. Additionally, a useful feature of the MSM that differentiates it from other physics-based simulation is to be independent of force measurements that may not be available to a robotic framework lacking force/torque sensors. In this work, we design a method to estimate deformability based on certain properties, such as volume conservation. Using the finite element method (FEM) we create the ground truth deformability for various settings to evaluate our method. The experimental evaluation shows that our approach is able to accurately identify the deformability of test objects, supporting the value of MSM for robotic applications.
Püren Güler, Alessandro Pieropan, Masatoshi Ishikawa, Danica Kragic
IROS4
2017 Caging and Path Non-existence: A Deterministic Sampling-Based Verification Algorithm
Anastasiia Varava, J. Frederico Carvalho 0001, Florian T. Pokorny, Danica Kragic
ISRR4
2017 Interactive Perception: Leveraging Action in Perception and Perception in Action
abstract
Recent approaches in robot perception follow the insight that perception is facilitated by interaction with the environment. These approaches are subsumed under the term Interactive Perception (IP). This view of perception provides the following benefits. First, interaction with the environment creates a rich sensory signal that would otherwise not be present. Second, knowledge of the regularity in the combined space of sensory data and action parameters facilitates the prediction and interpretation of the sensory signal. In this survey, we postulate this as a principle for robot perception and collect evidence in its support by analyzing and categorizing existing work in this area. We also provide an overview of the most important applications of IP. We close this survey by discussing remaining open questions. With this survey, we hope to help define the field of Interactive Perception and to provide a valuable resource for future research.
Jeannette Bohg, Karol Hausman, Bharath Sankaran, Oliver Brock, Danica Kragic, Stefan Schaal, Gaurav S. Sukhatme
IEEE Trans. Robotics5
2016 Social Affordance Tracking over Time - A Sensorimotor Account of False-Belief Tasks
Judith Bütepage, Hedvig Kjellström, Danica Kragic
CogSci3
2016 Analytic grasp success prediction with tactile feedback
abstract
Predicting grasp success is useful for avoiding failures in many robotic applications. Based on reasoning in wrench space, we address the question of how well analytic grasp success prediction works if tactile feedback is incorporated. Tactile information can alleviate contact placement uncertainties and facilitates contact modeling. We introduce a wrench-based classifier and evaluate it on a large set of real grasps. The key finding of this work is that exploiting tactile information allows wrench-based reasoning to perform on a level with existing methods based on learning or simulation. Different from these methods, the suggested approach has no need for training data, requires little modeling effort and is computationally efficient. Furthermore, our method affords task generalization by considering the capabilities of the grasping device and expected disturbance forces/moments in a physically meaningful way.
Robert Krug 0002, Achim J. Lilienthal, Danica Kragic, Yasemin Bekiroglu
ICRA3
2016 Adaptive control for pivoting with visual and tactile feedback
abstract
In this work we present an adaptive control approach for pivoting, which is an in-hand manipulation maneuver that consists of rotating a grasped object to a desired orientation relative to the robot's hand. We perform pivoting by means of gravity, allowing the object to rotate between the fingers of a one degree of freedom gripper and controlling the gripping force to ensure that the object follows a reference trajectory and arrives at the desired angular position. We use a visual pose estimation system to track the pose of the object and force measurements from tactile sensors to control the gripping force. The adaptive controller employs an update law that accommodates for errors in the friction coefficient, which is one of the most common sources of uncertainty in manipulation. Our experiments confirm that the proposed adaptive controller successfully pivots a grasped object in the presence of uncertainty in the object's friction parameters.
Francisco E. Vina, Yiannis Karayiannidis, Christian Smith, Danica Kragic
ICRA4
2016 Probabilistic consolidation of grasp experience
abstract
We present a probabilistic model for joint representation of several sensory modalities and action parameters in a robotic grasping scenario. Our non-linear probabilistic latent variable model encodes relationships between grasp-related parameters, learns the importance of features, and expresses confidence in estimates. The model learns associations between stable and unstable grasps that it experiences during an exploration phase. We demonstrate the applicability of the model for estimating grasp stability, correcting grasps, identifying objects based on tactile imprints and predicting tactile imprints from object-relative gripper poses. We performed experiments on a real platform with both known and novel objects, i.e., objects the robot trained with, and previously unseen objects. Grasp correction had a 75% success rate on known objects, and 73% on new objects. We compared our model to a traditional regression model that succeeded in correcting grasps in only 38% of cases.
Yasemin Bekiroglu, Andreas Damianou, Renaud Detry, Johannes A. Stork, Danica Kragic, Carl Henrik Ek
ICRA5
2016 Self-learning and adaptation in a sensorimotor framework
abstract
We present a general framework to autonomously achieve the task of finding a sequence of actions that result in a desired state. Autonomy is acquired by learning sensorimotor patterns of a robot, while it is interacting with its environment.
Ali Ghadirzadeh, Judith Bütepage, Danica Kragic, Mårten Björkman
ICRA3
2016 On the evolution of fingertip grasping manifolds
abstract
Efficient and accurate planning of fingertip grasps is essential for dexterous in-hand manipulation. In this work, we present a system for fingertip grasp planning that incrementally learns a heuristic for hand reachability and multi-fingered inverse kinematics. The system consists of an online execution module and an offline optimization module. During execution the system plans and executes fingertip grasps using Canny's grasp quality metric and a learned random forest based hand reachability heuristic. In the offline module, this heuristic is improved based on a grasping manifold that is incrementally learned from the experiences collected during execution. The system is evaluated both in simulation and on a Schunk-SDH dexterous hand mounted on a KUKA-KR5 arm. We show that, as the grasping manifold is adapted to the system's experiences, the heuristic becomes more accurate, which results in an improved performance of the execution module. The improvement is not only observed for experienced objects, but also for previously unknown objects of similar sizes.
Kaiyu Hang, Joshua A. Haustein, Miao Li 0002, Aude Billard, Christian Smith, Danica Kragic
ICRA6
2016 Integrated on-line robot-camera calibration and object pose estimation
abstract
We present a novel on-line approach for extrinsic robot-camera calibration, a process often referred to as hand-eye calibration, that uses object pose estimates from a real-time model-based tracking approach. While off-line calibration has seen much progress recently due to the incorporation of bundle adjustment techniques, on-line calibration still remains a largely open problem. Since we update the calibration in each frame, the improvements can be incorporated immediately in the pose estimation itself to facilitate object tracking. Our method does not require the camera to observe the robot or to have markers at certain fixed locations on the robot. To comply with a limited computational budget, it maintains a fixed size configuration set of samples. This set is updated each frame in order to maximize an observability criterion. We show that a set of size 20 is sufficient in real-world scenarios with static and actuated cameras. With this set size, only 100 microseconds are required to update the calibration in each frame, and we typically achieve accurate robot-camera calibration in 10 to 20 seconds. Together, these characteristics enable the incorporation of calibration in normal task execution.
Karl Pauwels, Danica Kragic
ICRA2
2016 Robust tracking of unknown objects through adaptive size estimation and appearance learning
abstract
This work employs an adaptive learning mechanism to perform tracking of an unknown object through RGBD cameras. We extend our previous framework to robustly track a wider range of arbitrarily shaped objects by adapting the model to the measured object size. The size is estimated as the object undergoes motion, which is done by fitting an inscribed cuboid to the measurements. The region spanned by this cuboid is used during tracking, to determine whether or not new measurements should be added to the object model. In our experiments we test our tracker with a set of objects of arbitrary shape and we show the benefit of the proposed model due to its ability to adapt to the object shape which leads to more robust tracking results.
Alessandro Pieropan, Niklas Bergström, Masatoshi Ishikawa, Danica Kragic, Hedvig Kjellström
ICRA4
2016 Topological trajectory clustering with relative persistent homology
abstract
Cloud Robotics techniques based on Learning from Demonstrations suggest promising alternatives to manual programming of robots and autonomous vehicles. One challenge is that demonstrated trajectories may vary dramatically: it can be very difficult, if not impossible, for a system to learn control policies unless the trajectories are clustered into meaningful consistent subsets. Metric clustering methods, based on a distance measure, require quadratic time to compute a pairwise distance matrix and do not naturally distinguish topologically distinct trajectories. This paper presents an algorithm for topological clustering based on relative persistent homology, which, for a fixed underlying simplicial representation and discretization of trajectories, requires only linear time in the number of trajectories. The algorithm incorporates global constraints formalized in terms of the topology of sublevel or superlevel sets of a function and can be extended to incorporate probabilistic motion models. In experiments with real automobile and ship GPS trajectories as well as pedestrian trajectories extracted from video, the algorithm clusters trajectories into meaningful consistent subsets and, as we show in an experiment with ship trajectories, results in a faster and more efficient clustering than a metric clustering by Fréchet distance.
Florian T. Pokorny, Kenneth Y. Goldberg, Danica Kragic
ICRA3
2016 High-dimensional Winding-Augmented Motion Planning with 2D topological task projections and persistent homology
abstract
Recent progress in motion planning has made it possible to determine homotopy inequivalent trajectories between an initial and terminal configuration in a robot configuration space. Current approaches have however either assumed the knowledge of differential one-forms related to a skeletonization of the collision space, or have relied on a simplicial representation of the free space. Both of these approaches are currently however not yet practical for higher dimensional configuration spaces. We propose 2D topological task projections (TTPs): mappings from the configuration space to 2-dimensional spaces where simplicial complex filtrations and persistent homology can identify topological properties of the high-dimensional free configuration space. Our approach only requires the availability of collision free samples to identify winding centers that can be used to determine homotopy inequivalent trajectories. We propose the Winding Augmented RRT and RRT* (WA-RRT/RRT*) algorithms using which homotopy inequivalent trajectories can be found. We evaluate our approach in experiments with configuration spaces of planar linkages with 2-10 degrees of freedom. Results indicate that our approach can reliably identify suitable topological task projections and our proposed WA-RRT and WA-RRT* algorithms were able to identify a collection of homotopy inequivalent trajectories in each considered configuration space dimension.
Florian T. Pokorny, Danica Kragic, Lydia E. Kavraki, Kenneth Y. Goldberg
ICRA2
2016 Active exploration using Gaussian Random Fields and Gaussian Process Implicit Surfaces
abstract
In this work we study the problem of exploring surfaces and building compact 3D representations of the environment surrounding a robot through active perception. We propose an online probabilistic framework that merges visual and tactile measurements using Gaussian Random Field and Gaussian Process Implicit Surfaces. The system investigates incomplete point clouds in order to find a small set of regions of interest which are then physically explored with a robotic arm equipped with tactile sensors. We show experimental results obtained using a PrimeSense camera, a Kinova Jaco2 robotic arm and Optoforce sensors on different scenarios. We then demostrate how to use the online framework for object detection and terrain classification.
Sergio Caccamo, Yasemin Bekiroglu, Carl Henrik Ek, Danica Kragic
IROS4
2016 A sensorimotor reinforcement learning framework for physical Human-Robot Interaction
abstract
Modeling of physical human-robot collaborations is generally a challenging problem due to the unpredictive nature of human behavior. To address this issue, we present a data-efficient reinforcement learning framework which enables a robot to learn how to collaborate with a human partner. The robot learns the task from its own sensorimotor experiences in an unsupervised manner. The uncertainty in the interaction is modeled using Gaussian processes (GP) to implement a forward model and an action-value function. Optimal action selection given the uncertain GP model is ensured by Bayesian optimization. We apply the framework to a scenario in which a human and a PR2 robot jointly control the ball position on a plank based on vision and force/torque data. Our experimental results show the suitability of the proposed method in terms of fast and data-efficient model learning, optimal action selection under uncertainty and equal role sharing between the partners.
Ali Ghadirzadeh, Judith Bütepage, Atsuto Maki, Danica Kragic, Mårten Björkman
IROS4
2016 The GRASP Taxonomy of Human Grasp Types
abstract
In this paper, we analyze and compare existing human grasp taxonomies and synthesize them into a single new taxonomy (dubbed “The GRASP Taxonomy” after the GRASP project funded by the European Commission). We consider only static and stable grasps performed by one hand. The goal is to extract the largest set of different grasps that were referenced in the literature and arrange them in a systematic way. The taxonomy provides a common terminology to define human hand configurations and is important in many domains such as human-computer interaction and tangible user interfaces where an understanding of the human is basis for a proper interface. Overall, 33 different grasp types are found and arranged into the GRASP taxonomy. Within the taxonomy, grasps are arranged according to 1) opposition type, 2) the virtual finger assignments, 3) type in terms of power, precision, or intermediate grasp, and 4) the position of the thumb. The resulting taxonomy incorporates all grasps found in the reviewed taxonomies that complied with the grasp definition. We also show that due to the nature of the classification, the 33 grasp types might be reduced to a set of 17 more general grasps if only the hand configuration is considered without the object shape/size.
Thomas Feix, Javier Romero 0002, Heinz-Bodo Schmiedmayer, Aaron M. Dollar, Danica Kragic
IEEE Trans. Hum. Mach. Syst.5
2016 Hierarchical Fingertip Space: A Unified Framework for Grasp Planning and In-Hand Grasp Adaptation
abstract
We present a unified framework for grasp planning and in-hand grasp adaptation using visual, tactile, and proprioceptive feedback. The main objective of the proposed framework is to enable fingertip grasping by addressing problems of changed weight of the object, slippage, and external disturbances. For this purpose we introduce the Hierarchical Fingertip Space as a representation enabling optimization for both efficient grasp synthesis and online finger gaiting. Grasp synthesis is followed by a grasp adaptation step that consists of both grasp force adaptation through impedance control and regrasping/finger gaiting when the former is not sufficient. Experimental evaluation is conducted on an Allegro hand mounted on a Kuka LWR arm.
Kaiyu Hang, Miao Li 0002, Johannes A. Stork, Yasemin Bekiroglu, Florian T. Pokorny, Aude Billard, Danica Kragic
IEEE Trans. Robotics7
2016 An Adaptive Control Approach for Opening Doors and Drawers Under Uncertainties
abstract
We study the problem of robot interaction with mechanisms that afford one degree of freedom motion, e.g., doors and drawers. We propose a methodology for simultaneous compliant interaction and estimation of constraints imposed by the joint. Our method requires no prior knowledge of the mechanisms' kinematics, including the type of joint, prismatic or revolute. The method consists of a velocity controller that relies on force/torque measurements and estimation of the motion direction, the distance, and the orientation of the rotational axis. It is suitable for velocity-controlled manipulators with force/torque sensor capabilities at the end-effector. Forces and torques are regulated within given constraints, while the velocity controller ensures that the end-effector of the robot moves with a task-related desired velocity. We give proof that the estimates converge to the true values under valid assumptions on the grasp, and error bounds for setups with inaccuracies in control, measurements, or modeling. The method is evaluated in different scenarios involving opening a representative set of door and drawer mechanisms found in household environments.
Yiannis Karayiannidis, Christian Smith, Francisco E. Vina, Petter Ögren, Danica Kragic
IEEE Trans. Robotics5
2016 Caging Grasps of Rigid and Partially Deformable 3-D Objects With Double Fork and Neck Features
abstract
Caging provides an alternative to point-contact-based rigid grasping, relying on reasoning about the global free configuration space of an object under consideration. While substantial progress has been made toward the analysis, verification, and synthesis of cages of polygonal objects in the plane, the use of caging as a tool for manipulating general complex objects in 3-D remains challenging. In this work, we introduce the problem of caging rigid and partially deformable 3-D objects, which exhibit geometric features we call double forks and necks. Our approach is based on the linking number-a classical topological invariant, allowing us to determine sufficient conditions for caging objects with these features even in the case when the object under consideration is partially deformable under a set of neck or double fork preserving deformations. We present synthesis and verification algorithms and demonstrations of applying these algorithms to cage 3-D meshes.
Anastasiia Varava, Danica Kragic, Florian T. Pokorny
IEEE Trans. Robotics2
2015 Learning the tactile signatures of prototypical object parts for robust part-based grasping of novel objects
abstract
We present a robotic agent that learns to derive object grasp stability from touch. The main contribution of our work is the use of a characterization of the shape of the part of the object that is enclosed by the gripper to condition the tactile-based stability model. As a result, the agent is able to express that a specific tactile signature may for instance indicate stability when grasping a cylinder, while cuing instability when grasping a box. We proceed by (1) discretizing the space of graspable object parts into a small set of prototypical shapes, via a data-driven clustering process, and (2) learning a touch-based stability classifier for each prototype. Classification is conducted through kernel logistic regression, applied to a low-dimensional approximation of the tactile data read from the robot's hand. We present an experiment that demonstrates the applicability of the method, yielding a success rate of 89%. Our experiment also shows that the distribution of tactile data differs substantially between grasps collected with different prototypes, supporting the use of shape cues in touch-based stability estimators.
Emil Hyttinen, Danica Kragic, Renaud Detry
ICRA2
2015 Learning Predictive State Representation for in-hand manipulation
abstract
We study the use of Predictive State Representation (PSR) for modeling of an in-hand manipulation task through interaction with the environment. We extend the original PSR model to a new domain of in-hand manipulation and address the problem of partial observability by introducing new kernel-based features that integrate both actions and observations. The model is learned directly from haptic data and is used to plan series of actions that rotate the object in the hand to a specific configuration by pushing it against a table. Further, we analyze the model's belief states using additional visual data and enable planning of action sequences when the observations are ambiguous. We show that the learned representation is geometrically meaningful by embedding labeled action-observation traces. Suitability for planning is demonstrated by a post-grasp manipulation example that changes the object state to multiple specified target configurations.
Johannes A. Stork, Carl Henrik Ek, Yasemin Bekiroglu, Danica Kragic
ICRA4
2015 Learning Human Priors for Task-Constrained Grasping
Martin Hjelm, Carl Henrik Ek, Renaud Detry, Danica Kragic
ICVS4
2015 In-hand manipulation using gravity and controlled slip
abstract
In this work we propose a sliding mode controller for in-hand manipulation that repositions a tool in the robot's hand by using gravity and controlling the slippage of the tool. In our approach, the robot holds the tool with a pinch grasp and we model the system as a link attached to the gripper via a passive revolute joint with friction, i.e., the grasp only affords rotational motions of the tool around a given axis of rotation. The robot controls the slippage by varying the opening between the fingers in order to allow the tool to move to the desired angular position following a reference trajectory. We show experimentally how the proposed controller achieves convergence to the desired tool orientation under variations of the tool's inertial parameters.
Francisco E. Vina, Yiannis Karayiannidis, Karl Pauwels, Christian Smith, Danica Kragic
IROS5
2015 SimTrack: A simulation-based framework for scalable real-time object pose detection and tracking
abstract
We propose a novel approach for real-time object pose detection and tracking that is highly scalable in terms of the number of objects tracked and the number of cameras observing the scene. Key to this scalability is a high degree of parallelism in the algorithms employed. The method maintains a single 3D simulated model of the scene consisting of multiple objects together with a robot operating on them. This allows for rapid synthesis of appearance, depth, and occlusion information from each camera viewpoint. This information is used both for updating the pose estimates and for extracting the low-level visual cues. The visual cues obtained from each camera are efficiently fused back into the single consistent scene representation using a constrained optimization method. The centralized scene representation, together with the reliability measures it enables, simplify the interaction between pose tracking and pose detection across multiple cameras. We demonstrate the robustness of our approach in a realistic manipulation scenario. We publicly release this work as a part of a general ROS software framework for real-time pose estimation, SimTrack, that can be integrated easily for different robotic applications.
Karl Pauwels, Danica Kragic
IROS2
2015 Learning Predictive State Representations for planning
abstract
Predictive State Representations (PSRs) allow modeling of dynamical systems directly in observables and without relying on latent variable representations. A problem that arises from learning PSRs is that it is often hard to attribute semantic meaning to the learned representation. This makes generalization and planning in PSRs challenging. In this paper, we extend PSRs and introduce the notion of PSRs that include prior information (P-PSRs) to learn representations which are suitable for planning and interpretation. By learning a low-dimensional embedding of test features we map belief points of similar semantic to the same region of a subspace. This facilitates better generalization for planning and semantical interpretation of the learned representation. In specific, we show how to overcome the training sample bias and introduce feature selection such that the resulting representation emphasizes observables related to the planning task. We show that our P-PSRs result in qualitatively meaningful representations and present quantitative results that indicate improved suitability for planning.
Johannes A. Stork, Carl Henrik Ek, Danica Kragic
IROS3
2015 Task-Based Robot Grasp Planning Using Probabilistic Inference
abstract
Grasping and manipulating everyday objects in a goal-directed manner is an important ability of a service robot. The robot needs to reason about task requirements and ground these in the sensorimotor information. Grasping and interaction with objects are challenging in real-world scenarios, where sensorimotor uncertainty is prevalent. This paper presents a probabilistic framework for the representation and modeling of robot-grasping tasks. The framework consists of Gaussian mixture models for generic data discretization, and discrete Bayesian networks for encoding the probabilistic relations among various task-relevant variables, including object and action features as well as task constraints. We evaluate the framework using a grasp database generated in a simulated environment including a human and two robot hand models. The generative modeling approach allows the prediction of grasping tasks given uncertain sensory data, as well as object and grasp selection in a task-oriented manner. Furthermore, the graphical model framework provides insights into dependencies between variables and features relevant for object grasping.
Dan Song 0002, Carl Henrik Ek, Kai Huebner, Danica Kragic
IEEE Trans. Robotics4
2014 Combinatorial optimization for hierarchical contact-level grasping
abstract
We address the problem of generating force-closed point contact grasps on complex surfaces and model it as a combinatorial optimization problem. Using a multilevel refinement metaheuristic, we maximize the quality of a grasp subject to a reachability constraint by recursively forming a hierarchy of increasingly coarser optimization problems. A grasp is initialized at the top of the hierarchy and then locally refined until convergence at each level. Our approach efficiently addresses the high dimensional problem of synthesizing stable point contact grasps while resulting in stable grasps from arbitrary initial configurations. Compared to a sampling-based approach, our method yields grasps with higher grasp quality. Empirical results are presented for a set of different objects. We investigate the number of levels in the hierarchy, the computational complexity, and the performance relative to a random sampling baseline approach.
Kaiyu Hang, Johannes A. Stork, Florian T. Pokorny, Danica Kragic
ICRA4
2014 Representations for cross-task, cross-object grasp transfer
abstract
We address the problem of transferring grasp knowledge across objects and tasks. This means dealing with two important issues: 1) the induction of possible transfers, i.e., whether a given object affords a given task, and 2) the planning of a grasp that will allow the robot to fulfill the task. The induction of object affordances is approached by abstracting the sensory input of an object as a set of attributes that the agent can reason about through similarity and proximity. For grasp execution, we combine a part-based grasp planner with a model of task constraints. The task constraint model indicates areas of the object that the robot can grasp to execute the task. Within these areas, the part-based planner finds a hand placement that is compatible with the object shape. The key contribution is the ability to transfer task parameters across objects while the part-based grasp planner allows for transferring grasp information across tasks. As a result, the robot is able to synthesize plans for previously unobserved task/object combinations. We illustrate our approach with experiments conducted on a real robot.
Martin Hjelm, Renaud Detry, Carl Henrik Ek, Danica Kragic
ICRA4
2014 Online contact point estimation for uncalibrated tool use
abstract
One of the big challenges for robots working outside of traditional industrial settings is the ability to robustly and flexibly grasp and manipulate tools for various tasks. When a tool is interacting with another object during task execution, several problems arise: a tool can be partially or completely occluded from the robot's view, it can slip or shift in the robot's hand - thus, the robot may lose the information about the exact position of the tool in the hand. Thus, there is a need for online calibration and/or recalibration of the tool. In this paper, we present a model-free online tool-tip calibration method that uses force/torque measurements and an adaptive estimation scheme to estimate the point of contact between a tool and the environment. An adaptive force control component guarantees that interaction forces are limited even before the contact point estimate has converged. We also show how to simultaneously estimate the location and normal direction of the surface being touched by the tool-tip as the contact point is estimated. The stability of the the overall scheme and the convergence of the estimated parameters are theoretically proven and the performance is evaluated in experiments on a real robot.
Yiannis Karayiannidis, Christian Smith, Francisco E. Vina, Danica Kragic
ICRA4
2014 ST-HMP: Unsupervised Spatio-Temporal feature learning for tactile data
abstract
Tactile sensing plays an important role in robot grasping and object recognition. In this work, we propose a new descriptor named Spatio-Temporal Hierarchical Matching Pursuit (ST-HMP) that captures properties of a time series of tactile sensor measurements. It is based on the concept of unsupervised hierarchical feature learning realized using sparse coding. The ST-HMP extracts rich spatio-temporal structures from raw tactile data without the need to predefine discriminative data characteristics. We apply it to two different applications: (1) grasp stability assessment and (2) object instance recognition, presenting its universal properties. An extensive evaluation on several synthetic and real datasets collected using the Schunk Dexterous, Schunk Parallel and iCub hands shows that our approach outperforms previously published results by a large margin.
Marianna Madry, Liefeng Bo, Danica Kragic, Dieter Fox
ICRA3
2014 Grasp moduli spaces and spherical harmonics
abstract
In this work, we present a novel representation which enables a robot to reason about, transfer and optimize grasps on various objects by representing objects and grasps on them jointly in a common space. In our approach, objects are parametrized using smooth differentiable functions which are obtained from point cloud data via a spectral analysis. We show how, starting with point cloud data of various objects, one can utilize this space consisting of grasps and smooth surfaces in order to continuously deform various surface/grasp configurations with the goal of synthesizing force closed grasps on novel objects. We illustrate the resulting shape space for a collection of real world objects using multidimensional scaling and show that our formulation naturally enables us to use gradient ascent approaches to optimize and simultaneously deform a grasp from a known object towards a novel object.
Florian T. Pokorny, Yasemin Bekiroglu, Danica Kragic
ICRA3
2014 What's in the container? Classifying object contents from vision and touch
abstract
Robots operating in household environments need to interact with food containers of different types. Whether a container is filled with milk, juice, yogurt or coffee may affect the way robots grasp and manipulate the container. In this paper, we concentrate on the problem of identifying what kind of content is in a container based on tactile and/or visual feedback in combination with grasping. In particular, we investigate the benefits of using unimodal (visual or tactile) or bimodal (visual-tactile) sensory data for this purpose. We direct our study toward cardboard containers with liquid or solid content or being empty. The motivation for using grasping rather than shaking is that we want to investigate the content prior to applying manipulation actions to a container. Our results show that we achieve comparable classification rates with unimodal data and that the visual and tactile data are complimentary.
Püren Güler, Yasemin Bekiroglu, Xavi Gratal, Karl Pauwels, Danica Kragic
IROS5
2014 Hierarchical Fingertip Space for multi-fingered precision grasping
abstract
Dexterous in-hand manipulation of objects benefits from the ability of a robot system to generate precision grasps. In this paper, we propose a concept of Fingertip Space and its use for precision grasp synthesis. Fingertip Space is a representation that takes into account both the local geometry of object surface as well as the fingertip geometry. As such, it is directly applicable to the object point cloud data and it establishes a basis for the grasp search space. We propose a model for a hierarchical encoding of the Fingertip Space that enables multilevel refinement for efficient grasp synthesis. The proposed method works at the grasp contact level while not neglecting object shape nor hand kinematics. Experimental evaluation is performed for the Barrett hand considering also noisy and incomplete point cloud data.
Kaiyu Hang, Johannes A. Stork, Danica Kragic
IROS3
2014 Learning of grasp adaptation through experience and tactile sensing
abstract
To perform robust grasping, a multi-fingered robotic hand should be able to adapt its grasping configuration, i.e., how the object is grasped, to maintain the stability of the grasp. Such a change of grasp configuration is called grasp adaptation and it depends on the controller, the employed sensory feedback and the type of uncertainties inherit to the problem. This paper proposes a grasp adaptation strategy to deal with uncertainties about physical properties of objects, such as the object weight and the friction at the contact points. Based on an object-level impedance controller, a grasp stability estimator is first learned in the object frame. Once a grasp is predicted to be unstable by the stability estimator, a grasp adaptation strategy is triggered according to the similarity between the new grasp and the training examples. Experimental results demonstrate that our method improves the grasping performance on novel objects with different physical properties from those used for training.
Miao Li 0002, Yasemin Bekiroglu, Danica Kragic, Aude Billard
IROS3
2014 Maximally satisfying LTL action planning
abstract
We focus on autonomous robot action planning problem from Linear Temporal Logic (LTL) specifications, where the action refers to a “simple” motion or manipulation task, such as “go from A to B” or “grasp a ball”. At the high-level planning layer, we propose an algorithm to synthesize a maximally satisfying discrete control strategy while taking into account that the robot's action executions may fail. Furthermore, we interface the high-level plan with the robot's low-level controller through a reactive middle-layer formalism called Behavior Trees (BTs). We demonstrate the proposed framework using a NAO robot capable of walking, ball grasping and ball dropping actions.
Jana Tumova, Alejandro Marzinotto, Dimos V. Dimarogonas, Danica Kragic
IROS4
2014 Mapping human intentions to robot motions via physical interaction through a jointly-held object
abstract
In this paper we consider the problem of human-robot collaborative manipulation of an object, where the human is active in controlling the motion, and the robot is passively following the human's lead. Assuming that the human grasp of the object only allows for transfer of forces and not torques, there is a disambiguity as to whether the human desires translation or rotation. In this paper, we analyze different approaches to this problem both theoretically and in experiment. This leads to the proposal of a control methodology that uses switching between two different admittance control modes based on the magnitude of measured force to achieve disambiguation of the rotation/translation problem.
Yiannis Karayiannidis, Christian Smith, Danica Kragic
RO-MAN3
2014 Detecting, segmenting and tracking unknown objects using multi-label MRF inference
Mårten Björkman, Niklas Bergström, Danica Kragic
Comput. Vis. Image Underst.3
2014 Data-Driven Grasp Synthesis - A Survey
abstract
We review the work on data-driven grasp synthesis and the methodologies for sampling and ranking candidate grasps. We divide the approaches into three groups based on whether they synthesize grasps for known, familiar, or unknown objects. This structure allows us to identify common object representations and perceptual processes that facilitate the employed data-driven grasp synthesis technique. In the case of known objects, we concentrate on the approaches that are based on object recognition and pose estimation. In the case of familiar objects, the techniques use some form of a similarity matching to a set of previously encountered objects. Finally, for the approaches dealing with unknown objects, the core part is the extraction of specific features that are indicative of good grasps. Our survey provides an overview of the different methodologies and discusses open problems in the area of robot grasping. We also draw a parallel to the classical approaches that rely on analytic formulations.
Jeannette Bohg, Antonio Morales, Tamim Asfour, Danica Kragic
IEEE Trans. Robotics4
2013 The Path Kernel
Andrea Baisero, Florian T. Pokorny, Danica Kragic, Carl Henrik Ek
ICPRAM3
2013 A probabilistic framework for task-oriented grasp stability assessment
abstract
We present a probabilistic framework for grasp modeling and stability assessment. The framework facilitates assessment of grasp success in a goal-oriented way, taking into account both geometric constraints for task affordances and stability requirements specific for a task. We integrate high-level task information introduced by a teacher in a supervised setting with low-level stability requirements acquired through a robot's self-exploration. The conditional relations between tasks and multiple sensory streams (vision, proprioception and tactile) are modeled using Bayesian networks. The generative modeling approach both allows prediction of grasp success, and provides insights into dependencies between variables and features relevant for object grasping.
Yasemin Bekiroglu, Dan Song 0002, Lu Wang 0006, Danica Kragic
ICRA4
2013 Learning a dictionary of prototypical grasp-predicting parts from grasping experience
abstract
We present a real-world robotic agent that is capable of transferring grasping strategies across objects that share similar parts. The agent transfers grasps across objects by identifying, from examples provided by a teacher, parts by which objects are often grasped in a similar fashion. It then uses these parts to identify grasping points onto novel objects. We focus our report on the definition of a similarity measure that reflects whether the shapes of two parts resemble each other, and whether their associated grasps are applied near one another. We present an experiment in which our agent extracts five prototypical parts from thirty-two real-world grasp examples, and we demonstrate the applicability of the prototypical parts for grasping novel objects.
Renaud Detry, Carl Henrik Ek, Marianna Madry, Danica Kragic
ICRA4
2013 Sparse summarization of robotic grasping data
abstract
We propose a new approach for learning a summarized representation of high dimensional continuous data. Our technique consists of a Bayesian non-parametric model capable of encoding high-dimensional data from complex distributions using a sparse summarization. Specifically, the method marries techniques from probabilistic dimensionality reduction and clustering. We apply the model to learn efficient representations of grasping data for two robotic scenarios.
Martin Hjelm, Carl Henrik Ek, Renaud Detry, Hedvig Kjellström, Danica Kragic
ICRA5
2013 Model-free robot manipulation of doors and drawers by means of fixed-grasps
abstract
This paper addresses the problem of robot interaction with objects attached to the environment through joints such as doors or drawers. We propose a methodology that requires no prior knowledge of the objects' kinematics, including the type of joint - either prismatic or revolute. The method consists of a velocity controller which relies on force/torque measurements and estimation of the motion direction, rotational axis and the distance from the center of rotation. The method is suitable for any velocity controlled manipulator with a force/torque sensor at the end-effector. The force/torque control regulates the applied forces and torques within given constraints, while the velocity controller ensures that the end-effector moves with a task-related desired tangential velocity. The paper also provides a proof that the estimates converge to the actual values. The method is evaluated in different scenarios typically met in a household environment.
Yiannis Karayiannidis, Christian Smith, Francisco E. Vina, Petter Ögren, Danica Kragic
ICRA5
2013 Language for learning complex human-object interactions
abstract
In this paper we use a Hierarchical Hidden Markov Model (HHMM) to represent and learn complex activities/task performed by humans/robots in everyday life. Action primitives are used as a grammar to represent complex human behaviour and learn the interactions and behaviour of human/robots with different objects. The main contribution is the use of a probabilistic model capable of representing behaviours at multiple levels of abstraction to support the proposed hypothesis. The hierarchical nature of the model allows decomposition of the complex task into simple action primitives. The framework is evaluated with data collected for tasks of everyday importance performed by a human user.
Carl Henrik Ek, Nikolaos Kyriazis, Antonis A. Argyros, Jaime Valls Miró, Danica Kragic
ICRA6
2013 Grasping objects with holes: A topological approach
abstract
This work proposes a topologically inspired approach for generating robot grasps on objects with `holes'. Starting from a noisy point-cloud, we generate a simplicial representation of an object of interest and use a recently developed method for approximating shortest homology generators to identify graspable loops. To control the movement of the robot hand, a topologically motivated coordinate system is used in order to wrap the hand around such loops. Finally, another concept from topology - namely the Gauss linking integral - is adapted to serve as evidence for secure caging grasps after a grasp has been executed. We evaluate our approach in simulation on a Barrett hand using several target objects of different sizes and shapes and present an initial experiment with real sensor data.
Florian T. Pokorny, Johannes A. Stork, Danica Kragic
ICRA3
2013 Predicting human intention in visual observations of hand/object interactions
abstract
The main contribution of this paper is a probabilistic method for predicting human manipulation intention from image sequences of human-object interaction. Predicting intention amounts to inferring the imminent manipulation task when human hand is observed to have stably grasped the object. Inference is performed by means of a probabilistic graphical model that encodes object grasping tasks over the 3D state of the observed scene. The 3D state is extracted from RGB-D image sequences by a novel vision-based, markerless hand-object 3D tracking framework. To deal with the high-dimensional state-space and mixed data types (discrete and continuous) involved in grasping tasks, we introduce a generative vector quantization method using mixture models and self-organizing maps. This yields a compact model for encoding of grasping actions, able of handling uncertain and partial sensory data. Experimentation showed that the model trained on simulated data can provide a potent basis for accurate goal-inference with partial and noisy observations of actual real-world demonstrations. We also show a grasp selection process, guided by the inferred human intention, to illustrate the use of the system for goal-directed grasp imitation.
Dan Song 0002, Nikolaos Kyriazis, Iasonas Oikonomidis, Chavdar Papazov, Antonis A. Argyros, Darius Burschka, Danica Kragic
ICRA7
2013 Enhancing visual perception of shape through tactile glances
abstract
Object shape information is an important parameter in robot grasping tasks. However, it may be difficult to obtain accurate models of novel objects due to incomplete and noisy sensory measurements. In addition, object shape may change due to frequent interaction with the object (cereal boxes, etc). In this paper, we present a probabilistic approach for learning object models based on visual and tactile perception through physical interaction with an object. Our robot explores unknown objects by touching them strategically at parts that are uncertain in terms of shape. The robot starts by using only visual features to form an initial hypothesis about the object shape, then gradually adds tactile measurements to refine the object model. Our experiments involve ten objects of varying shapes and sizes in a real setup. The results show that our method is capable of choosing a small number of touches to construct object models similar to real object shapes and to determine similarities among acquired models.
Mårten Björkman, Yasemin Bekiroglu, Virgile Hogman, Danica Kragic
IROS4
2013 Friction coefficients and grasp synthesis
abstract
We propose a new concept called friction sensitivity which measures how susceptible a specific grasp is to changes in the underlying friction coefficients. We develop algorithms for the synthesis of stable grasps with low friction sensitivity and for the synthesis of stable grasps in the case of small friction coefficients. We describe how grasps with low friction sensitivity can be used when a robot has an uncertain belief about friction coefficients and study the statistics of grasp quality under changes in those coefficients. We also provide a parametric estimate for the distribution of grasp qualities and friction sensitivities for a uniformly sampled set of grasps.
Kaiyu Hang, Florian T. Pokorny, Danica Kragic
IROS3
2013 Interactive object classification using sensorimotor contingencies
abstract
Understanding and representing objects and their function is a challenging task. Objects we manipulate in our daily activities can be described and categorized in various ways according to their properties or affordances, depending also on our perception of those. In this work, we are interested in representing the knowledge acquired through interaction with objects, describing these in terms of action-effect relations, i.e. sensorimotor contingencies, rather than static shape or appearance representations. We demonstrate how a robot learns sensorimotor contingencies through pushing using a probabilistic model. We show how functional categories can be discovered and how entropy-based action selection can improve object classification.
Virgile Hogman, Mårten Björkman, Danica Kragic
IROS3
2013 Online kinematics estimation for active human-robot manipulation of jointly held objects
abstract
This paper introduces a method for estimating the constraints imposed by a human agent on a jointly manipulated object. These estimates can be used to infer knowledge of where the human is grasping an object, enabling the robot to plan trajectories for manipulating the object while subject to the constraints. We describe the method in detail, motivate its validity theoretically, and demonstrate its use in co-manipulation tasks with a real robot.
Yiannis Karayiannidis, Christian Smith, Francisco E. Vina, Danica Kragic
IROS4
2013 Extracting essential local object characteristics for 3D object categorization
abstract
Most object classes share a considerable amount of local appearance and often only a small number of features are discriminative. The traditional approach to represent an object is based on a summarization of the local characteristics by counting the number of feature occurrences. In this paper we propose the use of a recently developed technique for summarizations that, rather than looking into the quantity of features, encodes their quality to learn a description of an object. Our approach is based on extracting and aggregating only the essential characteristics of an object class for a task. We show how the proposed method significantly improves on previous work in 3D object categorization. We discuss the benefits of the method in other scenarios such as robot grasping. We provide extensive quantitative and qualitative experiments comparing our approach to the state of the art to justify the described approach.
Marianna Madry, Heydar Maboudi Afkham, Carl Henrik Ek, Stefan Carlsson, Danica Kragic
IROS5
2013 Classical grasp quality evaluation: New algorithms and theory
abstract
This paper investigates theoretical properties of a well-known L1grasp quality measure Q whose approximation Q−lis commonly used for the evaluation of grasps and where the precision of Q−ldepends on an approximation of a cone by a convex polyhedral cone with l edges. We prove the Lipschitz continuity of Q and provide an explicit Lipschitz bound that can be used to infer the stability of grasps lying in a neighbourhood of a known grasp. We think of Q−las a lower bound estimate to Q and describe an algorithm for computing an upper bound Q+. We provide worst-case error bounds relating Q and Q−l. Furthermore, we develop a novel grasp hypothesis rejection algorithm which can exclude unstable grasps much faster than current implementations. Our algorithm is based on a formulation of the grasp quality evaluation problem as an optimization problem, and we show how our algorithm can be used to improve the efficiency of sampling based grasp hypotheses generation methods.
Florian T. Pokorny, Danica Kragic
IROS2
2013 Integrated motion and clasp planning with virtual linking
abstract
In this work, we address the problem of simultaneous clasp and motion planning on unknown objects with holes. Clasping an object enables a rich set of activities such as dragging, toting, pulling and hauling which can be applied to both soft and rigid objects. To this end, we define a virtual linking measure which characterizes the spacial relation between the robot hand and object. The measure utilizes a set of closed curves arising from an approximately shortest basis of the object's first homology group. We define task spaces to perform collision-free motion planing with respect to multiple prioritized objectives using a sampling-based planing method. The approach is tested in simulation using different robot hands and various real-world objects.
Johannes A. Stork, Florian T. Pokorny, Danica Kragic
IROS3
2013 Caging complex objects with geodesic balls
abstract
This paper proposes a novel approach for the synthesis of grasps of objects whose geometry can be observed only in the presence of noise. We focus in particular on the problem of generating caging grasps with a realistic robot hand simulation and show that our method can generate such grasps even on complex objects. We introduce the idea of using geodesic balls on the object's surface in order to approximate the maximal contact surface between a robotic hand and an object. We define two types of heuristics which extract information from approximate geodesic balls in order to identify areas on an object that can likely be used to generate a caging grasp. Our heuristics are based on two scoring functions. The first uses winding angles measuring how much a geodesic ball on the surface winds around a dominant axis, while the second explores using the total discrete Gaussian curvature of a geodesic ball to rank potential caging postures. We evaluate our approach with respect to variations in hand kinematics, for a selection of complex real-world objects and with respect to its robustness to noise.
Dmitry Zarubin, Florian T. Pokorny, Marc Toussaint, Danica Kragic
IROS4
2013 Non-parametric hand pose estimation with object context
Javier Romero 0002, Hedvig Kjellström, Carl Henrik Ek, Danica Kragic
Image Vis. Comput.4
2013 A Metric for Comparing the Anthropomorphic Motion Capability of Artificial Hands
abstract
We propose a metric for comparing the anthropomorphic motion capability of robotic and prosthetic hands. The metric is based on the evaluation of how many different postures or configurations a hand can perform by studying the reachable set of fingertip poses. To define a benchmark for comparison, we first generate data with human subjects based on an extensive grasp taxonomy. We then develop a methodology for comparison using generative, nonlinear dimensionality reduction techniques. We assess the performance of different hands with respect to the human hand and with respect to each other. The method can be used to compare other types of kinematic structures.
Thomas Feix, Javier Romero 0002, Carl Henrik Ek, Heinz-Bodo Schmiedmayer, Danica Kragic
IEEE Trans. Robotics5
2013 Extracting Postural Synergies for Robotic Grasping
abstract
We address the problem of representing and encoding human hand motion data using nonlinear dimensionality reduction methods. We build our work on the notion of postural synergies being typically based on a linear embedding of the data. In addition to addressing the encoding of postural synergies using nonlinear methods, we relate our work to control strategies of combined reaching and grasping movements. We show the drawbacks of the (commonly made) causality assumption and propose methods that model the data as being generated from an inferred latent manifold to cope with the problem. Another important contribution is a thorough analysis of the parameters used in the employed dimensionality reduction techniques. Finally, we provide an experimental evaluation that shows how the proposed methods outperform the standard techniques, both in terms of recognition and generation of motion patterns.
Javier Romero 0002, Thomas Feix, Carl Henrik Ek, Hedvig Kjellström, Danica Kragic
IEEE Trans. Robotics5
2012 Generalizing grasps across partly similar objects
abstract
The paper starts by reviewing the challenges associated to grasp planning, and previous work on robot grasping. Our review emphasizes the importance of agents that generalize grasping strategies across objects, and that are able to transfer these strategies to novel objects. In the rest of the paper, we then devise a novel approach to the grasp transfer problem, where generalization is achieved by learning, from a set of grasp examples, a dictionary of object parts by which objects are often grasped. We detail the application of dimensionality reduction and unsupervised clustering algorithms to the end of identifying the size and shape of parts that often predict the application of a grasp. The learned dictionary allows our agent to grasp novel objects which share a part with previously seen objects, by matching the learned parts to the current view of the new object, and selecting the grasp associated to the best-fitting part. We present and discuss a proof-of-concept experiment in which a dictionary is learned from a set of synthetic grasp examples. While prior work in this area focused primarily on shape analysis (parts identified, e.g., through visual clustering, or salient structure analysis), the key aspect of this work is the emergence of parts from both object shape and grasp examples. As a result, parts intrinsically encode the intention of executing a grasp.
Renaud Detry, Carl Henrik Ek, Marianna Madry, Justus H. Piater, Danica Kragic
ICRA5
2012 From object categories to grasp transfer using probabilistic reasoning
abstract
In this paper we address the problem of grasp generation and grasp transfer between objects using categorical knowledge. The system is built upon an i) active scene segmentation module, able of generating object hypotheses and segmenting them from the background in real time, ii) object categorization system using integration of 2D and 3D cues, and iii) probabilistic grasp reasoning system. Individual object hypotheses are first generated, categorized and then used as the input to a grasp generation and transfer system that encodes task, object and action properties. The experimental evaluation compares individual 2D and 3D categorization approaches with the integrated system, and it demonstrates the usefulness of the categorization in task-based grasping and grasp transfer.
Marianna Madry, Dan Song 0002, Danica Kragic
ICRA3
2012 Distributed cooperative object attitude manipulation
abstract
This paper proposes a local information based control law in order to solve the planar manipulation problem of rotating a grasped rigid object to a desired orientation using multiple mobile manipulators. We adopt a multi-agent systems theory approach and assume that: (i) the manipulators (agents) are capable of sensing the relative position to their neighbors at discrete time instances, (ii) neighboring agents may exchange information at discrete time instances, and (iii) the communication topology is connected. Control of the manipulators is carried out at a kinematic level in continuous time and utilizes inverse kinematics. The mobile platforms are assigned trajectory tracking tasks that adjust the positions of the manipulator bases in order to avoid singular arm configurations. Our main result concerns the stability of the proposed control law.
Johan Markdahl, Yiannis Karayiannidis, Xiaoming Hu 0001, Danica Kragic
ICRA4
2012 Improving generalization for 3D object categorization with Global Structure Histograms
abstract
We propose a new object descriptor for three dimensional data named the Global Structure Histogram (GSH). The GSH encodes the structure of a local feature response on a coarse global scale, providing a beneficial trade-off between generalization and discrimination. Encoding the structural characteristics of an object allows us to retain low local variations while keeping the benefit of global representativeness. In an extensive experimental evaluation, we applied the framework to category-based object classification in realistic scenarios. We show results obtained by combining the GSH with several different local shape representations, and we demonstrate significant improvements to other state-of-the-art global descriptors.
Marianna Madry, Carl Henrik Ek, Renaud Detry, Kaiyu Hang, Danica Kragic
IROS5
2012 YES - YEt another object segmentation: Exploiting camera movement
abstract
We address the problem of object segmentation in image sequences where no a-priori knowledge of objects is assumed. We take advantage of robots' ability to move, gathering multiple images of the scene. Our approach starts by extracting edges, uses a polar domain representation and performs integration over time based on a simple dilation operation. The proposed system can be used for providing reliable initial segmentation of unknown objects in scenes of varying complexity, allowing for recognition, categorization or physical interaction with the objects. The experimental evaluation on both self-captured and a publicly available dataset shows the efficiency and stability of the proposed method.
Lazaros Nalpantidis, Mårten Björkman, Danica Kragic
IROS3
2012 Learning and recognition of objects inspired by early cognition
abstract
In this paper, we present a unifying approach for learning and recognition of objects in unstructured environments through exploration. Taking inspiration from how young infants learn objects, we establish four principles for object learning. First, early object detection is based on an attention mechanism detecting salient parts in the scene. Second, motion of the object allows more accurate object localization. Next, acquiring multiple observations of the object through manipulation allows a more robust representation of the object. And last, object recognition benefits from a multi-modal representation. Using these principles, we developed a unifying method including visual attention, smooth pursuit of the object, and a multi-view and multi-modal object representation. Our results indicate the effectiveness of this approach and the improvement of the system when multiple observations are acquired from active object manipulation.
Maja Rudinac, Gert Kootstra, Danica Kragic, Pieter P. Jonker
IROS3
2012 Persistent Homology for Learning Densities with Bounded Support
abstract
We present a novel method for learning densities with bounded support which enables us to incorporate `hard' topological constraints. In particular, we show how emerging techniques from computational algebraic topology and the notion of Persistent Homology can be combined with kernel based methods from Machine Learning for the purpose of density estimation. The proposed formalism facilitates learning of models with bounded support in a principled way, and -- by incorporating Persistent Homology techniques in our approach -- we are able to encode algebraic-topological constraints which are not addressed in current state-of the art probabilistic models. We study the behaviour of our method on two synthetic examples for various sample sizes and exemplify the benefits of the proposed approach on a real-world data-set by learning a motion model for a racecar. We show how to learn a model which respects the underlying topological structure of the racetrack, constraining the trajectories of the car.
Florian T. Pokorny, Carl Henrik Ek, Hedvig Kjellström, Danica Kragic
NIPS4
2011 Integrating grasp planning with online stability assessment using tactile sensing
abstract
This paper presents an integration of grasp planning and online grasp stability assessment based on tactile data. We show how the uncertainty in grasp execution posterior to grasp planning can be dealt with using tactile sensing and machine learning techniques. The majority of the state-of-the art grasp planners demonstrate impressive results in simulation. However, these results are mostly based on perfect scene/object knowledge allowing for analytical measures to be employed. It is questionable how well these measures can be used in realistic scenarios where the information about the object and robot hand may be incomplete and/or uncertain. Thus, tactile and force-torque sensory information is necessary for successful online grasp stability assessment. We show how a grasp planner can be integrated with a probabilistic technique for grasp stability assessment in order to improve the hypotheses about suitable grasps on different types of objects. Experimental evaluation with a three-fingered robot hand equipped with tactile array sensors shows the feasibility and strength of the integrated approach.
Yasemin Bekiroglu, Kai Huebner, Danica Kragic
ICRA3
2011 Mind the gap - robotic grasping under incomplete observation
abstract
We consider the problem of grasp and manipulation planning when the state of the world is only partially observable. Specifically, we address the task of picking up unknown objects from a table top. The proposed approach to object shape prediction aims at closing the knowledge gaps in the robot's understanding of the world. A completed state estimate of the environment can then be provided to a simulator in which stable grasps and collision-free movements are planned. The proposed approach is based on the observation that many objects commonly in use in a service robotic scenario possess symmetries. We search for the optimal parameters of these symmetries given visibility constraints. Once found, the point cloud is completed and a surface mesh reconstructed. Quantitative experiments show that the predictions are valid approximations of the real object shape. By demonstrating the approach on two very different robotic platforms its generality is emphasized.
Jeannette Bohg, Matthew Johnson-Roberson, Beatriz León, Javier Felip, Xavi Gratal, Niklas Bergström, Danica Kragic, Antonio Morales
ICRA7
2011 Fast and bottom-up object detection, segmentation, and evaluation using Gestalt principles
abstract
In many scenarios, domestic robot will regularly encounter unknown objects. In such cases, top-down knowledge about the object for detection, recognition, and classification cannot be used. To learn about the object, or to be able to grasp it, bottom-up object segmentation is an important competence for the robot. Also when there is top-down knowledge, prior segmentation of the object can improve recognition and classification. In this paper, we focus on the problem of bottom-up detection and segmentation of unknown objects. Gestalt psychology studies the same phenomenon in human vision. We propose the utilization of a number of Gestalt principles. Our method starts by generating a set of hypotheses about the location of objects using symmetry. These hypotheses are then used to initialize the segmentation process. The main focus of the paper is on the evaluation of the resulting object segments using Gestalt principles to select segments with high figural goodness. The results show that the Gestalt principles can be successfully used for detection and segmentation of unknown objects. The results furthermore indicate that the Gestalt measures for the goodness of a segment correspond well with the objective quality of the segment. We exploit this to improve the overall segmentation performance.
Gert Kootstra, Danica Kragic
ICRA2
2011 Multivariate discretization for Bayesian Network structure learning in robot grasping
abstract
A major challenge in modeling with BNs is learning the structure from both discrete and multivariate continuous data. A common approach in such situations is to discretize continuous data before structure learning. However efficient methods to discretize high-dimensional variables are largely lacking. This paper presents a novel method specifically aiming at discretization of high-dimensional, high-correlated data. The method consists of two integrated steps: non-linear dimensionality reduction using sparse Gaussian process latent variable models, and discretization by application of a mixture model. The model is fully probabilistic and capable to facilitate structure learning from discretized data, while at the same time retain the continuous representation. We evaluate the effectiveness of the method in the domain of robot grasping. Compared with traditional discretization schemes, our model excels both in task classification and prediction of hand grasp configurations. Further, being a fully probabilistic model it handles uncertainty in the data and can easily be integrated into other frameworks in a principled manner.
Dan Song 0002, Carl Henrik Ek, Kai Huebner, Danica Kragic
ICRA4
2011 Scene Understanding through Autonomous Interactive Perception
Niklas Bergström, Carl Henrik Ek, Mårten Björkman, Danica Kragic
ICVS4
2011 Learning tactile characterizations of object- and pose-specific grasps
abstract
Our aim is to predict the stability of a grasp from the perceptions available to a robot before attempting to lift up and transport an object. The percepts we consider consist of the tactile imprints and the object-gripper configuration read before and until the robot's manipulator is fully closed around an object. Our robot is equipped with multiple tactile sensing arrays and it is able to track the pose of an object during the application of a grasp. We present a kernel-logistic-regression model of pose- and touch-conditional grasp success probability which we train on grasp data collected by letting the robot experience the effect on tactile and visual signals of grasps suggested by a teacher, and letting the robot verify which grasps can be used to rigidly control the object. We consider models defined on several subspaces of our input data - e.g., using tactile perceptions or pose information only. Our experiment demonstrates that joint tactile and pose-based perceptions carry valuable grasp-related information, as models trained on both hand poses and tactile parameters perform better than the models trained exclusively on one perceptual input.
Yasemin Bekiroglu, Renaud Detry, Danica Kragic
IROS3
2011 Generating object hypotheses in natural scenes through human-robot interaction
abstract
We propose a method for interactive modeling of objects and object relations based on real-time segmentation of video sequences. In interaction with a human, the robot can perform multi-object segmentation through principled modeling of physical constraints. The key contribution is an efficient multi-labeling framework, that allows object modeling and disambiguation in natural scenes. Object modeling and labeling is done in a real-time segmentation system, to which hypotheses and constraints denoting relations between objects can be added incrementally. Through instructions such as key presses or spoken words, a scene can be segmented in regions corresponding to multiple physical objects. The approach solves some of the difficult problems related to disambiguation of objects merged due to their direct physical contact. Results show that even a limited set of simple interactions with a human operator can substantially improve segmentation results.
Niklas Bergström, Mårten Björkman, Danica Kragic
IROS3
2011 Enhanced visual scene understanding through human-robot dialog
abstract
We propose a novel human-robot-interaction framework for robust visual scene understanding. Without any a-priori knowledge about the objects, the task of the robot is to correctly enumerate how many of them are in the scene and segment them from the background. Our approach builds on top of state-of-the-art computer vision methods, generating object hypotheses through segmentation. This process is combined with a natural dialog system, thus including a `human in the loop' where, by exploiting the natural conversation of an advanced dialog system, the robot gains knowledge about ambiguous situations. We present an entropy-based system allowing the robot to detect the poorest object hypotheses and query the user for arbitration. Based on the information obtained from the human-robot dialog, the scene segmentation can be re-seeded and thereby improved. We present experimental results on real data that show an improved segmentation performance compared to segmentation without interaction.
Matthew Johnson-Roberson, Jeannette Bohg, Gabriel Skantze, Joakim Gustafson, Rolf Carlson, Babak Rasolzadeh, Danica Kragic
IROS7
2011 Representing actions with Kernels
abstract
A long standing research goal is to create robots capable of interacting with humans in dynamic environments. To realise this a robot needs to understand and interpret the underlying meaning and intentions of a human action through a model of its sensory data. The visual domain provides a rich description of the environment and data is readily available in most system through inexpensive cameras. However, such data is very high-dimensional and extremely redundant making modeling challenging.
Guoliang Luo, Niklas Bergström, Carl Henrik Ek, Danica Kragic
IROS4
2011 Grasping unknown objects using an Early Cognitive Vision system for general scene understanding
abstract
For some time now machine learning methods have been widely used in perception for autonomous robots. While there have been many results describing the performance of machine learning techniques with regards to their accuracy or convergence rates, relatively little work has been done on developing theoretical performance guarantees about their stability and robustness. As a result, many machine learning techniques are still limited to being used in situations where safety and robustness are not critical for success. One way to overcome this difficulty is by using reachability analysis, which can be used to compute regions of the state space, known as reachable sets, from which the system can be guaranteed to remain safe over some time horizon regardless of the disturbances. In this paper we show how reachability analysis can be combined with machine learning in a scenario in which an aerial robot is attempting to learn the dynamics of a ground vehicle using a camera with a limited field of view. The resulting simulation data shows that by combining these two paradigms, one can create robotic systems that feature the best qualities of each, namely high performance and guaranteed safety.
Mila Popovic, Gert Kootstra, Jimmy A. Jørgensen, Danica Kragic, Norbert Krüger
IROS4
2011 Embodiment-specific representation of robot grasping using graphical models and latent-space discretization
abstract
We study embodiment-specific robot grasping tasks, represented in a probabilistic framework. The framework consists of a Bayesian network (BN) integrated with a novel multi-variate discretization model. The BN models the probabilistic relationships among tasks, objects, grasping actions and constraints. The discretization model provides compact data representation that allows efficient learning of the conditional structures in the BN. To evaluate the framework, we use a database generated in a simulated environment including examples of a human and a robot hand interacting with objects. The results show that the different kinematic structures of the hands affect both the BN structure and the conditional distributions over the modeled variables. Both models achieve accurate task classification, and successfully encode the semantic task requirements in the continuous observation spaces. In an imitation experiment, we demonstrate that the representation framework can transfer task knowledge between different embodiments, therefore is a suitable model for grasp planning and imitation in a goal-directed manner.
Dan Song 0002, Carl Henrik Ek, Kai Huebner, Danica Kragic
IROS4
2011 The Importance of Structure
Carl Henrik Ek, Danica Kragic
ISRR2
2011 Visual object-action recognition: Inferring object affordances from human demonstration
Hedvig Kjellström, Javier Romero 0002, Danica Kragic
Comput. Vis. Image Underst.3
2011 Tracking rigid objects using integration of model-based and model-free cues
Ville Kyrki, Danica Kragic
Mach. Vis. Appl.2
2011 Assessing Grasp Stability Based on Learning and Haptic Data
abstract
An important ability of a robot that interacts with the environment and manipulates objects is to deal with the uncertainty in sensory data. Sensory information is necessary to, for example, perform online assessment of grasp stability. We present methods to assess grasp stability based on haptic data and machine-learning methods, including AdaBoost, support vector machines (SVMs), and hidden Markov models (HMMs). In particular, we study the effect of different sensory streams to grasp stability. This includes object information such as shape; grasp information such as approach vector; tactile measurements from fingertips; and joint configuration of the hand. Sensory knowledge affects the success of the grasping process both in the planning stage (before a grasp is executed) and during the execution of the grasp (closed-loop online control). In this paper, we study both of these aspects. We propose a probabilistic learning framework to assess grasp stability and demonstrate that knowledge about grasp stability can be inferred using information from tactile sensors. Experiments on both simulated and real data are shown. The results indicate that the idea to exploit the learning approach is applicable in realistic scenarios, which opens a number of interesting venues for the future research.
Yasemin Bekiroglu, Janne Laaksonen, Jimmy A. Jørgensen, Ville Kyrki, Danica Kragic
IEEE Trans. Robotics5
2010 Active 3D Segmentation through Fixation of Previously Unseen Objects
abstract
We present an approach for active segmentation based on integration of several cues.It serves as a framework for generation of object hypotheses of previously unseen objectsin natural scenes. Using an approximate Expectation-Maximisation method, the appearance,3D shape and size of objects are modelled in an iterative manner, with fixation usedfor unsupervised initialisation. To better cope with situations where an object is hard tosegregate from the surface it is placed on, a flat surface model is added to the typical twohypotheses used in classical figure-ground segmentation. The framework is further extendedto include modelling over time, in order to exploit temporal consistency for bettersegmentation and to facilitate tracking.
Mårten Björkman, Danica Kragic
BMVC2
2010 Tracking people interacting with objects
abstract
While the problem of tracking 3D human motion has been widely studied, most approaches have assumed that the person is isolated and not interacting with the environment. Environmental constraints, however, can greatly constrain and simplify the tracking problem. The most studied constraints involve gravity and contact with the ground plane. We go further to consider interaction with objects in the environment. In many cases, tracking rigid environmental objects is simpler than tracking high-dimensional human motion. When a human is in contact with objects in the world, their poses constrain the pose of body, essentially removing degrees of freedom. Thus what would appear to be a harder problem, combining object and human tracking, is actually simpler. We use a standard formulation of the body tracking problem but add an explicit model of contact with objects. We find that constraints from the world make it possible to track complex articulated human motion in 3D from a monocular camera.
Hedvig Kjellström, Danica Kragic, Michael J. Black
CVPR2
2010 Using Symmetry to Select Fixation Points for Segmentation
abstract
For the interpretation of a visual scene, it is important for a robotic system to pay attention to the objects in the scene and segment them from their background. We focus on the segmentation of previously unseen objects in unknown scenes. The attention model therefore needs to be bottom-up and context-free. In this paper, we propose the use of symmetry, one of the Gestalt principles for figure-ground segregation, to guide the robot's attention. We show that our symmetry-saliency model outperforms the contrast-saliency model, proposed in. The symmetry model performs better in finding the objects of interest and selects a fixation point closer to the center of the object. Moreover, the objects are better segmented from the background when the initial points are selected on the basis of symmetry.
Gert Kootstra, Niklas Bergström, Danica Kragic
ICPR3
2010 Active 3D scene segmentation and detection of unknown objects
abstract
We present an active vision system for segmentation of visual scenes based on integration of several cues. The system serves as a visual front end for generation of object hypotheses for new, previously unseen objects in natural scenes. The system combines a set of foveal and peripheral cameras where, through a stereo based fixation process, object hypotheses are generated. In addition to considering the segmentation process in 3D, the main contribution of the paper is integration of different cues in a temporal framework and improvement of initial hypotheses over time.
Mårten Björkman, Danica Kragic
ICRA2
2010 Hands in action: real-time 3D reconstruction of hands in interaction with objects
abstract
This paper presents a method for vision based estimation of the pose of human hands in interaction with objects. Despite the fact that most robotics applications of human hand tracking involve grasping and manipulation of objects, the majority of methods in the literature assume a free hand, isolated from the surrounding environment. Our hand tracking method is non-parametric, performing a nearest neighbor search in a large database (100000 entries) of hand poses with and without grasped objects. The system operates in real time, it is robust to self occlusions, object occlusions and segmentation errors, and provides full hand pose reconstruction from markerless video. Temporal consistency in hand pose is taken into account, without explicitly tracking the hand in the high dimensional pose space.
Javier Romero 0002, Hedvig Kjellström, Danica Kragic
ICRA3
2010 Strategies for multi-modal scene exploration
abstract
We propose a method for multi-modal scene exploration where initial object hypothesis formed by active visual segmentation are confirmed and augmented through haptic exploration with a robotic arm. We update the current belief about the state of the map with the detection results and predict yet unknown parts of the map with a Gaussian Process. We show that through the integration of different sensor modalities, we achieve a more complete scene model. We also show that the prediction of the scene structure leads to a valid scene representation even if the map is not fully traversed. Furthermore, we propose different exploration strategies and evaluate them both in simulation and on our robotic platform.
Jeannette Bohg, Matthew Johnson-Roberson, Mårten Björkman, Danica Kragic
IROS4
2010 Attention-based active 3D point cloud segmentation
abstract
In this paper we present a framework for the segmentation of multiple objects from a 3D point cloud. We extend traditional image segmentation techniques into a full 3D representation. The proposed technique relies on a state-of-the-art min-cut framework to perform a fully 3D global multi-class labeling in a principled manner. Thereby, we extend our previous work in which a single object was actively segmented from the background. We also examine several seeding methods to bootstrap the graphical model-based energy minimization and these methods are compared over challenging scenes. All results are generated on real-world data gathered with an active vision robotic head. We present quantitive results over aggregate sets as well as visual results on specific examples.
Matthew Johnson-Roberson, Jeannette Bohg, Mårten Björkman, Danica Kragic
IROS4
2010 Representations for object grasping and learning from experience
abstract
We study two important problems in the area of robot grasping: i) the methodology and representations for grasp selection on known and unknown objects, and ii) learning from experience for grasping of similar objects. The core part of the paper is the study of different representations necessary for implementing grasping tasks on objects of different complexity. We show how to select a grasp satisfying force-closure, taking into account the parameters of the robot hand and collision-free paths. Our implementation takes also into account efficient computation at different levels of the system regarding representation, description and grasp hypotheses generation.
Óscar Jesús Rubio Martí, Kai Huebner, Danica Kragic
IROS3
2010 Spatio-temporal modeling of grasping actions
abstract
Understanding the spatial dimensionality and temporal context of human hand actions can provide representations for programming grasping actions in robots and inspire design of new robotic and prosthetic hands. The natural representation of human hand motion has high dimensionality. For specific activities such as handling and grasping of objects, the commonly observed hand motions lie on a lower-dimensional non-linear manifold in hand posture space. Although full body human motion is well studied within Computer Vision and Biomechanics, there is very little work on the analysis of hand motion with nonlinear dimensionality reduction techniques. In this paper we use Gaussian Process Latent Variable Models (GPLVMs) to model the lower dimensional manifold of human hand motions during object grasping. We show how the technique can be used to embed high-dimensional grasping actions in a lower-dimensional space suitable for modeling, recognition and mapping.
Javier Romero 0002, Thomas Feix, Hedvig Kjellström, Danica Kragic
IROS4
2010 Learning task constraints for robot grasping using graphical models
abstract
This paper studies the learning of task constraints that allow grasp generation in a goal-directed manner. We show how an object representation and a grasp generated on it can be integrated with the task requirements. The scientific problems tackled are (i) identification and modeling of such task constraints, and (ii) integration between a semantically expressed goal of a task and quantitative constraint functions defined in the continuous object-action domains. We first define constraint functions given a set of object and action attributes, and then model the relationships between object, action, constraint features and the task using Bayesian networks. The probabilistic framework deals with uncertainty, combines a-priori knowledge with observed data, and allows inference on target attributes given only partial observations. We present a system designed to structure data generation and constraint learning processes that is applicable to new tasks, embodiments and sensory data. The application of the task constraint model is demonstrated in a goal-directed imitation experiment.
Dan Song 0002, Kai Huebner, Ville Kyrki, Danica Kragic
IROS4
2010 Learning grasp stability based on tactile data and HMMs
abstract
In this paper, the problem of learning grasp stability in robotic object grasping based on tactile measurements is studied. Although grasp stability modeling and estimation has been studied for a long time, there are few robots today able of demonstrating extensive grasping skills. The main contribution of the work presented here is an investigation of probabilistic modeling for inferring grasp stability based on learning from examples. The main objective is classification of a grasp as stable or unstable before applying further actions on it, e.g. lifting. The problem cannot be solved by visual sensing which is typically used to execute an initial robot hand positioning with respect to the object. The output of the classification system can trigger a regrasping step if an unstable grasp is identified. An off-line learning process is implemented and used for reasoning about grasp stability for a three-fingered robotic hand using Hidden Markov models. To evaluate the proposed method, experiments are performed both in simulation and on a real robot system.
Yasemin Bekiroglu, Danica Kragic, Ville Kyrki
RO-MAN2
2009 Integration of Visual Cues for Robotic Grasping
Niklas Bergström, Jeannette Bohg, Danica Kragic
ICVS3
2008 Simultaneous Visual Recognition of Manipulation Actions and Manipulated Objects
Hedvig Kjellström, Javier Romero 0002, David Martínez Mercado, Danica Kragic
ECCV (2)4
2008 Minimum volume bounding box decomposition for shape approximation in robot grasping
abstract
Thinking about intelligent robots involves consideration of how such systems can be enabled to perceive, interpret and act in arbitrary and dynamic environments. While sensor perception and model interpretation focus on the robot's internal representation of the world rather passively, robot grasping capabilities are needed to actively execute tasks, modify scenarios and thereby reach versatile goals. These capabilities should also include the generation of stable grasps to safely handle even objects unknown to the robot. We believe that the key to this ability is not to select a good grasp depending on the identification of an object (e.g. as a cup), but on its shape (e.g. as a composition of shape primitives). In this paper, we envelop given 3D data points into primitive box shapes by a fit-and-split algorithm that is based on an efficient Minimum Volume Bounding Box implementation. Though box shapes are not able to approximate arbitrary data in a precise manner, they give efficient clues for planning grasps on arbitrary objects. We present the algorithm and experiments using the 3D grasping simulator Grasplt!.
Kai Huebner, Steffen Ruthotto, Danica Kragic
ICRA3
2008 Modeling and recognition of actions through motor primitives
abstract
We investigate modeling and recognition of object manipulation actions for the purpose of imitation based learning in robotics. To model the process, we are using a combination of discriminative (support vector machines, conditional random fields) and generative approaches (hidden Markov models). We examine the hypothesis that complex actions can be represented as a sequence of motion or action primitives. The experimental evaluation, performed with five object manipulation actions and 10 people, investigates the modeling approach of the primitive action structure and compares the performance of the considered generative and discriminative models.
David Martínez Mercado, Danica Kragic
ICRA2
2008 Dynamic time warping for binocular hand tracking and reconstruction
abstract
We show how matching and reconstruction of contour points can be performed using dynamic time warping (DTW) for the purpose of 3D hand contour tracking. We evaluate the performance of the proposed algorithm in object manipulation activities and perform comparison with the iterative closest point (ICP) method.
Javier Romero 0002, Danica Kragic, Ville Kyrki, Antonis A. Argyros
ICRA2
2008 Integration of Visual and Shape Attributes for Object Action Complexes
Kai Huebner, Mårten Björkman, Babak Rasolzadeh, Martina Schmidt, Danica Kragic
ICVS5
2008 Selection of robot pre-grasps using box-based shape approximation
abstract
Grasping is a central issue of various robot applications, especially when unknown objects have to be manipulated by the system. In earlier work, we have shown the efficiency of 3D object shape approximation by box primitives for the purpose of grasping. A point cloud was approximated by box primitives [1]. In this paper, we present a continuation of these ideas and focus on the box representation itself. On the number of grasp hypotheses from box face normals, we apply heuristic selection integrating task, orientation and shape issues. Finally, an off-line trained neural network is applied to chose a final best hypothesis as the final grasp. We motivate how boxes as one of the simplest representations can be applied in a more sophisticated manner to generate task-dependent grasps.
Kai Huebner, Danica Kragic
IROS2
2008 Visual recognition of grasps for human-to-robot mapping
abstract
This paper presents a vision based method for grasp classification. It is developed as part of a Programming by Demonstration (PbD) system for which recognition of objects and pick-and-place actions represent basic building blocks for task learning. In contrary to earlier approaches, no articulated 3D reconstruction of the hand over time is taking place. The indata consists of a single image of the human hand. A 2D representation of the hand shape, based on gradient orientation histograms, is extracted from the image. The hand shape is then classified as one of six grasps by finding similar hand shapes in a large database of grasp images. The database search is performed using Locality Sensitive Hashing (LSH), an approximate k-nearest neighbor approach. The nearest neighbors also give an estimated hand orientation with respect to the camera. The six human grasps are mapped to three Barret hand grasps. Depending on the type of robot grasp, a precomputed grasp strategy is selected. The strategy is further parameterized by the orientation of the hand relative to the object. To evaluate the potential for the method to be part of a robust vision system, experiments were performed, comparing classification results to a baseline of human classification performance. The experiments showed the LSH recognition performance to be comparable to human performance.
Hedvig Kjellström, Javier Romero 0002, Danica Kragic
IROS3
2007 Learning and Evaluation of the Approach Vector for Automatic Grasp Generation and Planning
abstract
In this paper, we address the problem of automatic grasp generation for robotic hands where experience and shape primitives are used in synergy so to provide a basis not only for grasp generation but also for a grasp evaluation process when the exact pose of the object is not available. One of the main challenges in automatic grasping is the choice of the object approach vector, which is dependent both on the object shape and pose as well as the grasp type. Using the proposed method, the approach vector is chosen not only based on the sensory input but also on experience that some approach vectors will provide useful tactile information that finally results in stable grasps. A methodology for developing and evaluating grasp controllers is presented where the focus lies on obtaining stable grasps under imperfect vision. The method is used in a teleoperation or a programming by demonstration setting where a human demonstrates to a robot how to grasp an object. The system first recognizes the object and grasp type which can then be used by the robot to perform the same action using a mapped version of the human grasping posture.
Staffan Ekvall, Danica Kragic
ICRA2
2007 Contour reconstruction using recursive smoothing splines - experimental validation
abstract
In this paper, a recursive smoothing spline approach for contour reconstruction is studied and evaluated. Periodic smoothing splines are used by a robot to approximate the contour of encountered obstacles in the environment. The splines are generated through minimizing a cost function subject to constraints imposed by a linear control system and accuracy is improved iteratively using a recursive spline algorithm. The filtering effect of the smoothing splines allows for usage of noisy sensor data and the method is robust to odometry drift. Experimental evaluation is performed for contour reconstruction of three objects using a SICK laser scanner mounted on a PowerBot from ActivMedia Robotics.
Giacomo Piccolo, Maja Karasalo, Danica Kragic, Xiaoming Hu 0001
IROS3
2007 Action Recognition and Understanding using Motor Primitives
abstract
We investigate modeling and recognition of arm manipulation actions of different levels of complexity. To model the process, we are using a combination of discriminative support vector machines and generative hidden Markov models. The experimental evaluation, performed with 10 people, investigates both definition and structure of primitive motions as well as the validity of the modeling approach taken.
Ville Kyrki, Isabel Serrano Vicente, Danica Kragic, Jan-Olof Eklundh
RO-MAN3
2007 Learning and Recognition of Object Manipulation Actions Using Linear and Nonlinear Dimensionality Reduction
abstract
In this work, we perform an extensive statistical evaluation for learning and recognition of object manipulation actions. We concentrate on single arm/hand actions but study the problem of modeling and dimensionality reduction for cases where actions are very similar to each other in terms of arm motions. For this purpose, we evaluate a linear and a nonlinear dimensionality reduction techniques: principal component analysis and spatio-temporal isomap. Classification of query sequences is based on different variants of Nearest Neighbor classification. We thoroughly describe and evaluate different parameters that affect the modeling strategies and perform the evaluation with a training set of 20 people.
Isabel Serrano Vicente, Danica Kragic, Jan-Olof Eklundh
RO-MAN2
2006 A Framework for Vision Based bearing only 3D SLAM
abstract
This paper presents a framework for 3D vision based bearing only SLAM using a single camera, an interesting setup for many real applications due to its low cost. The focus in is on the management of the features to achieve real-time performance in extraction, matching and loop detection. For matching image features to map landmarks a modified, rotationally variant SIFT descriptor is used in combination with a Harris-Laplace detector. To reduce the complexity in the map estimation while maintaining matching performance only a few, high quality, image features are used for map landmarks. The rest of the features are used for matching. The framework has been combined with an EKF implementation for SLAM. Experiments performed in indoor environments are presented. These experiments demonstrate the validity and effectiveness of the approach. In particular they show how the robot is able to successfully match current image features to the map when revisiting an area
Patric Jensfelt, Danica Kragic, John Folkesson, Mårten Björkman
ICRA2
2006 Tracking Unobservable Rotations by Cue Integration
abstract
Model based object tracking has earned significant importance in areas such as augmented reality, surveillance, visual servoing, robotic object manipulation and grasping. Although an active research area, there are still few systems that perform robustly in realistic settings. The key problems to robust and precise object tracking are outliers caused by occlusion, self-occlusion, cluttered background, and reflections. Two most common solutions to the above problems have been the use of robust estimators and the integration of visual cues. The tracking system considered in this paper achieves robustness by integrating model-based and model-free cues. As model-based cues, we consider a CAD model of the object known a priori and as model-free cues, automatically generated corner features are used. The main idea is to account for relative object motion between consecutive frames using integration of the two cues. The particular contribution of this work is the integration framework where not only polyhedral objects are considered. In particular, we deal with spherical, cylindrical and conical objects for which the complete pose cannot be estimate using only CAD like models. Using the integration with the model-free features, we show how a full pose estimate can be obtained. Experimental evaluation demonstrates robust system performance in realistic settings with highly textured objects
Ville Kyrki, Danica Kragic
ICRA2
2006 Nonholonomic Epipolar Visual Servoing
abstract
A significant amount of work has been reported in the area of visual servoing during the last decade. However, most of the contributions are applied in cases of holonomic robots. More recently, the use of visual feedback for control of nonholonomic vehicles has been reported. Some of the examples are docking and parallel parking maneuvers of cars or vision-based stabilization of a mobile manipulator to a desired pose with respect to a target of interest. Still, many of the approaches are mostly interested in the control part of visual servoing loop considering very simple vision algorithms based on artificial markers. In this paper, we present an approach for nonholonomic visual servoing based on epipolar geometry. The method facilitates a classical teach-by-showing approach where a reference image is used to define the desired pose (position and orientation) of the robot. The major contribution of the paper is the design of the control law that considers nonholonomic constraints of the robot as well as the robust feature detection and matching process based on scale and rotation invariant image features. An extensive experimental evaluation has been performed in a realistic indoor setting and the results are summarized in the paper
Gonzalo López-Nicolás, Carlos Sagüés, Josechu J. Guerrero, Danica Kragic, Patric Jensfelt
ICRA4
2006 Robust Statistics for 3D Object Tracking
abstract
This paper focuses on methods that enhance performance of a model based 3D object tracking system. Three statistical methods and an improved edge detector are discussed and compared. The evaluation is performed on a number of characteristic sequences incorporating shift, rotation, texture, weak illumination and occlusion. Considering the deviations of the pose parameters from ground truth, it is shown that improving the measurements' accuracy in the detection step yields better results than improving contaminated measurements with statistical means
Peter Preisig, Danica Kragic
ICRA2
2006 Strategies for Object Manipulation using Foveal and Peripheral Vision
abstract
Visual feedback is used extensively in robotics and application areas range from human-robot interaction to object grasping and manipulation. There have been a number of examples of how to develop different components required by the above applications and very few general vision systems capable of performing a variety of tasks. In this paper, we concentrate on vision strategies for robotic manipulation tasks in a domestic environment. In particular, given fetch-and-carry type of tasks, the issues related to the whole detect-approach-grasp loop are considered. We deal with the problem of flexibility and robustness by using monocular and binocular visual cues and their integration. We demonstrate real-time disparity estimation, object recognition and pose estimation. We also show how a combination of foveal and peripheral vision system can be combined in order to provide a wide, low resolution and narrow, high resolution field of view.
Danica Kragic, Mårten Björkman
ICVS1
2006 Layered HMM for Motion Intention Recognition
abstract
Acquiring, representing and modeling human skills is one of the key research areas in teleoperation, programming-by-demonstration and human-machine collaborative settings. One of the common approaches is to divide the task that the operator is executing into several subtask in order to provide manageable modeling. In this paper we consider the use of a layered hidden Markov model (LHMM) to model human skills. We evaluate a gestem classifier that classifies motions into basic action-primitives, or gestems. The gestem classifiers are then used in a LHMM to model a simulated teleoperated task. We investigate the online and offline classification performance with respect to noise, number of gestems, type of HMM and the available number of training sequences. We also apply the LHMM to data recorded during the execution of a trajectory-tracking task in 2D and 3D with a robotic manipulator in order to give qualitative as well as quantitative results for the proposed approach. The results indicate that the LHMM is suitable for modeling teleoperative trajectory-tracking tasks and that the difference in classification performance between one and multi-dimensional HMMs for gestem classification are small. It can also be seen that the LHMM is robust w.r.t misclassifications in the underlying gestem classifiers
Daniel Aarno, Danica Kragic
IROS2
2006 Integrating Active Mobile Robot Object Recognition and SLAM in Natural Environments
abstract
Linking semantic and spatial information has become an important research area in robotics since, for robots interacting with humans and performing tasks in natural environments, it is of foremost importance to be able to reason beyond simple geometrical and spatial levels. In this paper, we consider this problem in a service robot scenario where a mobile robot autonomously navigates in a domestic environment, builds a map as it moves along, localizes its position in it, recognizes objects on its way and puts them in the map. The experimental evaluation is performed in a realistic setting where the main concentration is put on the synergy of object recognition and simultaneous localization and mapping systems
Staffan Ekvall, Patric Jensfelt, Danica Kragic
IROS3
2006 Integration of Tracking and Adaptive Gaussian Mixture Models for Posture Recognition
abstract
In this paper, we present a system for continuous posture recognition. The main contributions of the proposed approach are the integration of an adaptive color model with a tracking system that allows for robust continuous posture recognition based on principal component analysis. The adaptive color model uses Gaussian mixture models for skin and background color representation, Bayesian framework for classification and Kalman filter for tracking hands and head of a person that interacts with the robot. Experimental evaluation shows that the integration of tracking and an adaptive color model supports the robustness and flexibility of the system when illumination changes occur
Jesus Ignacio Bueno, Danica Kragic
RO-MAN2
2006 Task Learning Using Graphical Programming and Human Demonstrations
abstract
The next generation of robots will have to learn new tasks or refine the existing ones through direct interaction with the environment or through a teaching/coaching process in programming by demonstration (PbD) and learning by instruction frameworks. In this paper, we propose to extend the classical PbD approach with a graphical language that makes robot coaching easier. The main idea is based on graphical programming where the user designs complex robot tasks by using a set of low-level action primitives. Different to other systems, our action primitives are made general and flexible so that the user can train them online and therefore easily design high level tasks
Staffan Ekvall, Daniel Aarno, Danica Kragic
RO-MAN3
2006 Learning Task Models from Multiple Human Demonstrations
abstract
In this paper, we present a novel method for learning robot tasks from multiple demonstrations. Each demonstrated task is decomposed into subtasks that allow for segmentation and classification of the input data. The demonstrated tasks are then merged into a flexible task model, describing the task goal and its constraints. The two main contributions of the paper are the state generation and contraints identification methods. We also present a task level planner, that is used to assemble a task plan at run-time, allowing the robot to choose the best strategy depending on the current world state
Staffan Ekvall, Danica Kragic
RO-MAN2
2006 Augmenting SLAM with Object Detection in a Service Robot Framework
abstract
In a service robot scenario, we are interested in a task of building maps of the environment that include automatically recognized objects. Most systems for simultaneous localization and mapping (SLAM) build maps that are only used for localizing the robot. Such maps are typically based on grids or different types of features such as point and lines. Here, we augment the process with an object recognition system that detects objects in the environment and puts them in the map generated by the SLAM system. During task execution, the robot can use this information to reason about objects, places and their relationships. The metric map is also split into topological entities corresponding to rooms. In this way, the user can command the robot to retrieve an object from a particular room or get help from a robot when searching for a certain object
Patric Jensfelt, Staffan Ekvall, Danica Kragic, Daniel Aarno
RO-MAN3
2006 Online task recognition and real-time adaptive assistance for computer-aided machine control
abstract
Segmentation and recognition of operator-generated motions are commonly facilitated to provide appropriate assistance during task execution in teleoperative and human-machine collaborative settings. The assistance is usually provided in a virtual fixture framework where the level of compliance can be altered online, thus improving the performance in terms of execution time and overall precision. However, the fixtures are typically inflexible, resulting in a degraded performance in cases of unexpected obstacles or incorrect fixture models. In this paper, we present a method for online task tracking and propose the use of adaptive virtual fixtures that can cope with the above problems. Here, rather than executing a predefined plan, the operator has the ability to avoid unforeseen obstacles and deviate from the model. To allow this, the probability of following a certain trajectory (subtask) is estimated and used to automatically adjusts the compliance, thus providing the online decision of how to fixture the movement
Staffan Ekvall, Daniel Aarno, Danica Kragic
IEEE Trans. Robotics3
2005 Adaptive Virtual Fixtures for Machine-Assisted Teleoperation Tasks
abstract
It has been demonstrated in a number of robotic areas how the use of virtual fixtures improves task performance both in terms of execution time and overall precision, [1]. However, the fixtures are typically inflexible, resulting in a degraded performance in cases of unexpected obstacles or incorrect fixture models. In this paper, we propose the use of adaptive virtual fixtures that enable us to cope with the above problems. A teleoperative or human machine collaborative setting is assumed with the core idea of dividing the task, that the operator is executing, into several subtasks. The operator may remain in each of these subtasks as long as necessary and switch freely between them. Hence, rather than executing a predefined plan, the operator has the ability to avoid unforeseen obstacles and deviate from the model. In our system, the probability that the user is following a certain trajectory (subtask) is estimated and used to automatically adjusts the compliance. Thus, an on-line decision of how to fixture the movement is provided.
Daniel Aarno, Staffan Ekvall, Danica Kragic
ICRA3
2005 Adaptive Virtual Fixtures for Machine-Assisted Teleoperation Tasks
abstract
It has been demonstrated in a number of robotic areas how the use of virtual fixtures improves task performance both in terms of execution time and overall precision, [1]. However, the fixtures are typically inflexible, resulting in a degraded performance in cases of unexpected obstacles or incorrect fixture models. In this paper, we propose the use of adaptive virtual fixtures that enable us to cope with the above problems. A teleoperative or human machine collaborative setting is assumed with the core idea of dividing the task, that the operator is executing, into several subtasks. The operator may remain in each of these subtasks as long as necessary and switch freely between them. Hence, rather than executing a predefined plan, the operator has the ability to avoid unforeseen obstacles and deviate from the model. In our system, the probability that the user is following a certain trajectory (subtask) is estimated and used to automatically adjusts the compliance. Thus, an on-line decision of how to fixture the movement is provided.
Daniel Aarno, Staffan Ekvall, Danica Kragic
ICRA3
2005 Robust Real-Time Visual Tracking: Comparison, Theoretical Analysis and Performance Evaluation
abstract
In this paper, two real-time pose tracking algorithms for rigid objects are compared. Both methods are 3D-model based and are capable of calculating the pose between the camera and an object with a monocular vision system. Here, special consideration has been put into defining and evaluating different performance criteria such as computational efficiency, accuracy and robustness. Both methods are described and a unifying framework is derived. The main advantage of both algorithms lie in their real-time capabilities (on standard hardware) whilst being robust to miss-tracking, occlusion and changes in illumination.
Andrew I. Comport, Danica Kragic, Éric Marchand, François Chaumette
ICRA2
2005 Grasp Recognition for Programming by Demonstration
abstract
The demand for flexible and re-programmable robots has increased the need for programming by demonstration systems. In this paper, grasp recognition is considered in a programming by demonstration framework. Three methods for grasp recognition are presented and evaluated. The first method uses Hidden Markov Models to model the hand posture sequence during the grasp sequence, while the second method relies on the hand trajectory and hand rotation. The third method is a hybrid method, in which both the first two methods are active in parallel. The particular contribution is that all methods rely on the grasp sequence and not just the final posture of the hand. This facilitates grasp recognition before the grasp is completed. Also, by analyzing the entire sequence and not just the final grasp, the decision is based on more information and increased robustness of the overall system is achieved. The experimental results show that both arm trajectory and final hand posture provide important information for grasp classification. By combining them, the recognition rate of the overall system is increased.
Staffan Ekvall, Danica Kragic
ICRA2
2005 Integration of Model-based and Model-free Cues for Visual Object Tracking in 3D
abstract
Vision is one of the most powerful sensory modalities in robotics, allowing operation in dynamic envi ronments. One of our long-term research interests is mobile manipulation, where precise location of the target object is commonly required during task execution. Recently, a number of approaches have been proposed for real-time 3D tracking and most of them utilize an edge (wireframe) model of the target. However, the use of an edge model has significant problems in complex scenes due to occlusions and multiple responses, especially in terms of initialization. In this paper, we propose a new tracking method based on integration of model-based cues with automatically generated model-free cues, in order to improve tracking accuracy and to avoid weaknesses of edge based tracking. The integration is performed in a Kalman filter framework that operates in real-time. Experimental evaluation shows that the inclusion of model-free cues offers superior performance.
Ville Kyrki, Danica Kragic
ICRA2
2005 Receptive field cooccurrence histograms for object detection
abstract
Object recognition is one of the major research topics in the field of computer vision. In robotics, there is often a need for a system that can locate certain objects in the environment - the capability which we denote as 'object detection'. In this paper, we present a new method for object detection. The method is especially suitable for detecting objects in natural scenes, as it is able to cope with problems such as complex background, varying illumination and object occlusion. The proposed method uses the receptive field representation where each pixel in the image is represented by a combination of its color and response to different filters. Thus, the cooccurrence of certain filter responses within a specific radius in the image serves as information basis for building the representation of the object. The specific goal in this paper is the development of an online learning scheme that is effective after just one training example but still has the ability to improve its performance with more time and new examples. We describe the details behind the algorithm and demonstrate its strength with an extensive experimental evaluation.
Staffan Ekvall, Danica Kragic
IROS2
2005 Object recognition and pose estimation using color cooccurrence histograms and geometric modeling
Staffan Ekvall, Danica Kragic, Frank Hoffmann 0001
Image Vis. Comput.2
2004 Artificial Potential Biased Probabilistic Roadmap Method
abstract
Probabilistic roadmap methods (PRM) have been successfully used to solve difficult path planning problems but their efficiency is limited when the free space contains narrow passages through which the robot must pass. This paper presents a new sampling scheme that aims to increase the probability of finding paths through narrow passages. Here, a biased sampling scheme is used to increase the distribution of nodes in narrow regions of the free space. A partial computation of the artificial potential field is used to bias the distribution of nodes.
Daniel Aarno, Danica Kragic, Henrik I. Christensen
ICRA2
2004 Combination of Foveal and Peripheral Vision for Object Recognition and Pose Estimation
abstract
In this paper, we present a real-time vision system that integrates a number of algorithms using monocular and binocular cues to achieve robustness in realistic settings, for tasks such as object recognition, tracking and pose estimation. The system consists of two sets of binocular cameras; a peripheral set for disparity based attention and a foveal one for higher level processes. Thus the conflicting requirements of a wide field of view and high resolution can be overcome. One important property of the system is that the step from task specification through object recognition to pose estimation is completely automatic, combining both appearance and geometric models. Experimental evaluation is performed in a realistic indoor environment with occlusions, clutter, changing lighting and background conditions.
Mårten Björkman, Danica Kragic
ICRA2
2004 Interactive Grasp Learning based on Human Demonstration
abstract
We describe our effort in development of an artificial cognitive system, able of performing complex manipulation tasks in a teleoperated or collaborative manner. Some of the work is motivated by human control strategies that, in general, involve comparison between sensory feedback and a-priori known, internal models. According to recent neuroscientific findings, predictions help to reduce the delays in obtaining the sensory information and to perform more complex tasks. This paper deals with the issue of robotic manipulation and grasping in particular. Two main contributions of the paper are: i) evaluation, recognition and modeling of human grasps during the arm transportation sequence, and ii) learning and representation of grasp strategies for different robotic hands.
Staffan Ekvall, Danica Kragic
ICRA2
2004 Measurement Errors in Visual Servoing
abstract
In recent years, a number of hybrid visual servoing control algorithms have been proposed and evaluated. For some time now, it has been clear that classical control approaches-image and position based-have some inherent problems. Hybrid approaches try to combine them to overcome these problems. However, most of the proposed approaches concentrate on the design of the control law, neglecting the issue of errors resulting from the sensory system. This paper addresses the issue of measurement errors in visual servoing. The particular contribution is the analysis of the propagation of image error through pose estimation and visual servoing control law. We have chosen to investigate the properties of the vision system and their effect to the performance of the control system. Two approaches are evaluated: i) position, and ii) 2 1/2 D visual servoing. We believe that our evaluation offers a tool to build and analyze hybrid control systems based on, for example, switching or partitioning.
Ville Kyrki, Danica Kragic, Henrik I. Christensen
ICRA2
2004 An Interactive Interface for Service Robots
abstract
In this paper, we present an initial design of an interactive interface for a service robot based on multisensor fusion. We show how the integration of speech, vision and laser range data can be performed using a high level of abstraction. Guided by a number of scenarios commonly used in a service robot framework, the experimental evaluation will show the benefit of sensory integration which allows the design of a robust and natural interaction system using a set of simple perceptual algorithms.
Elin Anna Topp, Danica Kragic, Patric Jensfelt, Henrik I. Christensen
ICRA2
2004 New shortest-path approaches to visual servoing
abstract
In recent years, a number of visual servo control algorithms have been proposed. Most approaches try to solve the inherent problems of image-based and position based servoing by partitioning the control between image and Cartesian spaces. However, partitioning of the control often causes the Cartesian path to become more complex, which might result in operation close to the joint limits. A solution to avoid the joint limits is to use a shortest-path approach, which avoids the limits in most cases. In this paper, two new shortest-path approaches to visual servoing are presented. First, a position-based approach is proposed that guarantees both shortest Cartesian trajectory and object visibility. Then, a variant is presented, which avoids the use of a 3D model of the target object by using homography based partial pose estimation.
Ville Kyrki, Danica Kragic, Henrik I. Christensen
IROS2
2003 Confluence of parameters in model based tracking
abstract
During the last decade, model based tracking of objects and its necessity in visual servoing and manipulation has been advocated in a number of systems. Most of these systems demonstrate robust performance for cases where either the background or the object are relatively uniform in color. In terms of manipulation, our basic interest is handling of everyday objects in domestic environments such as a home or an office. In this paper, we consider a number of different parameters that effect the performance of a model-based tracking system. Parameters such as color channels, feature detection, validation gates, outliers rejection and feature selection are considered here and their affect to the overall system performance is discussed. Experimental evaluation shows how some of these parameters can successfully be evaluated (learned) on-line and consequently improve the performance of the system.
Danica Kragic, Henrik I. Christensen
ICRA1
2003 Vision and tactile sensing for real world tasks
abstract
Robotic fetch-and-carry tasks are commonly facilitated to demonstrate a number of research directions such as navigation, mobile manipulation, systems integration, etc. As a part of an integrated system in terms of a service robot framework, this paper describes a set of methods for real-world object manipulation tasks. We concentrate here on two particular parts of a manipulation sequence: i) robust visual servoing, and ii) grasping strategies. In terms of visual servoing we discuss the handling of singularities during a manipulation sequence. For grasping, we present a biologically motivated strategy using tactile feedback.
Danica Kragic, S. Crinier, Dietrich Brunn, Henrik I. Christensen
ICRA1
2003 A Framework for Visual Servoing
Danica Kragic, Henrik I. Christensen
ICVS1
2003 Object recognition and pose estimation for robotic manipulation using color cooccurrence histograms
abstract
Robust techniques for object recognition, image segmentation and pose elimination are essential for robotic manipulation and grasping. We present a novel approach for object recognition and pose estimation based on color cooccurrence histograms (CCHs). Consequently, two problems addressed in this paper are: i) robust recognition and segmentation of the object in the scene, and ii) object's pose estimation using an appearance based approach. The proposed recognition scheme is based on the CCHs used in a classical learning framework that facilitates a "winner-takes-all" strategy across different scales. The detected "window of attention" is compared with training images of the object for which the pose is known. The orientation of the object is estimated as the weighted average among competitive poses, in which the weight increases proportional to the degree of matching between the training and the segmented image histograms. The major advantages of the proposed two-step appearance based method are its robustness and invariance towards scaling and translations. The method is also computationally efficient since both recognition and pose estimation rely on the same representation of the object.
Staffan Ekvall, Frank Hoffmann 0001, Danica Kragic
IROS3
2003 Biologically motivated visual servoing and grasping for real world tasks
abstract
Hand-eye coordination involves four tasks: i) identification of the object to be manipulated, ii) ballistic arm motion to the vicinity of the object, iii) preshaping and alignment of the hand, and finally iv) manipulation or grasping of the object. Motivated by the operation of biological systems and utilizing some constraints for each of the above mentioned tasks, we are aiming at design of a robust, robotic hand-eye coordination system. Hand-eye coordination tasks we consider here are of the basic fetch-and-carry type useful for service robots operating in everyday environments. Objects to be manipulated are, for example, food items that are simple in shape (polyhedral, cylindrical) but with complex surface texture. To achieve the required robustness and flexibility, we integrate both geometric and appearance based information to solve the task at hand. We show how the research in human visuo-motor system can be facilitated to design a fully operational, visually guided object manipulation system.
Danica Kragic, Henrik I. Christensen
IROS1
2003 Task modeling and specification for modular sensory based human-machine cooperative systems
abstract
This paper is directed towards developing human-machine cooperative systems (HCMS) for augmented surgical manipulation tasks. These tasks are commonly repetitive, sequential, and consist of simple steps. The transitions between these steps can be driven either by the surgeon's input or sensory information. Consequently, complex tasks can be effectively modeled using a set of basic primitives, where each primitive defines some basic type of motion (e.g. translational motion along a line, rotation about an axis, etc.). These steps can be "open-loop" (simply complying to user's demands) or "closed-loop, in which case external sensing is used to define a nominal reference trajectory. The particular research problem considered here is the development of a system that supports simple design of complex surgical procedures from a set of basic control primitives. The three system levels considered are: i) task graph generation which allows the user to easily design or model a task, ii) task graph execution which executes the task graph, and iii) at the lowest level, the specification of primitives which allows the user to easily specify new types of primitive motions. The system has been developed and validated using the JHU Steady Hand Robot as an experimental platform.
Danica Kragic, Gregory D. Hager
IROS1
2003 Erratum: Human-Machine Collaborative Systems for Microsurgical Applications
Danica Kragic, Panadda Marayong, Ming Li 0052, Allison M. Okamura, Gregory D. Hager
ISRR1
2002 Weak Models and Cue Integration for Real-Time Tracking
abstract
Traditionally, fusion of visual information for tracking has been based on explicit models for uncertainty and integration. Most of the approaches use some form of Bayesian statistics where strong models are employed. We argue that for cases where a large number of visual features are available, weak models for integration may be employed. We analyze integration by voting where two methods are proposed and evaluated: (i) response and (ii) action fusion. The methods differ in the choice of voting space: the former integrates visual information in image space and latter in velocity space. We also evaluate four weighting techniques for integration.
Danica Kragic, Henrik I. Christensen
ICRA1
2002 Systems Integration for Real-World Manipulation Tasks
abstract
A system developed to demonstrate integration of a number of key research areas such as localization, recognition, visual tracking, visual servoing and grasping is presented together with the underlying methodology adopted to facilitate the integration. Through sequencing of basic skills, provided by the above mentioned competencies, the system has the potential to carry out flexible grasping for fetch and carry in realistic environments. Through careful fusion of reactive and deliberative control and use of multiple sensory modalities a significant flexibility is achieved. Experimental verification of the integrated system is presented.
Lars Petersson, Patric Jensfelt, Dennis Tell, M. Strandberg, Danica Kragic, Henrik I. Christensen
ICRA5
2002 Model based techniques for robotic servoing and grasping
abstract
A robotic manipulation of objects typically involves object detection/recognition, servoing to the object, alignment and grasping. To perform fine alignment and final grasping, it is usually necessary to estimate the position and orientation (pose) of the object. In this paper we present a model based tracking system used to estimate and continuously update the pose of the object to be manipulated. Here, a wire-frame model is used to find and track features in the consequent images. One of the important parts of the system is the ability to automatically initiate the tracking process. The strength of the system is the ability to operate in an domestic environment (living room) with changing lighting and background conditions.
Danica Kragic, Henrik I. Christensen
IROS1
2001 Real-time Tracking Meets Online Grasp Planning
abstract
Describes a synergistic integration of a grasping simulator and a real-time visual tracking system, that work in concert to (1) find an object's pose, (2) plan grasps and movement trajectories, and (3) visually monitor task execution. Starting with a CAD model of an object to be grasped, the system can find the object's pose through vision which then synchronizes the state of the robot workcell with an online, model-based grasp planning and visualization system we have developed called GraspIt. GraspIt can then plan a stable grasp for the object, and direct the robotic hand system to perform the grasp. It can also generate trajectories for the movement of the grasped object, which are used by the visual control system to monitor the task and compare the actual grasp and trajectory with the planned ones. We present experimental results using typical grasping tasks.
Danica Kragic, Andrew T. Miller, Peter K. Allen
ICRA1
2001 Cue integration for visual servoing
abstract
The robustness and reliability of vision algorithms is, nowadays, the key issue in robotic research and industrial applications. To control a robot in a closed-loop fashion, different tracking systems have been reported in the literature. A common approach to increased robustness of a tracking system is the use of different models (CAD model of the object, motion model) known a priori. Our hypothesis is that fusion of multiple features facilitates robust detection and tracking of objects in scenes of realistic complexity. A particular application is the estimation of a robot's end-effector position in a sequence of images. The research investigates the following two different approaches to cue integration: 1) voting and 2) fuzzy logic-based fusion. The two approaches have been tested in association with scenes of varying complexity. Experimental results clearly demonstrate that fusion of cues results in a tracking system with a robust performance. The robustness is in particular evident for scenes with multiple moving objects and partial occlusion of the tracked object.
Danica Kragic, Henrik I. Christensen
IEEE Trans. Robotics Autom.1
2000 Tracking Techniques for Visual Servoing Tasks
abstract
Many of today's visual servoing systems rely on the use of markers on the object to provide features for control. There is thus a need for a visual system that provides control features regardless of the appearance of the object. Region based tracking is a natural approach since it does not require any special type of features. In this paper we present two different approaches to region based tracking: 1) a multi-resolution gradient based approach (using optical flow); and 2) a discrete feature based search approach. We present experiments conducted with both techniques for different types of image motions. Finally, the performance, drawbacks and limitations of used techniques are discussed.
Danica Kragic, Henrik I. Christensen
ICRA1
2000 High-level control of a mobile manipulator for door opening
abstract
In this paper, off-the-shelf algorithms for force/torque control are used in the context of mobile manipulation, in particular, the task of opening a door is studied. To make the solution robust, as few assumptions as possible are made. By using relaxation of forces as the basic level of control more complex information can be derived from the resulting motion. In our system, the radius and centre of rotation of the door are estimated online. This enables the complete system to have a higher degree of autonomy in an unknown environment. In addition, the redundancy of the robot is exploited in such a way to drive the system towards a desired configuration. The framework of hybrid dynamic systems is used to implement the algorithm which gives a theoretically sound framework for analysing the system with respect to safety and functionality. The integration of the above approaches results in a system which can robustly locate and grasp the handle and then open the door.
Lars Petersson, David J. Austin, Danica Kragic
IROS3
1999 A Person Following Behaviour for a Mobile Robot
abstract
In this paper, a person following behaviour for a mobile robot is presented. The head of the person is located using skin colour detection. Then, a control loop is fed with the camera movements required to put the upper part of the person in the center of the image. The algorithm was tested in different rooms of a research lab. It performed well in all lighting conditions except in direct sunlight. Since the background and lighting cannot be controlled, the vision algorithm must be robust to such changes. However, since the computing power is quite limited, the algorithm must have as low complexity as possible.
Hedvig Kjellström, Danica Kragic, Henrik I. Christensen
ICRA2
1999 Integration of visual cues for active tracking of an end-effector
abstract
We describe and test how information from multiple sources can be combined into a robust visual servoing system. The main objective is integration of visual cues to provide smooth pursuit in a cluttered environment using a minimum or no calibration. For that purpose, voting schema and fuzzy logic command fusion are investigated. It is shown that the integration permits detection and rejection of measurement outliers.
Danica Kragic, Henrik I. Christensen
IROS1