Guillem Alenyà

dblp:50/377 · DBLP profile ↗
← Back
65ranked-venue papers
6as first author
23since 2021 · last 2026
0000-0002-6018-154XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 53 · 5 first-author · 18 since 2021Systems, architecture and hardware · 22 · 4 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 18 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Ontological foundations for contrastive explanatory narration of robot plans
abstract
Mutual understanding of artificial agents' decisions is key to ensuring a trustworthy and successful human-robot interaction. Hence, robots are expected to make reasonable decisions and communicate them to humans when needed. In this article, the focus is on an approach to modeling and reasoning about the comparison of two competing plans, so that robots can later explain the divergent result. First, a novel ontological model is proposed to formalize and reason about the differences between competing plans, enabling the classification of the most appropriate one (e.g., the shortest, the safest, the closest to human preferences, etc.). This work also investigates the limitations of a baseline algorithm for ontology-based explanatory narration. To address these limitations, a novel algorithm is presented, leveraging divergent knowledge between plans and facilitating the construction of contrastive narratives. Through empirical evaluation, it is observed that the explanations excel beyond the baseline method.
Alberto Olivares Alarcos, Sergi Foix, Júlia Borràs Sol, Gerard Canal, Guillem Alenyà
Inf. Sci.5
2026 Impact of Design Transparency on Trust and Data Sharing during Human-Robot Interactions in Public Places
abstract
The prevalence of social robots is increasing, with examples such as customer service robots in malls and airports. This trend highlights the importance of transparency, particularly in data-sharing interactions with social robots operating in public spaces, where users may be asked to provide personal information to receive personalized experiences. This article investigates how design transparency influences user trust and data-sharing behavior in human-robot interactions. We conducted an experiment with 143 participants who interacted with the social robot ARI under two transparency conditions: low and high transparency. In the low-transparency condition, participants were informed about the data being collected and could choose to save or delete it. In the high-transparency condition, the robot additionally indicated the sensitivity level of each data item: low (e.g., scenario preference), medium (e.g., name and e-mail), and high (e.g., religious beliefs), allowing participants to make more informed decisions. Participants were presented with two scenarios: exploring city events and discovering local attractions. They received personalized recommendations based on their preferences, with the option to provide personal data (name, phone number, e-mail) for possible future communication. After the interaction, participants decided whether to save or delete the data they had shared. The results indicated that while transparency did not significantly affect trust in the robot, it influenced data-sharing behavior. In particular, participants in the high-transparency condition demonstrated more cautious behavior, opting to save less data and delete more. Furthermore, the results showed that both sensitivity levels and transparency influenced the participants’ data-sharing choices. Low-level sensitivity data led to the highest rates of saving and the lowest rates of deleting, while medium-level sensitivity data showed the opposite pattern. These findings highlight the need to align data categorization with user perceptions to address data sharing concerns more effectively.
Azra Aryania, Sabarathinam Chockalingam, Hanne Kristine Rødsethol, Guillem Alenyà
ACM Trans. Hum. Robot Interact.4
2024 Planning for Human-Robot Collaboration Scenarios with Heterogeneous Costs and Durations
abstract
This paper looks at human-robot collaboration (HRC) scenarios, in particular where the durations and costs of the actions are heterogeneous between agents, reflecting the agents’ capabilities as well as environmental constraints. We explore the use of temporal PDDL planning as a means of finding over-arching task plans for such collaborative scenarios, and apply suitable heuristics and search algorithms to improve the extent to which plans can be found that are sensitive to combined duration and cost metrics. An evaluation in a kitchen scenario shows our approach is effective, finding cost-effective task plans compared to those from existing planners, and a hand-crafted baseline.
Silvia Izquierdo-Badiola, Gerard Canal, Guillem Alenyà, Carlos Rizzo, Andrew Coles
ECAI3
2024 Standardization of Cloth Objects and its Relevance in Robotic Manipulation
abstract
The field of robotics faces inherent challenges in manipulating deformable objects, particularly in understanding and standardising fabric properties like elasticity, stiffness, and friction. While the significance of these properties is evident in the realm of cloth manipulation, accurately categorising and comprehending them in real-world applications remains elusive. This study sets out to address two primary objectives: (1) to provide a framework suitable for robotics applications to characterise cloth objects, and (2) to study how these properties influence robotic manipulation tasks. Our preliminary results validate the framework’s ability to characterise cloth properties and compare cloth sets, and reveal the influence that different properties have on the outcome of five manipulation primitives. We believe that, in general, results on the manipulation of clothes should be reported along with a better description of the garments used in the evaluation. This paper proposes a set of these measures.
Irene Garcia-Camacho, Alberta Longhini, Michael C. Welle, Guillem Alenyà, Danica Kragic, Júlia Borràs Sol
ICRA4
2024 PlanCollabNL: Leveraging Large Language Models for Adaptive Plan Generation in Human-Robot Collaboration
abstract
"Hey, robot. Let’s tidy up the kitchen. By the way, I have back pain today". How can a robotic system devise a shared plan with an appropriate task allocation from this abstract goal and agent condition? Classical AI task planning has been explored for this purpose, but it involves a tedious definition of an inflexible planning problem. Large Language Models (LLMs) have shown promising generalisation capabilities in robotics decision-making through knowledge extraction from Natural Language (NL). However, the translation of NL information into constrained robotics domains remains a challenge. In this paper, we use LLMs as translators between NL information and a structured AI task planning problem, targeting human-robot collaborative plans. The LLM generates information that is encoded in the planning problem, including specific subgoals derived from an NL abstract goal, as well as recommendations for subgoal allocation based on NL agent conditions. The framework, PlanCollabNL, is evaluated for a number of goals and agent conditions, and the results show that correct and executable plans are found in most cases. With this framework, we intend to add flexibility and generalisation to HRC plan generation, eliminating the need for a manual and laborious definition of restricted planning problems and agent models.
Silvia Izquierdo-Badiola, Gerard Canal, Carlos Rizzo, Guillem Alenyà
ICRA4
2024 How do people intend to disclose personal information to a social robot in public spaces?
abstract
Social robots interacting with people in public spaces may access and collect their personal information, which raises privacy concerns regarding the disclosure of personal information. This paper aims to investigate factors impacting individuals’ intention to disclose personal information to a social robot in public spaces and evaluate the actual disclosure during the interaction with the robot. For this purpose, a model is proposed to predict people’s intentions to disclose information to a social robot. We conducted our experiment at a public festival with more than 100 participants using the social robot ARI. The findings reveal the substantial impact of factors including risk beliefs, trusting beliefs, perceived enjoyment, and social influence on the intention to disclose personal information. Moreover, they reveal that although only a small percentage (6.20%) of people had the intention to disclose information to the social robot, most participants (98.00%) finally disclosed their personal information.
Azra Aryania, Ruben Huertas-Garcia, Santiago Forgas-Coll, Cecilio Angulo, Guillem Alenyà
RO-MAN5
2024 Exploring the Potential of a Robot-Assisted Frailty Assessment System for Elderly Care
abstract
Frailty assessment plays a pivotal role in providing older adults care. However, the current process is time-consuming and only measures patients’ completion time for each test. This paper introduces a set of algorithms to be used in robots to autonomously perform frailty assessments. In doing so we aim at reducing therapists’ burden and provide additional frailty-related metrics that can enhance the effectiveness of diagnosis. We conducted a pilot study with 22 elderly participants and compared our system’s performance with that of medical professionals to assess its precision. The results demonstrate that our approach achieved performances close to that of its human counterpart. This research represents an important step forward in the integration of social robotics in healthcare, offering potential benefits for patient care and clinical decision-making.
Aniol Civit, Antonio Andriella, Maite Antonio, Casimiro Javierre, Concepción Boqué, Guillem Alenyà
RO-MAN6
2024 What Would I Do If...? Promoting Understanding in HRI through Real-Time Explanations in the Wild
abstract
As robots become more and more integrated in human spaces, it is increasingly important for them to be able to explain their decisions to the people they interact with. These explanations need to be generated automatically and in real-time in response to decisions taken in dynamic and often unstructured environments. However, most research in explainable human-robot interaction only considers explanations (often manually selected) presented in controlled environments. We present an explanation generation method based on counterfactuals and demonstrate its use in an "in-the-wild" experiment using automatically generated and selected explanations of autonomous interactions with real people to assess the effect of these explanations on participants’ ability to predict the robot’s behaviour in hypothetical scenarios. Our results suggest that explanations aid one’s ability to predict the robot’s behaviour, but also that the addition of counterfactual statements may add some burden and counteract this beneficial effect.
Tamlin Love, Antonio Andriella, Guillem Alenyà
RO-MAN3
2024 LayerNet: High-Resolution Semantic 3D Reconstruction of Clothed People
abstract
In this article, we introduce SMPLicit, a novel generative model to jointly represent body pose, shape and clothing geometry; and LayerNet, a deep network that given a single image of a person simultaneously performs detailed 3D reconstruction of body and clothes. In contrast to existing learning-based approaches that require training specific models for each type of garment, SMPLicit can represent in a unified manner different garment topologies (e.g. from sleeveless tops to hoodies and open jackets), while controlling other properties like garment size or tightness/looseness. LayerNet follows a coarse-to-fine multi-stage strategy by first predicting smooth cloth geometries from SMPLicit, which are then refined by an image-guided displacement network that gracefully fits the body recovering high-frequency details and wrinkles. LayerNet achieves competitive accuracy in the task of 3D reconstruction against current 'garment-agnostic' state of the art for images of people in up-right positions and controlled environments, and consistently surpasses these methods on challenging body poses and uncontrolled settings. Furthermore, the semantically rich outcome of our approach is suitable for performing Virtual Try-on tasks directly on 3D, a task which, so far, has only been addressed in the 2D domain.
Enric Corona, Guillem Alenyà, Gerard Pons-Moll, Francesc Moreno-Noguer
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Robot explanatory narratives of collaborative and adaptive experiences
abstract
In the future, robots are expected to autonomously interact and/or collaborate with humans, who will increase the uncertainty during the execution of tasks, provoking online adaptations of robots' plans. Hence, trustworthy robots must be able to store, retrieve and narrate important knowledge about their collaborations and adaptations. In this article, it is proposed a sound methodology that integrates three main elements. First, an ontology for collaborative robotics and adaptation to model the domain knowledge. Second, an episodic memory for time-indexed knowledge storage and retrieval. Third, a novel algorithm to extract the relevant knowledge and generate textual explanatory narratives. The algorithm produces three different types of outputs, varying the specificity, for diverse uses and preferences. A pilot study was conducted to evaluate the usefulness of the narratives, obtaining promising results. Finally, we discuss how the methodology can be generalized to other ontologies and experiences. This work boosts robot explainability, especially in cases where robots need to narrate the details of their short and long-term past experiences.
Alberto Olivares Alarcos, Antonio Andriella, Sergi Foix, Guillem Alenyà
ICRA4
2023 User Interactions and Negative Examples to Improve the Learning of Semantic Rules in a Cognitive Exercise Scenario
abstract
Enabling a robot to perform new tasks is a complex endeavor, usually beyond the reach of non-technical users. For this reason, research efforts that aim at empowering end-users to teach robots new abilities using intuitive modes of interaction are valuable. In this article, we present INtuitive PROgramming 2 (INPRO2), a learning framework that allows inferring planning actions from demonstrations given by a human teacher. INPRO2 operates in an assistive scenario, in which the robot may learn from a healthcare professional (a therapist or caregiver) new cognitive exercises that can be later administered to patients with cognitive impairment. INPRO2 features significant improvements over previous work, namely: (1) exploitation of negative examples; (2) proactive interaction with the teacher to ask questions about the legality of certain movements; and (3) learning goals in addition to legal actions. Through simulations, we show the performance of different proactive strategies for gathering negative examples. Real-world experiments with human teachers and a TIAGo robot are also presented to qualitatively illustrate INPRO2.
Alejandro Suárez-Hernández, Antonio Andriella, Carme Torras, Guillem Alenyà
IROS4
2023 Teaching a Robot Where Doors and Drawers Are and How To Handle Them
abstract
We address the problem of teaching a service robot to detect doors and drawers in indoor environments. We propose a robust and accurate method in which a human demonstrates to the robot how to open doors and drawers that the robot is expected to operate in its future use. The proposed algorithm creates a model of a door or drawer from a sequence of RGB-D images and inserts it into an environment map. The model contains information about the size of the door panel or drawer front, as well as the position and orientation of the joint axis. This augmented environment map is then used by the robot to detect the target object in its environment and estimate its state.
Robert Cupec, Ivan Vidovic, Valentin Simundic, Petra Pejic, Sergi Foix, Guillem Alenyà
RO-MAN6
2023 Adaptive Human-Robot Collaboration: Evolutionary Learning of Action Costs Using an Action Outcome Simulator
abstract
One of the main challenges for successful human-robot collaborative applications lies in adapting the plan to the human agent’s changing state and preferences. A promising solution is to bridge the gap between agent modelling and AI task planning, which can be done by integrating the agent state as action costs in the task planning domain. This allows for the plan to be adapted to different partners, by influencing the action allocation. The difficulty then lies in setting appropriate action costs. This paper presents a novel framework to learn a set of planning action costs considering the preferred actions for an agent based on their state. An evolutionary optimisation algorithm is used for this purpose, and an action outcome simulator is developed to act as the black-box function, based on both an agent model and an action type model. This addresses the challenge of collecting data in HRC real-world scenarios, accelerating the learning for posterior fine-tuning in real applications. The coherence of the models and the simulator is proven through a conducted survey, and the learning algorithm is shown to learn appropriate action costs, producing plans that satisfy both the agents’ preferences and the prioritised plan requisites. The resulting system is a generic learning framework integrating components that can be easily extended to a wide range of applications, models and planning formalisms.
Silvia Izquierdo-Badiola, Guillem Alenyà, Carlos Rizzo
RO-MAN2
2023 Introducing CARESSER: A framework for in situ learning robot social assistance from expert knowledge and demonstrations
abstract
Abstract Socially assistive robots have the potential to augment and enhance therapist’s effectiveness in repetitive tasks such as cognitive therapies. However, their contribution has generally been limited as domain experts have not been fully involved in the entire pipeline of the design process as well as in the automatisation of the robots’ behaviour. In this article, we present aCtive leARning agEnt aSsiStive bEhaviouR (CARESSER), a novel framework that actively learns robotic assistive behaviour by leveraging the therapist’s expertise (knowledge-driven approach) and their demonstrations (data-driven approach). By exploiting that hybrid approach, the presented method enables in situ fast learning, in a fully autonomous fashion, of personalised patient-specific policies. With the purpose of evaluating our framework, we conducted two user studies in a daily care centre in which older adults affected by mild dementia and mild cognitive impairment ( N = 22) were requested to solve cognitive exercises with the support of a therapist and later on of a robot endowed with CARESSER. Results showed that: (i) the robot managed to keep the patients’ performance stable during the sessions even more so than the therapist; (ii) the assistance offered by the robot during the sessions eventually matched the therapist’s preferences. We conclude that CARESSER, with its stakeholder-centric design, can pave the way to new AI approaches that learn by leveraging human–human interactions along with human expertise, which has the benefits of speeding up the learning process, eliminating the need for the design of complex reward functions, and finally avoiding undesired states.
Antonio Andriella, Carme Torras, Carla Abdelnour, Guillem Alenyà
User Model. User Adapt. Interact.4
2023 Generating predicate suggestions based on the space of plans: an example of planning with preferences
abstract
Abstract Task planning in human–robot environments tends to be particularly complex as it involves additional uncertainty introduced by the human user. Several plans, entailing few or various differences, can be obtained to solve the same given task. To choose among them, the usual least-cost plan criteria is not necessarily the best option, because here, human constraints and preferences come into play. Knowing these user preferences is very valuable to select an appropriate plan, but the preference values are usually hard to obtain. In this context, we propose the Space-of-Plans-based Suggestions (SoPS) algorithms that can provide suggestions for some planning predicates, which are used to define the state of the environment in a task planning problem where actions modify the predicates. We denote these predicates as suggestible predicates, of which user preferences are a particular case. The first algorithm is able to analyze the potential effect of the unknown predicates and provide suggestions to values for these unknown predicates that may produce better plans. The second algorithm is able to suggest changes to already known values that potentially improve the obtained reward. The proposed approach utilizes a Space of Plans Tree structure to represent a subset of the space of plans. The tree is traversed to find the predicates and the values that would most increase the reward, and output them as a suggestion to the user. Our evaluation in three preference-based assistive robotics domains shows how the proposed algorithms can improve task performance by suggesting the most effective predicate values first.
Gerard Canal, Carme Torras, Guillem Alenyà
User Model. User Adapt. Interact.3
2022 Learned Vertex Descent: A New Direction for 3D Human Model Fitting
Enric Corona, Gerard Pons-Moll, Guillem Alenyà, Francesc Moreno-Noguer
ECCV (2)3
2022 Improved Task Planning through Failure Anticipation in Human-Robot Collaboration
abstract
Human-Robot Collaboration (HRC) has become a major trend in robotics in recent years with the idea of combining the strengths from both humans and robots. In order to share the work to be done, many task planning approaches have been implemented. However, they don't fully satisfy the required adaptability in human-robot collaborative tasks, with most approaches not considering neither the state of the human partner nor the possibility of adapting the collaborative plan during execution or even anticipating failures. In this paper, we present a planning system for human-robot collaborative plans that takes into account the agents' states and deals with unforeseen human behaviour, by replanning in anticipation when the human state changes to prevent action failure. The human state is defined in terms of capacity, knowledge and motivation. The system has been implemented in a standardised environment using the Planning Domain Definition Language (PDDL) and the modular ROSPlan framework, and we have validated the approach in multiple simulation settings. Our results show that using the human model fosters an appropriate task allocation while allowing failure anticipation, replanning in time to prevent it.
Silvia Izquierdo-Badiola, Gerard Canal, Carlos Rizzo, Guillem Alenyà
ICRA4
2022 Evaluating the Effect of Theory of Mind on People's Trust in a Faulty Robot
abstract
The success of human-robot interaction is strongly affected by the people’s ability to infer others’ intentions and behaviours, and the level of people’s trust that others will abide by their same principles and social conventions to achieve a common goal. The ability of understanding and reasoning about other agents’ mental states is known as Theory of Mind (ToM). ToM and trust, therefore, are key factors in the positive outcome of human-robot interaction. We believe that a robot endowed with a ToM is able to gain people’s trust, even when this may occasionally make errors.In this work, we present a user study in the field in which participants (N=123) interacted with a robot that may or may not have a ToM, and may or may not exhibit erroneous behaviour. Our findings indicate that a robot with ToM is perceived as more reliable, and they trusted it more than a robot without a ToM even when the robot made errors. Finally, ToM results to be a key driver for tuning people’s trust in the robot even when the initial condition of the interaction changed (i.e., loss and regain of trust in a longer relationship).
Alessandra Rossi 0001, Antonio Andriella, Silvia Rossi 0002, Carme Torras, Guillem Alenyà
RO-MAN5
2021 Online Action Recognition
abstract
Recognition in planning seeks to find agent intentions, goals or activities given a set of observations and a knowledge library (e.g. goal states, plans or domain theories). In this work we introduce the problem of Online Action Recognition. It consists in recognizing, in an open world, the planning action that best explains a partially observable state transition from a knowledge library of first-order STRIPS actions, which is initially empty. We frame this as an optimization problem, and propose two algorithms to address it: Action Unification (AU) and Online Action Recognition through Unification (OARU). The former builds on logic unification and generalizes two input actions using weighted partial MaxSAT. The latter looks for an action within the library that explains an observed transition. If there is such action, it generalizes it making use of AU, building in this way an AU hierarchy. Otherwise, OARU inserts a Trivial Grounded Action (TGA) in the library that explains just that transition. We report results on benchmarks from the International Planning Competition and PDDLGym, where OARU recognizes actions accurately with respect to expert knowledge, and shows real-time performance.
Alejandro Suárez-Hernández, Javier Segovia-Aguas, Carme Torras, Guillem Alenyà
AAAI4
2021 SMPLicit: Topology-Aware Generative Model for Clothed People
abstract
In this paper we introduce SMPLicit, a novel generative model to jointly represent body pose, shape and clothing geometry. In contrast to existing learning-based approaches that require training specific models for each type of garment, SMPLicit can represent in a unified manner different garment topologies (e.g. from sleeveless tops to hoodies and to open jackets), while controlling other properties like the garment size or tightness/looseness. We show our model to be applicable to a large variety of garments including T-shirts, hoodies, jackets, shorts, pants, skirts, shoes and even hair. The representation flexibility of SMPLicit builds upon an implicit model conditioned with the SMPL human body parameters and a learnable latent space which is semantically interpretable and aligned with the clothing attributes. The proposed model is fully differentiable, allowing for its use into larger end-to-end trainable systems. In the experimental section, we demonstrate SMPLicit can be readily used for fitting 3D scans and for 3D reconstruction in images of dressed people. In both cases we are able to go beyond state of the art, by retrieving complex garment geometries, handling situations with multiple clothing layers and providing a tool for easy outfit editing. To stimulate further research in this direction, we will make our code and model publicly available at http://www.iri.upc.edu/people/ecorona/smplicit/.
Enric Corona, Albert Pumarola, Guillem Alenyà, Gerard Pons-Moll, Francesc Moreno-Noguer
CVPR3
2021 Self-Supervised Policy Adaptation during Deployment
Nicklas Hansen 0001, Rishabh Jangir, Yu Sun 0020, Guillem Alenyà, Pieter Abbeel, Alexei A. Efros, Lerrel Pinto, Xiaolong Wang 0004
ICLR4
2021 Automatic Learning of Cognitive Exercises for Socially Assistive Robotics
abstract
In this paper, we present a learning approach to facilitate the teaching of new board exercises to assistive robotic systems. We formulate the problem as the learning of action models using Boolean predicates, disjunctive preconditions, and existential quantifiers from demonstrations of successful exercise executions. To be able to cope with exercises whose rules depend on a set of features that are initialized at the beginning of each play-out, we introduce the concept of dynamic context. Furthermore, we show how the learnt knowledge can be represented intuitively in a graphical interface that helps the caregiver understand what the system has learnt. As validation, we conducted a user study in which we evaluated whether and to which extent different types of feedback can affect the subjects’ performance while teaching three types of exercises: (1) sorting numbers; (2) arranging letters; and (3) reproducing shapes sequences in reversed order. The results suggest that textual and graphical feedback are beneficial.
Alejandro Suárez-Hernández, Antonio Andriella, Aleksandar Taranovic, Javier Segovia-Aguas, Carme Torras, Guillem Alenyà
RO-MAN6
2021 Are Preferences Useful for Better Assistance?: A Physically Assistive Robotics User Study
abstract
Assistive Robots have an inherent need of adapting to the user they are assisting. This is crucial for the correct development of the task, user safety, and comfort. However, adaptation can be performed in several manners. We believe user preferences are key to this adaptation. In this article, we evaluate the use of preferences for Physically Assistive Robotics tasks in a Human-Robot Interaction user evaluation. Three assistive tasks have been implemented consisting of assisted feeding, shoe-fitting, and jacket dressing, where the robot performs each task in a different manner based on user preferences. We assess the ability of the users to determine which execution of the task used their chosen preferences (if any). The obtained results show that most of the users were able to successfully guess the cases where their preferences were used even when they had not seen the task before. We also observe that their satisfaction with the task increases when the chosen preferences are employed. Finally, we also analyze the user’s opinions regarding assistive tasks and preferences, showing promising expectations as to the benefits of adapting the robot behavior to the user through preferences.
Gerard Canal, Carme Torras, Guillem Alenyà
ACM Trans. Hum. Robot Interact.3
2020 Context-Aware Human Motion Prediction
abstract
The problem of predicting human motion given a sequence of past observations is at the core of many applications in robotics and computer vision. Current state-of-the-art formulate this problem as a sequence-to-sequence task, in which a historical of 3D skeletons feeds a Recurrent Neural Network (RNN) that predicts future movements, typically in the order of 1 to 2 seconds. However, one aspect that has been obviated so far, is the fact that human motion is inherently driven by interactions with objects and/or other humans in the environment. In this paper, we explore this scenario using a novel context-aware motion prediction architecture. We use a semantic-graph model where the nodes parameterize the human and objects in the scene and the edges their mutual interactions. These interactions are iteratively learned through a graph attention layer, fed with the past observations, which now include both object and human body motions. Once this semantic graph is learned, we inject it to a standard RNN to predict future movements of the human/s and object/s. We consider two variants of our architecture, either freezing the contextual interactions in the future of updating them. A thorough evaluation in the “Whole-Body Human Motion Database” [29] shows that in both cases, our context-aware networks clearly outperform baselines in which the context information is not considered.
Enric Corona, Albert Pumarola, Guillem Alenyà, Francesc Moreno-Noguer
CVPR3
2020 GanHand: Predicting Human Grasp Affordances in Multi-Object Scenes
abstract
The rise of deep learning has brought remarkable progress in estimating hand geometry from images where the hands are part of the scene. This paper focuses on a new problem not explored so far, consisting in predicting how a human would grasp one or several objects, given a single RGB image of these objects. This is a problem with enormous potential in e.g. augmented reality, robotics or prosthetic design. In order to predict feasible grasps, we need to understand the semantic content of the image, its geometric structure and all potential interactions with a hand physical model. To this end, we introduce a generative model that jointly reasons in all these levels and 1) regresses the 3D shape and pose of the objects in the scene; 2) estimates the grasp types; and 3) refines the 51-DoF of a 3D hand model that minimize a graspability loss. To train this model we build the YCB-Affordance dataset, that contains more than 133k images of 21 objects in the YCB-Video dataset. We have annotated these images with more than 28M plausible 3D human grasps according to a 33-class taxonomy. A thorough evaluation in synthetic and real images shows that our model can robustly predict realistic grasps, even in cluttered scenes with multiple objects in close contact.
Enric Corona, Albert Pumarola, Guillem Alenyà, Francesc Moreno-Noguer, Grégory Rogez
CVPR3
2020 Dynamic Cloth Manipulation with Deep Reinforcement Learning
abstract
In this paper we present a Deep Reinforcement Learning approach to solve dynamic cloth manipulation tasks. Differing from the case of rigid objects, we stress that the followed trajectory (including speed and acceleration) has a decisive influence on the final state of cloth, which can greatly vary even if the positions reached by the grasped points are the same. We explore how goal positions for non-grasped points can be attained through learning adequate trajectories for the grasped points. Our approach uses few demonstrations to improve control policy learning, and a sparse reward approach to avoid engineering complex reward functions. Since perception of textiles is challenging, we also study different state representations to assess the minimum observation space required for learning to succeed. Finally, we compare different combinations of control policy encodings, demonstrations, and sparse reward learning techniques, and show that our proposed approach can learn dynamic cloth manipulation in an efficient way, i.e., using a reduced observation space, a few demonstrations, and a sparse reward.
Rishabh Jangir, Guillem Alenyà, Carme Torras
ICRA2
2020 Leveraging Multiple Environments for Learning and Decision Making: a Dismantling Use Case
abstract
Learning is usually performed by observing real robot executions. Physics-based simulators are a good alternative for providing highly valuable information while avoiding costly and potentially destructive robot executions. We present a novel approach for learning the probabilities of symbolic robot action outcomes. This is done leveraging different environments, such as physics-based simulators, in execution time. To this end, we propose MENID (Multiple Environment Noise Indeterministic Deictic) rules, a novel representation able to cope with the inherent uncertainties present in robotic tasks. MENID rules explicitly represent each possible outcomes of an action, keep memory of the source of the experience, and maintain the probability of success of each outcome. We also introduce an algorithm to distribute actions among environments, based on previous experiences and expected gain. Before using physics-based simulations, we propose a methodology for evaluating different simulation settings and determining the least time-consuming model that could be used while still producing coherent results. We demonstrate the validity of the approach in a dismantling use case, using a simulation with reduced quality as simulated system, and a simulation with full resolution where we add noise to the trajectories and some physical parameters as a representation of the real system.
Alejandro Suárez-Hernández, Thierry Gaugry, Javier Segovia-Aguas, Antonin Bernardin, Carme Torras, Maud Marchal, Guillem Alenyà
IROS7
2020 Discovering SOCIABLE: Using a Conceptual Model to Evaluate the Legibility and Effectiveness of Backchannel Cues in an Entertainment Scenario
abstract
Robots are expected to become part of everyday life. However, while there have been important breakthroughs during the recent decades in terms of technological advances, the ability of robots to interact with humans intuitively and effectively is still an open challenge. In this paper, we aim to evaluate how humans interpret and leverage backchannel cues exhibited by a robot which interacts with them in an entertainment context. To do so, a conceptual model was designed to investigate the legibility and the effectiveness of a designed social cue, called SOCial ImmediAcy BackchanneL cuE (SOCIABLE), on participant's performance. In addition, user's attitude and cognitive capability were integrated into the model as an estimator of participants' motivation and ability to process the cue. In working toward such a goal, we conducted a two-day long user study (N=114) at an international event with untrained participants who were not aware of the social cue the robot was able to provide. The results showed that participants were able to perceive the social signal generated from SOCIABLE and thus, they benefited from it. Our findings provide some important insights for the design of effective and instantaneous backchannel cues and the methodology for evaluating them in social robots.
Antonio Andriella, Ruben Huertas-Garcia, Santiago Forgas-Coll, Carme Torras, Guillem Alenyà
RO-MAN5
2020 Recurrent Neural Networks for Inferring Intentions in Shared Tasks for Industrial Collaborative Robots
abstract
Industrial robots are evolving to work closely with humans in shared spaces. Hence, robotic tasks are increasingly shared between humans and robots in collaborative settings. To enable a fluent human robot collaboration, robots need to predict and respond in real-time to worker's intentions. We present a method for early decision using force infor-mation. Forces are provided naturally by the user through the manipulation of a shared object in a collaborative task. The proposed algorithm uses a recurrent neural network to recognize operator's intentions. The algorithm is evaluated in terms of action recognition on a force dataset. It excels at detecting intentions when partial data is provided, enabling early detection and facilitating a quick robot reaction.
Marc Maceira, Alberto Olivares Alarcos, Guillem Alenyà
RO-MAN3
2020 A Grasping-Centered Analysis for Cloth Manipulation
abstract
Compliant and soft hands have gained a lot of attention in the past decade because of their ability to adapt to the shape of the objects, increasing their effectiveness for grasping. However, when it comes to grasping highly flexible objects such as textiles, we face the dual problem: it is the object that will adapt to the shape of the hand or gripper. In this context, the classic grasp analysis or grasping taxonomies are not suitable for describing textile objects grasps. This article proposes a novel definition of textile object grasps that abstracts from the robotic embodiment or hand shape and recovers concepts from the early neuroscience literature on hand prehension skills. This framework enables us to identify what grasps have been used in literature until now to perform robotic cloth manipulation, and allows for a precise definition of all the tasks that have been tackled in terms of manipulation primitives based on regrasps. In addition, we also review what grippers have been used. Our analysis shows how the vast majority of cloth manipulations have relied only on one type of grasp, and at the same time we identify several tasks that need more variety of grasp types to be executed successfully. Our framework is generic, provides a classification of cloth manipulation primitives and can inspire gripper design and benchmark construction for cloth manipulation.
Júlia Borràs Sol, Guillem Alenyà, Carme Torras
IEEE Trans. Robotics2
2019 Learning Robot Policies Using a High-Level Abstraction Persona-Behaviour Simulator
abstract
Collecting data in Human-Robot Interaction for training learning agents might be a hard task to accomplish. This is especially true when the target users are older adults with dementia since this usually requires hours of interactions and puts quite a lot of workload on the user. This paper addresses the problem of importing the Personas technique from HRI to create fictional patients' profiles. We propose a Persona-Behaviour Simulator tool that provides, with high-level abstraction, user's actions during an HRI task, and we apply it to cognitive training exercises for older adults with dementia. It consists of a Persona Definition that characterizes a patient along four dimensions and a Task Engine that provides information regarding the task complexity. We build a simulated environment where the high-level user's actions are provided by the simulator and the robot initial policy is learned using a Q-learning algorithm. The results show that the current simulator provides a reasonable initial policy for a defined Persona profile. Moreover, the learned robot assistance has proved to be robust to potential changes in the user's behaviour. In this way, we can speed up the fine-tuning of the rough policy during the real interactions to tailor the assistance to the given user. We believe the presented approach can be easily extended to account for other types of HRI tasks; for example, when input data is required to train a learning algorithm, but data collection is very expensive or unfeasible. We advocate that simulation is a convenient tool in these cases.
Antonio Andriella, Carme Torras, Guillem Alenyà
RO-MAN3
2018 Joining High-Level Symbolic Planning with Low-Level Motion Primitives in Adaptive HRI: Application to Dressing Assistance
abstract
For a safe and successful daily living assistance, far from the highly controlled environment of a factory, robots should be able to adapt to ever-changing situations. Programming such a robot is a tedious process that requires expert knowledge. An alternative is to rely on a high-level planner, but the generic symbolic representations used are not well suited to particular robot executions. Contrarily, motion primitives encode robot motions in a way that can be easily adapted to different situations. This paper presents a combined framework that exploits the advantages of both approaches. The number of required symbolic states is reduced, as motion primitives provide “smart actions” that take the current state and cope online with variations. Symbolic actions can include interactions (e.g., ask and inform) that are difficult to demonstrate. We show that the proposed framework can adapt to the user preferences (in terms of robot speed and robot verbosity), can readjust the trajectories based on the user movements, and can handle unforeseen situations. Experiments are performed in a shoe-dressing scenario. This scenario is particularly interesting because it involves a sufficient number of actions, and the human-robot interaction requires the handling of user preferences and unexpected reactions.
Gerard Canal, Emmanuel Pignat, Guillem Alenyà, Sylvain Calinon, Carme Torras
ICRA3
2018 Interleaving Hierarchical Task Planning and Motion Constraint Testing for Dual-Arm Manipulation
abstract
In recent years the topic of combining motion and symbolic planning to perform complex tasks in the field of robotics has received a lot of attention. The underlying idea is to have access at once to the reasoning capabilities of a task planner and to the ability of the motion planner to verify that the plan is feasible from a physical and geometrical point of view. The present work describes a framework to perform manipulation tasks that require the use of two robotic manipulators. To do so we employ a Hierarchical Task Network (HTN) planner interleaved with geometric constraint verification. In this framework we also consider observation actions and handle noisy perceptions from a probabilistic perspective. These ideas are put into practice by means of an experimental set-up in which two Barrett WAM robots have to cooperatively solve a geometric puzzle. Our findings provide further evidence that considering explicitly physical constraints during task planning, rather than deferring their validation to the moment of execution, is advantageous in terms of execution time and breadth of situations that can be handled.
Alejandro Suárez-Hernández, Guillem Alenyà, Carme Torras
IROS2
2018 Deciding the different robot roles for patient cognitive training
Antonio Andriella, Guillem Alenyà, Joan Hernández-Farigola, Carme Torras
Int. J. Hum. Comput. Stud.2
2018 Active garment recognition and target grasping point detection using deep learning
Enric Corona, Guillem Alenyà, Antonio Gabas, Carme Torras
Pattern Recognit.2
2018 Robot motion adaptation through user intervention and reinforcement learning
Aleksandar Jevtic, Adria Colome, Guillem Alenyà, Carme Torras
Pattern Recognit. Lett.3
2018 Teaching a Robot the Semantics of Assembly Tasks
abstract
We present a three-level cognitive system in a learning by demonstration context. The system allows for learning and transfer on the sensorimotor level as well as the planning level. The fundamentally different data structures associated with these two levels are connected by an efficient mid-level representation based on so-called “semantic event chains.” We describe details of the representations and quantify the effect of the associated learning procedures for each level under different amounts of noise. Moreover, we demonstrate the performance of the overall system by three demonstrations that have been performed at a project review. The described system has a technical readiness level (TRL) of 4, which in an ongoing follow-up project will be raised to TRL 6.
Thiusius Rajeeth Savarimuthu, Anders Glent Buch, Christian Schlette, Nils Wantia, Jürgen Roßmann, David Martínez Martínez, Guillem Alenyà, Carme Torras, Ales Ude, Bojan Nemec, Aljaz Kramberger, Florentin Wörgötter, Eren Erdal Aksoy, Jeremie Papon, Simon Haller, Justus H. Piater, Norbert Krüger
IEEE Trans. Syst. Man Cybern. Syst.7
2017 A taxonomy of preferences for physically assistive robots
abstract
Assistive devices and technologies are getting common and some commercial products are starting to be available. However, the deployment of robots able to physically interact with a person in an assistive manner is still a challenging problem. Apart from the design and control, the robot must be able to adapt to the user it is attending in order to become a useful tool for caregivers. This robot behavior adaptation comes through the definition of user preferences for the task such that the robot can act in the user's desired way. This article presents a taxonomy of user preferences for assistive scenarios, including physical interactions, that may be used to improve robot decision-making algorithms. The taxonomy categorizes the preferences based on their semantics and possible uses. We propose the categorization in two levels of application (global and specific) as well as two types (primary and modifier). Examples of real preference classifications are presented in three assistive tasks: feeding, shoe fitting and coat dressing.
Gerard Canal, Guillem Alenyà, Carme Torras
RO-MAN2
2017 Relational reinforcement learning with guided demonstrations
David Martínez Martínez, Guillem Alenyà, Carme Torras
Artif. Intell.2
2017 Relational Reinforcement Learning for Planning with Exogenous Effects
abstract
Probabilistic planners have improved recently to the point that they can solve difficult tasks with complex and expressive models. In contrast, learners cannot tackle yet the expressive models that planners do, which forces complex models to be mostly handcrafted. We propose a new learning approach that can learn relational probabilistic models with both action effects and exogenous effects. The proposed learning approach combines a multi-valued variant of inductive logic programming for the generation of candidate models, with an optimization method to select the best set of planning operators to model a problem. We also show how to combine this learner with reinforcement learning algorithms to solve complete problems. Finally, experimental validation is provided that shows improvements over previous work in both simulation and a robotic task. The robotic task involves a dynamic scenario with several agents where a manipulator robot has to clear the tableware on a table. We show that the exogenous effects learned by our approach allowed the robot to clear the table in a more efficient way.
David Martínez Martínez, Guillem Alenyà, Tony Ribeiro, Katsumi Inoue, Carme Torras
J. Mach. Learn. Res.2
2016 A 3D descriptor to detect task-oriented grasping points in clothing
Arnau Ramisa, Guillem Alenyà, Francesc Moreno-Noguer, Carme Torras
Pattern Recognit.2
2015 V-MIN: Efficient Reinforcement Learning through Demonstrations and Relaxed Reward Demands
David Martínez Martínez, Guillem Alenyà, Carme Torras
AAAI2
2015 3D Sensor planning framework for leaf probing
abstract
Modern plant phenotyping requires active sensing technologies and particular exploration strategies. This article proposes a new method for actively exploring a 3D region of space with the aim of localizing special areas of interest for manipulation tasks over plants. In our method, exploration is guided by a multi-layer occupancy grid map. This map, together with a multiple-view estimator and a maximum-information-gain gathering approach, incrementally provides a better understanding of the scene until a task termination criterion is reached. This approach is designed to be applicable for any task entailing 3D object exploration where some previous knowledge of its general shape is available. Its suitability is demonstrated here for an eye-in-hand arm configuration in a leaf probing application.
Sergi Foix, Guillem Alenyà, Carme Torras
IROS2
2015 Safe robot execution in model-based reinforcement learning
abstract
Task learning in robotics requires repeatedly executing the same actions in different states to learn the model of the task. However, in real-world domains, there are usually sequences of actions that, if executed, may produce unrecoverable errors (e.g. breaking an object). Robots should avoid repeating such errors when learning, and thus explore the state space in a more intelligent way. This requires identifying dangerous action effects to avoid including such actions in the generated plans, while at the same time enforcing that the learned models are complete enough for the planner not to fall into dead-ends. We thus propose a new learning method that allows a robot to reason about dead-ends and their causes. Some such causes may be dangerous action effects (i.e., leading to unrecoverable errors if the action were executed in the given state) so that the method allows the robot to skip the exploration of risky actions and guarantees the safety of planned actions. If a plan might lead to a dead-end (e.g., one that includes a dangerous action effect), the robot tries to find an alternative safe plan and, if not found, it actively asks a teacher whether the risky action should be executed. This method permits learning safe policies as well as minimizing unrecoverable errors during the learning process. Experimental validation of the approach is provided in two different scenarios: a robotic task and a simulated problem from the international planning competition. Our approach greatly increases success ratios in problems where previous approaches had high probabilities of failing.
David Martínez Martínez, Guillem Alenyà, Carme Torras
IROS2
2015 Planning robot manipulation to clean planar surfaces
David Martínez Martínez, Guillem Alenyà, Carme Torras
Eng. Appl. Artif. Intell.2
2014 The HumanoidLab - Involving Students in a Research Centre Through an Educational Initiative
abstract
Abstract: The HumanoidLab is a more than 5 year old activity aimed to use educational robots to approach students to our Research Centre. Different commercial educative humanoid platforms have been used to introduce students to different aspects of robotics using projects and offering guidance and assistance. About 40 students have performed small mechanics, electronics or programming projects that are used to improve the robots by
Guillem Alenyà, José Luis Rivero, Aleix Rull, Patrick Grosch
CSEDU (2)1
2014 Realtime tracking and grasping of a moving object from range video
abstract
In this paper we present an automated system that is able to track and grasp a moving object within the workspace of a manipulator using range images acquired with a Microsoft Kinect sensor. Realtime tracking is achieved by a geometric particle filter on the affine group. Based on the tracked output, the pose of a 7-DoF WAM robotic arm is continuously updated using dynamic motor primitives until a distance measure between the tracked object and the gripper mounted on the arm is below a threshold. Then, it closes its three fingers and grasps the object. The tracker works in realtime and is robust to noise and partial occlusions. Using only the depth data makes our tracker independent of texture which is one of the key design goals in our approach. An experimental evaluation is provided along with a comparison of the proposed tracker with state-of-the-art approaches, including the OpenNI-tracker. The developed system is integrated with ROS and made available as part of IRI's ROS stack.
Farzad Husain, Adria Colome, Babette Dellen, Guillem Alenyà, Carme Torras
ICRA4
2014 Active learning of manipulation sequences
abstract
We describe a system allowing a robot to learn goal-directed manipulation sequences such as steps of an assembly task. Learning is based on a free mix of exploration and instruction by an external teacher, and may be active in the sense that the system tests actions to maximize learning progress and asks the teacher if needed. The main component is a symbolic planning engine that operates on learned rules, defined by actions and their pre- and postconditions. Learned by model-based reinforcement learning, rules are immediately available for planning. Thus, there are no distinct learning and application phases. We show how dynamic plans, replanned after every action if necessary, can be used for automatic execution of manipulation sequences, for monitoring of observed manipulation sequences, or a mix of the two, all while extending and refining the rule base on the fly. Quantitative results indicate fast convergence using few training examples, and highly effective teacher intervention at early stages of learning.
David Martínez Martínez, Guillem Alenyà, Pablo Jiménez, Carme Torras, Jürgen Roßmann, Nils Wantia, Eren Erdal Aksoy, Simon Haller, Justus H. Piater
ICRA2
2014 Learning RGB-D descriptors of garment parts for informed robot grasping
Arnau Ramisa, Guillem Alenyà, Francesc Moreno-Noguer, Carme Torras
Eng. Appl. Artif. Intell.2
2013 External force estimation during compliant robot manipulation
abstract
This paper presents a method to estimate external forces exerted on a manipulator during motion, avoiding the use of a sensor. The method is based on task-oriented dynamics model learning and a robust disturbance state observer. The combination of both leads to an efficient torque observer that can be incorporated to any control scheme. The use of a learning-based approach avoids the need of analytical models of joints' friction or Coriolis dynamics effects.
Adria Colome, Diego Pardo, Guillem Alenyà, Carme Torras
ICRA3
2013 FINDDD: A fast 3D descriptor to characterize textiles for robot manipulation
abstract
Most current depth sensors provide 2.5D range images in which depth values are assigned to a rectangular 2D array. In this paper we take advantage of this structured information to build an efficient shape descriptor which is about two orders of magnitude faster than competing approaches, while showing similar performance in several tasks involving deformable object recognition. Given a 2D patch surrounding a point and its associated depth values, we build the descriptor for that point, based on the cumulative distances between their normals and a discrete set of normal directions. This processing is made very efficient using integral images, even allowing to compute descriptors for every range image pixel in a few seconds. The discriminative power of our descriptor, dubbed FINDDD, is evaluated in three different scenarios: recognition of specific cloth wrinkles, instance recognition from geometry alone, and detection of reliable and informed grasping points.
Arnau Ramisa, Guillem Alenyà, Francesc Moreno-Noguer, Carme Torras
IROS2
2012 Information-Gain View Planning for Free-Form Object Reconstruction with a 3D ToF Camera
Sergi Foix, Simon Kriegel, Stefan Fuchs, Guillem Alenyà, Carme Torras
ACIVS4
2012 Single image 3D human pose estimation from noisy observations
abstract
Markerless 3D human pose detection from a single image is a severely underconstrained problem because different 3D poses can have similar image projections. In order to handle this ambiguity, current approaches rely on prior shape models that can only be correctly adjusted if 2D image features are accurately detected. Unfortunately, although current 2D part detector algorithms have shown promising results, they are not yet accurate enough to guarantee a complete disambiguation of the 3D inferred shape. In this paper, we introduce a novel approach for estimating 3D human pose even when observations are noisy. We propose a stochastic sampling strategy to propagate the noise from the image plane to the shape space. This provides a set of ambiguous 3D shapes, which are virtually undistinguishable from their image projections. Disambiguation is then achieved by imposing kinematic constraints that guarantee the resulting pose resembles a 3D human shape. We validate the method on a variety of situations in which state-of-the-art 2D detectors yield either inaccurate estimations or partly miss some of the body parts.
Edgar Simo-Serra, Arnau Ramisa, Guillem Alenyà, Carme Torras, Francesc Moreno-Noguer
CVPR3
2012 Using depth and appearance features for informed robot grasping of highly wrinkled clothes
abstract
Detecting grasping points is a key problem in cloth manipulation. Most current approaches follow a multiple re-grasp strategy for this purpose, in which clothes are sequentially grasped from different points until one of them yields to a desired configuration. In this paper, by contrast, we circumvent the need for multiple re-graspings by building a robust detector that identifies the grasping points, generally in one single step, even when clothes are highly wrinkled. In order to handle the large variability a deformed cloth may have, we build a Bag of Features based detector that combines appearance and 3D geometry features. An image is scanned using a sliding window with a linear classifier, and the candidate windows are refined using a non-linear SVM and a “grasp goodness” criterion to select the best grasping point. We demonstrate our approach detecting collars in deformed polo shirts, using a Kinect camera. Experimental results show a good performance of the proposed method not only in identifying the same trained textile object part under severe deformations and occlusions, but also the corresponding part in other clothes, exhibiting a degree of generalization.
Arnau Ramisa, Guillem Alenyà, Francesc Moreno-Noguer, Carme Torras
ICRA2
2012 POMDP approach to robotized clothes separation
abstract
Rigid object manipulation with robots has mainly relied on precise, expensive models and deterministic sequences. Given the great complexity of accurately modeling deformable objects, their manipulation seems to call for a rather different approach. This paper proposes a probabilistic planner, based on a Partially Observable Markov Decision Process (POMDP), targeted at reducing the inherent uncertainty of deformable object sorting. It is shown that a small set of unreliable actions and inaccurate perceptions suffices to accomplish the task, provided faithful statistics on both of them are collected beforehand. The planner has been applied to a clothes sorting task in a real case context with a depth and color sensor and a robotic arm. Experimental results show the promise of the approach since more than 95% certainty of having isolated a piece of clothing is reached in an average of four steps for quite entangled initial clothing configurations.
Pol Monso, Guillem Alenyà, Carme Torras
IROS2
2011 3D modelling of leaves from color and ToF data for robotized plant measuring
abstract
Supervision of long-lasting extensive botanic experiments is a promising robotic application that some recent technological advances have made feasible. Plant modelling for this application has strong demands, particularly in what concerns 3D information gathering and speed. This paper shows that Time-of-Flight (ToF) cameras achieve a good compromise between both demands, providing a suitable complement to color vision. A new method is proposed to segment plant images into their composite surface patches by combining hierarchical color segmentation with quadratic surface fitting using ToF depth data. Experimentation shows that the interpolated depth maps derived from the obtained surfaces fit well the original scenes. Moreover, candidate leaves to be approached by a measuring instrument are ranked, and then robot-mounted cameras move closer to them to validate their suitability to being sampled. Some ambiguities arising from leaves overlap or occlusions are cleared up in this way. The work is a proof-of-concept that dense color data combined with sparse depth as provided by a ToF camera yields a good enough 3D approximation for automated plant measuring at the high throughput imposed by the application.
Guillem Alenyà, Babette Dellen, Carme Torras
ICRA1
2011 Segmenting color images into surface patches by exploiting sparse depth data
abstract
We present a new method for segmenting color images into their composite surfaces by combining color segmentation with model-based fitting utilizing sparse depth data, acquired using time-of-flight (Swissranger, PMD CamCube) and stereo techniques. The main target of our work is the segmentation of plant structures, i.e., leaves, from color-depth images, and the extraction of color and 3D shape information for automating manipulation tasks. Since segmentation is performed in the dense color space, even sparse, incomplete, or noisy depth information can be used. This kind of data often represents a major challenge for methods operating in the 3D data space directly. To achieve our goal, we construct a three-stage segmentation hierarchy by segmenting the color image with different resolutions-assuming that “true” surface boundaries must appear at some point along the segmentation hierarchy. 3D surfaces are then fitted to the color-segment areas using depth data. Those segments which minimize the fitting error are selected and used to construct a new segmentation. Then, an additional region merging and a growing stage are applied to avoid over-segmentation and label previously unclustered points. Experimental results demonstrate that the method is successful in segmenting a variety of domestic objects and plants into quadratic surfaces. At the end of the procedure, the sparse depth data is completed using the extracted surface models, resulting in dense depth maps. For stereo, the resulting disparity maps are compared with ground truth and the average error is computed.
Babette Dellen, Guillem Alenyà, Sergi Foix, Carme Torras
WACV2
2010 Planning Stacking Operations with an Unknown Number of Objects
Lluis Trilla, Guillem Alenyà
ICINCO (2)2
2010 Object modeling using a ToF camera under an uncertainty reduction approach
abstract
Time-of-Flight (ToF) cameras deliver 3D images at 25 fps, offering great potential for developing fast object modeling algorithms. Surprisingly, this potential has not been extensively exploited up to now. A reason for this is that, since the acquired depth images are noisy, most of the available registration algorithms are hardly applicable. A further difficulty is that the transformations between views are in general not accurately known, a circumstance that multi-view object modeling algorithms do not handle properly under noisy conditions. In this work, we take into account both uncertainty sources (in images and camera poses) to generate spatially consistent 3D object models fusing multiple views with a probabilistic approach. We propose a method to compute the covariance of the registration process, and apply an iterative state estimation method to build object models under noisy conditions.
Sergi Foix, Guillem Alenyà, Juan Andrade-Cetto, Carme Torras
ICRA2
2010 Camera motion estimation by tracking contour deformation: Precision analysis
Guillem Alenyà, Carme Torras
Image Vis. Comput.1
2009 A comparison of three methods for measure of Time to Contact
abstract
Time to contact (TTC) is a biologically inspired method for obstacle detection and reactive control of motion that does not require scene reconstruction or 3D depth estimation. Estimating TTC is difficult because it requires a stable and reliable estimate of the rate of change of distance between image features. In this paper we propose a new method to measure time to contact, active contour affine scale (ACAS). We experimentally and analytically compare ACAS with two other recently proposed methods: scale invariant ridge segments (SIRS), and image brightness derivatives (IBD). Our results show that ACAS provides a more accurate estimation of TTC when the image flow may be approximated by an affine transformation, while SIRS provides an estimate that is generally valid, but may not always be as accurate as ACAS, and IBD systematically over-estimate time to contact.
Guillem Alenyà, Amaury Nègre, James L. Crowley
IROS1
2008 Recovering epipolar direction from two affine views of a planar object
Maria Alberich-Carramiñana, Guillem Alenyà, Juan Andrade-Cetto, Elisa Martínez Marroquín, Carme Torras
Comput. Vis. Image Underst.2
2007 Depth from the visual motion of a planar target induced by zooming
abstract
Robot egomotion can be estimated from an acquired video stream up to the scale of the scene. To remove this uncertainty (and obtain true egomotion), a distance within the scene needs to be known. If no a priori knowledge on the scene is assumed, the usual solution is to derive "in some way" the initial distance from the camera to a target object. This paper proposes a new, very simple way to obtain such a distance, when a zooming camera is available and there is a planar target in the scene. Similarly to "two-grid calibration" algorithms, no estimation of the camera parameters is required, and no assumption on the optical axis stability between the different focal lengths is needed. Quite the reverse, the non stability of the optical axis between the different focal lengths is the key ingredient that enables to derive our depth estimate, by applying a result in projective geometry. Experiments carried out on a mobile robot platform show the promise of the approach.
Guillem Alenyà, Maria Alberich-Carramiñana, Carme Torras
ICRA1
2006 Affine Epipolar Direction from Two Views of a Planar Contour
Maria Alberich-Carramiñana, Guillem Alenyà, Juan Andrade-Cetto, Elisa Martínez Marroquín, Carme Torras
ACIVS2
2005 Using Laser and Vision to Locate a Robot in an Industrial Environment: A Practical Experience
abstract
The fully flexible navigation of autonomous vehicles in industrial environments is still unsolved. It is hard to conciliate strict precision requirements with quick adaptivity to new settings without undergoing costly rearrangements. We are pursuing a research project trying to combine the precision of laser-based local positioning with the flexibility of vision-based robot motion estimation. An enhanced circle approach to dynamic triangulation combining laser and odometric signals has been used to improve positioning accuracy. As regards to vision, a novel technique relating the deformation of contours in an image sequence to the 3D motion underwent by the camera has been developed. Interestingly, contours are fitted to objects already present in the environment, without requiring any presetting. In this paper, we describe a practical experience conducted in the warehouse of a beer production factory in Barcelona. A database containing the laser readings, image sequences and robot odometry along several trajectories was compiled, and subsequently processed off-line in order to assess the accuracies of both techniques under a variety of circumstances. In all, vision-based estimation turned out to be about one order of magnitude less precise than laser-based positioning, which qualifies the vision-based technique as a promising alternative to accomplish robot transfers across long distances, such as those needed in a warehouse, while backing up on laser-based positioning when accurate docking for loading and unloading operations is needed.
Guillem Alenyà, Josep Escoda, Antonio B. Martínez, Carme Torras
ICRA1