Lorenzo Natale

dblp:17/4667 · DBLP profile ↗
← Back
91ranked-venue papers
0as first author
31since 2021 · last 2026
0000-0002-8777-5233ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 79 · 26 since 2021Systems, architecture and hardware · 63 · 17 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 7 since 2021Human-computer interaction and ubiquitous computing · 10 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 A Baseline Study and Benchmark for Few-Shot Open-Set Action Recognition with Feature Residual Discrimination
Stefano Berti, Giulia Pasquale, Lorenzo Natale
ICPR (4)3
2026 Gaze Estimation Learning Architecture as Support to Affective, Social and Cognitive Studies in Natural Human-Robot Interaction
abstract
Gaze is a crucial social cue in any interacting scenario and drives many mechanisms of social cognition (joint and shared attention, predicting human intention and coordinating tasks). Gaze is an indication of social and emotional functions affecting the way the emotions are perceived. Evidence shows that embodied humanoid robots endowed with social abilities can be seen as sophisticated stimuli to study several mechanisms of human social cognition while increasing engagement and ecological validity. In this context, building a robotic perception system to automatically estimate the human gaze only relying on robot’s sensors is still demanding. Main goal of the article is to propose a learning robotic architecture estimating the human gaze direction in table-top scenarios without any external hardware. Table-top tasks are largely used in experimental psychology because they are suitable to implement numerous face-to-face collaborative scenarios. Such an architecture can provide a valuable support in studies where external hardware might represent an obstacle to spontaneous human behaviour, especially in environments less controlled than the laboratory (e.g., in clinical settings). A novel dataset was also collected with the humanoid robot iCub, including images annotated from 24 participants in different gaze conditions.
Maria Lombardi, Elisa Maiettini, Agnieszka Wykowska, Lorenzo Natale
ACM Trans. Hum. Robot Interact.4
2025 Embodied Image Captioning: Self-Supervised Learning Agents for Spatially Coherent Image Descriptions
abstract
We present a self-supervised method to improve an agent's abilities in describing arbitrary objects while actively exploring a generic environment. This is a challenging problem, as current models struggle to obtain coherent image captions due to different camera viewpoints and clutter. We propose a three-phase framework to fine-tune existing captioning models that enhances caption accuracy and consistency across views via a consensus mechanism. First, an agent explores the environment, collecting noisy image-caption pairs. Then, a consistent pseudo-caption for each object instance is distilled via consensus using a large language model. Finally, these pseudo-captions are used to fine-tune an off-the-shelf captioning model, with the addition of contrastive learning. We analyse the performance of the combination of captioning models, exploration policies, pseudo-labeling methods, and fine-tuning strategies, on our manually labeled test set. Results show that a policy can be trained to mine samples with higher disagreement compared to classical baselines. Our pseudo-captioning method, in combination with all policies, has a higher semantic similarity compared to other existing methods, and fine-tuning improves caption accuracy and consistency by a significant margin. Code and test set annotations available at https://hsp-iit.github.io/embodied-captioning/
Tommaso Galliena, Tommaso Apicella, Stefano Rosa, Pietro Morerio, Alessio Del Bue, Lorenzo Natale
ICCV6
2025 FeelAnyForce: Estimating Contact Force Feedback from Tactile Sensation for Vision-Based Tactile Sensors
abstract
In this paper, we tackle the problem of estimating 3D contact forces using vision-based tactile sensors. In particular, our goal is to estimate contact forces over a large range (up to 15 N) on any objects while generalizing across different vision-based tactile sensors. Thus, we collected a dataset of over 200K indentations using a robotic arm that pressed various indenters onto a GelSight Mini sensor mounted on a force sensor and then used the data to train a multi-head transformer for force regression. Strong generalization is achieved via accurate data collection and multi-objective optimization that leverages depth contact images. Despite being trained only on primitive shapes and textures, the regressor achieves a mean absolute error of 4% on a dataset of unseen real-world objects. We further evaluate our approach's generalization capability to other GelSight mini and DIGIT sensors, and propose a reproducible calibration procedure for other sensors. Finally, the method was evaluated on real-world tasks, including weighing objects and controlling the deformation of delicate objects. Supplementary material and demo are available at http://prg.cs.umd.edu/FeelAnyForce.
Amir-Hossein Shahidzadeh, Gabriele M. Caddeo, Koushik Alapati, Lorenzo Natale, Cornelia Fermüller, Yiannis Aloimonos
ICRA4
2025 Bring Your Own Grasp Generator: Leveraging Robot Grasp Generation for Prosthetic Grasping
abstract
One of the most important research challenges in upper-limb prosthetics is enhancing the user-prosthesis communication to closely resemble the experience of a natural limb. As prosthetic devices become more complex, users often struggle to control the additional degrees of freedom. In this context, leveraging shared-autonomy principles can significantly improve the usability of these systems. In this paper, we present a novel eye-in-hand prosthetic grasping system that follows these principles. Our system initiates the approach-to-grasp action based on user's command and automatically configures the DoFs of a prosthetic hand. First, it reconstructs the 3D geometry of the target object without the need of a depth camera. Then, it tracks the hand motion during the approach-to-grasp action and finally selects a candidate grasp configuration according to user's intentions. We deploy our system on the Hannes prosthetic hand and test it on able-bodied subjects and amputees to validate its effectiveness. We compare it with a multi-DoF prosthetic control baseline and find that our method enables faster grasps, while simplifying the user experience. Code and demo videos are available online at this https URL.
Giuseppe Stracquadanio, Federico Vasile, Elisa Maiettini, Nicoló Boccardo, Lorenzo Natale
ICRA5
2025 Continuous Wrist Control on the Hannes Prosthesis: A Vision-Based Shared Autonomy Framework
abstract
Most control techniques for prosthetic grasping focus on dexterous fingers control, but overlook the wrist motion. This forces the user to perform compensatory movements with the elbow, shoulder and hip to adapt the wrist for grasping. We propose a computer vision-based system that leverages the collaboration between the user and an automatic system in a shared autonomy framework, to perform continuous control of the wrist degrees of freedom in a prosthetic arm, promoting a more natural approach-to-grasp motion. Our pipeline allows to seamlessly control the prosthetic wrist to follow the target object and finally orient it for grasping according to the user intent. We assess the effectiveness of each system component through quantitative analysis and finally deploy our method on the Hannes prosthetic arm. Code and videos: https: //hsp-iit.github.io/hannes-wrist-control.
Federico Vasile, Elisa Maiettini, Giulia Pasquale, Nicoló Boccardo, Lorenzo Natale
ICRA5
2025 HannesImitation: Grasping with the Hannes Prosthetic Hand via Imitation Learning
abstract
Recent advancements in control of prosthetic hands have focused on increasing autonomy through the use of cameras and other sensory inputs. These systems aim to reduce the cognitive load on the user by automatically controlling certain degrees of freedom. In robotics, imitation learning has emerged as a promising approach for learning grasping and complex manipulation tasks while simplifying data collection. Its application to the control of prosthetic hands remains, however, largely unexplored. Bridging this gap could enhance dexterity restoration and enable prosthetic devices to operate in more unconstrained scenarios, where tasks are learned from demonstrations rather than relying on manually annotated sequences. To this end, we present HannesImitationPolicy, an imitation learning-based method to control the Hannes prosthetic hand, enabling object grasping in unstructured environments. Moreover, we introduce the HannesImitationDataset comprising grasping demonstrations in table, shelf, and human-to-prosthesis handover scenarios. We leverage such data to train a single diffusion policy and deploy it on the prosthetic hand to predict the wrist orientation and hand closure for grasping. Experimental evaluation demonstrates successful grasps across diverse objects and conditions. Finally, we show that the policy outperforms a segmentation-based visual servo controller in unstructured scenarios. Additional material is provided on our project page: https://hsp-iit.github.io/HannesImitation.
Carlo Alessi, Federico Vasile, Federico Ceola, Giulia Pasquale, Nicoló Boccardo, Lorenzo Natale
IROS6
2025 Code Generation and Monitoring for Deliberation Components in Autonomous Robots
abstract
Hand-coded deliberation components are prone to flaws that may not be discovered before deployment and that can be harmful to the robot and its execution environment, including the people within it. To reduce development effort and at the same time increase confidence in robot’s safety, we propose to model deliberation components at a conceptual level, to automatically generate code from such models and also to monitor their execution during robot operation. We present two tools, one which compiles models of deliberation components into executable code, and one which generates runtime monitors from the models. We have tested them in simulation, to demonstrate the usefulness of combining together model-based development, code generation, and monitoring.
Stefano Bernagozzi, Sofia Faraci, Enrico Ghiorzi, K. Pedemonte, Lorenzo Natale, Armando Tacchella
IROS5
2025 6-DoF Object Tracking with Event-based Optical Flow and Frames
abstract
Tracking the position and orientation of objects in space (i.e., in 6-DoF) in real time is a fundamental problem in robotics for environment interaction. It becomes more challenging when objects move at high-speed due to frame rate limitations in conventional cameras and motion blur. Event cameras are characterized by high temporal resolution, low latency and high dynamic range, that can potentially overcome the impacts of motion blur. Traditional RGB cameras provide rich visual information that is more suitable for the challenging task of single-shot object pose estimation. In this work, we propose using event-based optical flow combined with an RGB based global object pose estimator for 6-DoF pose tracking of objects at high-speed, exploiting the core advantages of both types of vision sensors. Specifically, we propose an event-based optical flow algorithm for object motion measurement to implement an object 6-DoF velocity tracker. By integrating the tracked object 6-DoF velocity with low frequency estimated pose from the global pose estimator, the method can track pose when objects move at high-speed. The proposed algorithm is tested and validated on both synthetic and real world data, demonstrating its effectiveness, especially in high-speed motion scenarios.
Arren Glover, Chiara Bartolozzi, Lorenzo Natale
IROS4
2025 Would you let a humanoid play storytelling with your child? A usability study on LLM-powered narrative Humanoid-Robot Interaction
abstract
A key challenge in human-robot interaction research lies in developing robotic systems that can effectively perceive and interpret social cues, facilitating natural and adaptive interactions. In this work, we present a novel framework for enhancing the attention of the iCub humanoid robot by integrating advanced perceptual abilities to recognise social cues, understand surroundings through generative models, such as ChatGPT, and respond with contextually appropriate social behaviour. Specifically, we propose an interaction task implementing a narrative protocol (storytelling task) in which the human and the robot create a short imaginary story together, exchanging in turn cubes with creative images placed on them. To validate the protocol and the framework, experiments were performed to quantify the degree of usability and the quality of experience perceived by participants interacting with the system. Such a system can be beneficial in promoting effective humanrobot collaborations, especially in assistance, education and rehabilitation scenarios where the social awareness and the robot responsiveness play a pivotal role.
Maria Lombardi, Carmela Calabrese, Davide Ghiglino, Caterina Foglino, Davide De Tommaso, Giulia Da Lisca, Lorenzo Natale, Agnieszka Wykowska
IROS7
2025 Gaussian-Augmented Physics Simulation and System Identification with Complex Colliders
abstract
System identification involving the geometry, appearance, and physical properties from video observations is a challenging task with applications in robotics and graphics. Recent approaches have relied on fully differentiable Material Point Method (MPM) and rendering for simultaneous optimization of these properties. However, they are limited to simplified object-environment interactions with planar colliders and fail in more challenging scenarios where objects collide with non-planar surfaces. We propose AS-DiffMPM, a differentiable MPM framework that enables physical property estimation with arbitrarily shaped colliders. Our approach extends existing methods by incorporating a differentiable collision handling mechanism, allowing the target object to interact with complex rigid bodies while maintaining end-to-end optimization. We show AS-DiffMPM can be easily interfaced with various novel view synthesis methods as a framework for system identification from visual observations.
Federico Vasile, Ri-Zhao Qiu, Lorenzo Natale
NeurIPS3
2025 The duration of robot gaze affects people's attitudes towards humanoid robots
abstract
Gaze plays a crucial role in human social behavior. Notably, the same applies also to interactions between humans and robots, as gaze can communicate intentions and express interest or aversion similarly to what happens among humans. Besides the direction of gaze (direct vs. averted gaze), its temporal characteristics, such as duration, significantly affect our perception and interpretation of the other’s behavior. In the context of Human-Robot Interaction (HRI), this is still poorly investigated. Thus, the present study aimed to investigate whether, and how, the duration of the robot direct gaze impacts participants’ attitudes towards robots. To do so, participants observed the humanoid robot iCub, whose direct gaze varied in duration between 1 and 8 seconds. Then, they used three Likert scales to rate to what extent the robot gaze made them feel i) comfortable, ii) trustful, and iii) threatened, with participants’ rating operationalizing their attitudes towards the robot. Results showed that, overall, a positive relationship emerged between the duration of the robot gaze and participants’ attitudes, i.e., longer gaze duration led to higher ratings for all three Likert scales.
Cecilia Roselli, Maria Lombardi, Lorenzo Natale, Agnieszka Wykowska
RO-MAN3
2024 ConCon-Chi: Concept-Context Chimera Benchmark for Personalized Vision-Language Tasks
abstract
While recent Vision-Language (VL) models excel at open-vocabulary tasks, it is unclear how to use them with specific or uncommon concepts. Personalized Text-to-Image Retrieval (TIR) or Generation (TIG) are recently introduced tasks that represent this challenge, where the VL model has to learn a concept from few images and respectively discriminate or generate images of the target concept in arbitrary contexts. We identify the ability to learn new meanings and their compositionality with known ones as two key properties of a personalized system. We show that the available benchmarks offer a limited validation of personalized textual concept learning from images with respect to the above properties and introduce ConCon-Chi as a benchmark for both personalized TIR and TIG, designed to fill this gap. We modelled the new-meaning concepts by crafting chimeric objects and formulating a large, varied set of contexts where we photographed each object. To promote the compositionality assessment of the learned concepts with known contexts, we combined different contexts with the same concept, and vice-versa. We carry out a thorough evaluation of state-of-the-art methods on the resulting dataset. Our study suggests that future work on personalized TIR and TIG methods should focus on the above key properties, and we propose principles and a dataset for their performance assessment. Dataset: https://doi.org/10.48557/QJ1166 and code: https://github.com/hsp-iit/concon-chi_benchmark.
Andrea Rosasco, Stefano Berti, Giulia Pasquale, Damiano Malafronte, Shogo Sato, Hiroyuki Segawa, Tetsugo Inada, Lorenzo Natale
CVPR8
2024 Look Around and Learn: Self-training Object Detection by Exploration
Gianluca Scarpellini, Stefano Rosa, Pietro Morerio, Lorenzo Natale, Alessio Del Bue
ECCV (56)4
2024 Sim2Real Bilevel Adaptation for Object Surface Classification using Vision-Based Tactile Sensors
abstract
In this paper, we address the Sim2Real gap in the field of vision-based tactile sensors for classifying object surfaces. We train a Diffusion Model to bridge this gap using a relatively small dataset of real-world images randomly collected from unlabeled everyday objects via the DIGIT sensor. Subsequently, we employ a simulator to generate images by uniformly sampling the surface of objects from the YCB Model Set. These simulated images are then translated into the real domain using the Diffusion Model and automatically labeled to train a classifier. During this training, we further align features of the two domains using an adversarial procedure. Our evaluation is conducted on a dataset of tactile images obtained from a set of ten 3D-printed YCB objects. The results reveal a total accuracy of 81.9%, a significant improvement compared to the 34.7% achieved by the classifier trained solely on simulated images. This demonstrates the effectiveness of our approach. We further validate our approach using the classifier on a 6D object pose estimation task from tactile data.
Gabriele M. Caddeo, Andrea Maracani, Paolo Didier Alfano, Nicola A. Piga, Lorenzo Rosasco, Lorenzo Natale
ICRA6
2024 Mind the Error! Detection and Localization of Instruction Errors in Vision-and-Language Navigation
abstract
Vision-and-Language Navigation in Continuous Environments (VLN-CE) is one of the most intuitive yet challenging embodied AI tasks. Agents are tasked to navigate towards a target goal by executing a set of low-level actions, following a series of natural language instructions. All VLN-CE methods in the literature assume that language instructions are exact. However, in practice, instructions given by humans can contain errors when describing a spatial environment due to inaccurate memory or confusion. Current VLN-CE benchmarks do not address this scenario, making the state-of-the-art methods in VLN-CE fragile in the presence of erroneous instructions from human users. For the first time, we propose a novel benchmark dataset that introduces various types of instruction errors considering potential human causes. This benchmark provides valuable insight into the robustness of VLN systems in continuous environments. We observe a noticeable performance drop (up to −25%) in Success Rate when evaluating the state-of-the-art VLN-CE methods on our benchmark. Moreover, we formally define the task of Instruction Error Detection and Localization, and establish an evaluation protocol on top of our benchmark dataset. We also propose an effective method, based on a cross-modal transformer architecture, that achieves the best performance in error detection and localization, compared to baselines. Surprisingly, our proposed method has revealed errors in the validation set of the two commonly used datasets for VLN-CE, i.e., R2R-CE and RxR-CE, demonstrating the utility of our technique in other tasks.
Francesco Taioli, Stefano Rosa, Alberto Castellini, Lorenzo Natale, Alessio Del Bue, Alessandro Farinelli, Marco Cristani, Yiming Wang 0002
IROS4
2024 The impact of Compositionality in Zero-shot Multi-label action recognition for Object-based tasks
abstract
Addressing multi-label action recognition in videos represents a significant challenge for robotic applications in dynamic environments, especially when the robot is required to cooperate with humans in tasks that involve objects. Existing methods still struggle to recognize unseen actions or require extensive training data. To overcome these problems, we propose Dual-VCLIP, a unified approach for zero-shot multi-label action recognition. Dual-VCLIP enhances VCLIP, a zero-shot action recognition method, with the DualCoOp method for multi-label image classification. The strength of our method is that at training time it only learns two prompts, and it is therefore much simpler than other methods. We validate our method on the Charades dataset that includes a majority of object-based actions, demonstrating that - despite its simplicity - our method performs favorably with respect to existing methods on the complete dataset, and promising performance when tested on unseen actions. Our contribution emphasizes the impact of verb-object class-splits during robots’ training for new cooperative tasks, highlighting the influence on the performance and giving insights into mitigating biases. Dataset splits and code are publicly available on the project’s repository1.
Carmela Calabrese, Stefano Berti, Giulia Pasquale, Lorenzo Natale
RO-MAN4
2024 I2EDL: Interactive Instruction Error Detection and Localization
abstract
In the Vision-and-Language Navigation in Continuous Environments (VLN-CE) task, the human user guides an autonomous agent to reach a target goal via a series of low-level actions following a textual instruction in natural language. However, most existing methods do not address the likely case where users may make mistakes when providing such instruction (e.g., "turn left" instead of "turn right"). In this work, we address a novel task of Interactive VLN in Continuous Environments (IVLN-CE), which allows the agent to interact with the user during the VLN-CE navigation to verify any doubts regarding the instruction errors. We propose an Interactive Instruction Error Detector and Localizer (I2EDL) that triggers the user-agent interaction upon the detection of instruction errors during the navigation. We leverage a pre-trained module to detect instruction errors and pinpoint them in the instruction by cross-referencing the textual input and past observations. In such way, the agent is able to query the user for a timely correction, without demanding the user's cognitive load, as we locate the probable errors to a precise part of the instruction. We evaluate the proposed I2EDL on a dataset of instructions containing errors, and further devise a novel metric, the Success weighted by Interaction Number (SIN), to reflect both the navigation performance and the interaction effectiveness. We show how the proposed method can ask focused requests for corrections to the user, which in turn increases the navigation success, while minimizing the interactions.
Francesco Taioli, Stefano Rosa, Alberto Castellini, Lorenzo Natale, Alessio Del Bue, Alessandro Farinelli, Marco Cristani, Yiming Wang 0002
RO-MAN4
2023 Large-scale trialing of the B5G technology for eHealth and Emergency domains
abstract
5G is being deployed and B5G connectivity is under study and standardization. Benefits brought by the 5G/B5G air interface are numerous and 5G is more than just an evolution of radio technology since it consists of innovative concepts: the application of network softwarization and programmability paradigms to the overall network design, the reduced latency promised by edge computing, or the concept of network slicing. These innovations open the door to new vertical-specific services, even capable of saving more lives. The paper describes four use cases to demonstrate the large-scale trialing of the B5G technology specifically devoted to eHealth and Emergency domains, by supporting the B5G applications in large-scale environments (e.g., hospitals) and bringing novel applications (e.g., Remote Proctoring and Smart Ambulance) and on societal benefits in eHealth and Emergency areas through the development of innovative B5G/6G applications. The work is a part of a more complete behavior, TrialsNet project, within SNS JU European Commission Programme, considering other field of application of B5G connectivity.
Andrea Di Giglio, Marco Laurino, Giancarlo Sacco, Gianna Karanasiou, Sergio Berti, Elisa Maiettini, Mara Piccinino, Vera Stavroulaki, Simona Celi, Nicoló Boccardo, Paola Iovanna, Aruna Prem Bianzino, Chiara Benvenuti, Lorenzo Natale, Giulio Bottari
HealthCom14
2023 Collision-aware In-hand 6D Object Pose Estimation using Multiple Vision-based Tactile Sensors
abstract
In this paper, we address the problem of estimating the in-hand 6D pose of an object in contact with multiple vision-based tactile sensors. We reason on the possible spatial configurations of the sensors along the object surface. Specifically, we filter contact hypotheses using geometric reasoning and a Convolutional Neural Network (CNN), trained on simulated object-agnostic images, to promote those that better comply with the actual tactile images from the sensors. We use the selected sensors configurations to optimize over the space of 6D poses using a Gradient Descent-based approach. We finally rank the obtained poses by penalizing those that are in collision with the sensors. We carry out experiments in simulation using the DIGIT vision-based sensor with several objects, from the standard YCB model set. The results demonstrate that our approach estimates object poses that are compatible with actual object-sensor contacts in 87.5% of cases while reaching an average positional error in the order of 2 centimeters. Our analysis also includes qualitative results of experiments with a real DIGIT sensor.
Gabriele M. Caddeo, Nicola A. Piga, Fabrizio Bottarel, Lorenzo Natale
ICRA4
2023 A Grasp Pose is All You Need: Learning Multi-Fingered Grasping with Deep Reinforcement Learning from Vision and Touch
abstract
Multi-fingered robotic hands have potential to enable robots to perform sophisticated manipulation tasks. However, teaching a robot to grasp objects with an anthropomorphic hand is an arduous problem due to the high dimensionality of state and action spaces. Deep Reinforcement Learning (DRL) offers techniques to design control policies for this kind of problems without explicit environment or hand modeling. However, state-of-the-art model-free algorithms have proven inefficient for learning such policies. The main problem is that the exploration of the environment is unfeasible for such high-dimensional problems, thus hampering the initial phases of policy optimization. One possibility to address this is to rely on off-line task demonstrations, but, oftentimes, this is too demanding in terms of time and computational resources. To address these problems, we propose the A Grasp Pose is All You Need (G-PAYN) method for the anthropomorphic hand of the iCub humanoid. We develop an approach to automatically collect task demonstrations to initialize the training of the policy. The proposed grasping pipeline starts from a grasp pose generated by an external algorithm, used to initiate the movement. Then a control policy (previously trained with the proposed G-PAYN) is used to reach and grab the object. We deployed the iCub into the MuJoCo simulator and use it to test our approach with objects from the YCB-Video dataset. Results show that G-PAYN outperforms current DRL techniques in the considered setting in terms of success rate and execution time with respect to the baselines. The code to reproduce the experiments is released together with the paper with an open source license11https://github.com/hsp-iit/rl-icub-dexterous-manipulation.
Federico Ceola, Elisa Maiettini, Lorenzo Rosasco, Lorenzo Natale
IROS4
2023 Hybrid Object Tracking with Events and Frames
abstract
Robust object pose tracking plays an important role in robot manipulation, but it is still an open issue for quickly moving targets as motion blur and low frequency detection can reduce pose estimation accuracy even for state-of-the-art RGB-D-based methods. An event-camera is a low-latency vision sensor that can act complementary to RGB-D. Specifically, its sub-millisecond temporal resolution can be exploited to correct for pose estimation inaccuracies due to low frequency RGB-D based detection. To do so, we propose a dual Kalman filter: the first filter estimates an object's velocity from the spatiotemporal patterns of “events”, the second filter fuses the tracked object velocity with a low-frequency object pose estimated from a deep neural network using RGB-D data. The full system outputs high frequency, accurate object poses also for fast moving objects. The proposed method works towards low-power robotics by replacing high-cost GPU-based optical flow used in prior work with event-cameras that inherently extract the required signal without costly processing. The proposed algorithm achieves comparable or better performance when compared to two state-of-the-art 6-DoF object pose estimation algorithms and one hybrid event/RGB-D algorithm on benchmarks with simulated and real data. We discuss the benefits and tradeoffs for using the event-camera and contribute algorithm, code, and datasets to the community. The code and datasets are available at https://github.com/event-driven-robotics/Hybrid-object-tracking-with-events-and-frames.
Nicola A. Piga, Franco Di Pietro, Massimiliano Iacono, Arren Glover, Lorenzo Natale, Chiara Bartolozzi
IROS6
2022 Grasp Pre-shape Selection by Synthetic Training: Eye-in-hand Shared Control on the Hannes Prosthesis
abstract
We consider the task of object grasping with a prosthetic hand capable of multiple grasp types. In this setting, communicating the intended grasp type often requires a high user cognitive load which can be reduced adopting shared autonomy frameworks. Among these, so-called eye-in-hand systems automatically control the hand pre-shaping before the grasp, based on visual input coming from a camera on the wrist. In this paper, we present an eye-in-hand learning-based approach for hand pre-shape classification from RGB sequences. Differently from previous work, we design the system to support the possibility to grasp each considered object part with a different grasp type. In order to overcome the lack of data of this kind and reduce the need for tedious data collection sessions for training the system, we devise a pipeline for rendering synthetic visual sequences of hand trajectories. We develop a sensorized setup to acquire real human grasping sequences for benchmarking and show that, compared on practical use cases, models trained with our synthetic dataset achieve better generalization performance than models trained on real data. We finally integrate our model on the Hannes prosthetic hand and show its practical effectiveness. We make publicly available the code and dataset to reproduce the presented results11https://github.com/hsp-iit/prosthetic-grasping-simulation.
Federico Vasile, Elisa Maiettini, Giulia Pasquale, Astrid Florio, Nicoló Boccardo, Lorenzo Natale
IROS6
2022 From Handheld to Unconstrained Object Detection: a Weakly-supervised On-line Learning Approach
abstract
Deep Learning (DL) based methods for object detection achieve remarkable performance at the cost of computationally expensive training and extensive data labeling. Robots embodiment can be exploited to mitigate this burden by acquiring automatically annotated training data via a natural interaction with a human showing the object of interest, hand-held. However, learning solely from this data may introduce biases (the so-called domain shift), and prevents adaptation to novel tasks. While Weakly-supervised Learning offers a well-established set of techniques to cope with these problems in general-purpose Computer Vision, its adoption in challenging robotic domains is still at a preliminary stage. In this work, we target the scenario of a robot trained in a teacher-learner setting to detect handheld objects. The aim is to improve detection performance in different settings by letting the robot explore the environment with a limited human labeling budget. We compare several techniques for WSL in detection pipelines to reduce model re-training costs without compromising accuracy, proposing solutions which target the considered robotic scenario. We show that the robot can improve adaptation to novel domains, either by interacting with a human teacher (Active Learning) or with an autonomous supervision (Semi-supervised Learning). We integrate our strategies into an on-line detection method, achieving efficient model update capabilities with few labels. We experimentally benchmark our method on challenging robotic object detection tasks under domain shift1.
Elisa Maiettini, Andrea Maracani, Raffaello Camoriano, Giulia Pasquale, Vadim Tikhanoff, Lorenzo Rosasco, Lorenzo Natale
RO-MAN7
2022 Learn Fast, Segment Well: Fast Object Segmentation Learning on the iCub Robot
abstract
The visual system of a robot has different requirements depending on the application: it may require high accuracy or reliability, be constrained by limited resources, or need fast adaptation to dynamically changing environments. In this article, we focus on the instance segmentation task and provide a comprehensive study of different techniques that allow adapting an object segmentation model in the presence of novel objects or different domains. We propose a pipeline for fast instance segmentation learning designed for robotic applications where data come in stream. It is based on an hybrid method leveraging on a pre-trained convolutional neural network for feature extraction and fast-to-train Kernel-based classifiers. We also propose a training protocol that allows to shorten the training time by performing feature extraction during the data acquisition. We benchmark the proposed pipeline on two robotics datasets and we deploy it on a real robot, i.e., the iCub humanoid. To this aim, we adapt our method to an incremental setting in which novel objects are learned online by the robot. The code to reproduce the experiments is publicly available on GitHub.11[Online]. Available:https://github.com/hsp-iit/online-detection
Federico Ceola, Elisa Maiettini, Giulia Pasquale, Giacomo Meanti, Lorenzo Rosasco, Lorenzo Natale
IEEE Trans. Robotics6
2022 Handling Concurrency in Behavior Trees
abstract
This article addresses the concurrency issues affecting behavior trees (BTs), a popular tool to model the behaviors of autonomous agents in the video game and the robotics industry. BT designers can easily build complex behaviors composing simpler ones, which represents a key advantage of BTs. The parallel composition of BTs expresses a way to combine concurrent behaviors that has high potential, since composing pre-existing BTs in parallel results easier than composing in parallel classical control architectures, as finite state machines or teleo-reactive programs. However, BT designers rarely use such composition due to the underlying concurrency problems similar to the ones faced in concurrent programming. As a result, the parallel composition, despite its potential, finds application only in the composition of simple behaviors or where the designer can guarantee the absence of conflicts by design. In this article, we define two new BT nodes to tackle the concurrency problems in BTs and we show how to exploit them to create predictable behaviors. In addition, we introduce measures to assess execution performance and show how different design choices affect them. We validate our approach in both simulations and the real world. Simulated experiments provide statistically significant data, whereas real-world experiments show the applicability of our method on real robots. We provided an open-source implementation of the novel BT formulation and published all the source code to reproduce the numerical examples and experiments.
Michele Colledanchise, Lorenzo Natale
IEEE Trans. Robotics2
2021 Fast Object Segmentation Learning with Kernel-based Methods for Robotics
abstract
Object segmentation is a key component in the visual system of a robot that performs tasks like grasping and object manipulation, especially in presence of occlusions. Like many other computer vision tasks, the adoption of deep architectures has made available algorithms that perform this task with remarkable performance. However, adoption of such algorithms in robotics is hampered by the fact that training requires large amount of computing time and it cannot be performed on-line.In this work, we propose a novel architecture for object segmentation, that overcomes this problem and provides comparable performance in a fraction of the time required by the state-of-the-art methods. Our approach is based on a pre-trained Mask R-CNN, in which various layers have been replaced with a set of classifiers and regressors that are retrained for a new task. We employ an efficient Kernel-based method that allows for fast training on large scale problems. Our approach is validated on the YCB-Video dataset which is widely adopted in the computer vision and robotics community, demonstrating that we can achieve and even surpass performance of the state-of-the-art, with a significant reduction (~6×) of the training time.The code to reproduce the experiments is publicly available on GitHub1.
Federico Ceola, Elisa Maiettini, Giulia Pasquale, Lorenzo Rosasco, Lorenzo Natale
ICRA5
2021 In Situ Translational Hand-Eye Calibration of Laser Profile Sensors using Arbitrary Objects
abstract
Hand-eye calibration of laser profile sensors is the process of extracting the homogeneous transformation between the laser profile sensor frame and the end-effector frame of a robot in order to express the data extracted by the sensor in the robot’s global coordinate system. For laser profile scanners this is a challenging procedure, as they provide data only in two dimensions and state-of-the-art calibration procedures require the use of specialised calibration targets. This paper presents a novel method to extract the translation-part of the hand-eye calibration matrix with rotation-part known a priori in a target-agnostic way. Our methodology is applicable to any 2D image or 3D object as a calibration target and can also be performed in situ in the final application. The method is experimentally validated on a real robot-sensor setup with 2D and 3D targets.
Prajval Kumar Murali, Ines Sorrentino, Angelo Rendiniello, Claudio Fantacci, Enrico Villagrossi, Andrea Polo, Alessandro Ardesi, Marco Maggiali, Lorenzo Natale, Daniele Pucci, Silvio Traversaro
ICRA9
2021 Score to Learn: A Comparative Analysis of Scoring Functions for Active Learning in Robotics
Riccardo Grigoletto, Elisa Maiettini, Lorenzo Natale
ICVS3
2021 Formalizing the Execution Context of Behavior Trees for Runtime Verification of Deliberative Policies
abstract
In this paper, we enable automated property verification of deliberative components in robot control architectures. We focus on formalizing the execution context of Behavior Trees (BTs) to provide a scalable, yet formally grounded, methodology to enable runtime verification and prevent unexpected robot behaviors. To this end, we consider a message-passing model that accommodates both synchronous and asynchronous composition of parallel components, in which BTs and other components execute and interact according to the communication patterns commonly adopted in robotic software architectures. We introduce a formal property specification language to encode requirements and build runtime monitors. We performed a set of experiments, both on simulations and on the real robot, demonstrating the feasibility of our approach in a realistic application and its integration in a typical robot software architecture. We also provide an OS-level virtualization environment to reproduce the experiments in the simulated scenario.
Michele Colledanchise, Giuseppe Cicala, Daniele Domenichelli, Lorenzo Natale, Armando Tacchella
IROS4
2021 Active Perception for Ambiguous Objects Classification
abstract
Recent visual pose estimation and tracking solutions provide notable results on popular datasets such as T-LESS and YCB. However, in the real world, we can find ambiguous objects that do not allow exact classification and detection from a single view. In this work, we propose a framework that, given a single view of an object, provides the coordinates of a next viewpoint to discriminate the object against similar ones, if any, and eliminates ambiguities. We also describe a complete pipeline from a real object’s scans to the viewpoint selection and classification. We validate our approach with a Franka Emika Panda robot and common household objects featured with ambiguities. We released the source code to reproduce our experiments.
Evgenii Safronov, Nicola A. Piga, Michele Colledanchise, Lorenzo Natale
IROS4
2020 A Flexible Software Architecture for Robotic Industrial Applications
abstract
The paper introduce a robotics software control architecture suitable for the development of complete robotic industrial applications. The architecture fuse the state-of-the-art software technologies in a single standalone platform to provide an easy integration between all the software components necessary to control a robotic application, i.e. PLC logic, robot motion program. The main goal is to provide an architecture as much as possible hardware agnostic to develop easily portable software.
Angelo Rendiniello, Alberto Remus, Ines Sorrentino, Prajval Kumar Murali, Daniele Pucci, Marco Maggiali, Lorenzo Natale, Silvio Traversaro, Enrico Villagrossi, Andrea Polo, Alessandro Ardesi
ETFA7
2020 Act, Perceive, and Plan in Belief Space for Robot Localization
abstract
In this paper, we outline an interleaved acting and planning technique to rapidly reduce the uncertainty of the estimated robot's pose by perceiving relevant information from the environment, as recognizing an object or asking someone for a direction. Generally, existing localization approaches rely on low-level geometric features such as points, lines, and planes. While these approaches provide the desired accuracy, they may require time to converge, especially with incorrect initial guesses. In our approach, a task planner computes a sequence of action and perception tasks to actively obtain relevant information from the robot's perception system. We validate our approach in large state spaces, to show how the approach scales, and in real environments, to show the applicability of our method on real robots. We prove that our approach is sound, probabilistically complete, and tractable in practical cases.
Michele Colledanchise, Damiano Malafronte, Lorenzo Natale
ICRA3
2020 Task Planning with Belief Behavior Trees
abstract
In this paper, we propose Belief Behavior Trees (BBTs), an extension to Behavior Trees (BTs) that allows to automatically create a policy that controls a robot in partially observable environments. We extend the semantic of BTs to account for the uncertainty that affects both the conditions and action nodes of the BT. The tree gets synthesized following a planning strategy for BTs proposed recently: from a set of goal conditions we iteratively select a goal and find the action, or in general the subtree, that satisfies it. Such action may have preconditions that do not hold. For those preconditions, we find an action or subtree in the same fashion. We extend this approach by including, in the planner, actions that have the purpose to reduce the uncertainty that affects the value of a condition node in the BT (for example, turning on the lights to have better lighting conditions). We demonstrate that BBTs allows task planning with non-deterministic outcomes for actions. We provide experimental validation of our approach in a real robotic scenario and - for sake of reproducibility - in a simulated one.
Evgenii Safronov, Michele Colledanchise, Lorenzo Natale
IROS3
2019 Analysis and Exploitation of Synchronized Parallel Executions in Behavior Trees
abstract
Behavior Trees (BTs) are becoming a popular tool to model the behaviors of autonomous agents in the computer game and the robotics industry. One of the key advantages of BTs lies in their composability, where complex behaviors can be built by composing simpler ones. The parallel composition is the one with the highest potential since the complexity of composing pre-existing behaviors in parallel is much lower than the one needed using classical control architectures as finite state machines. However, the parallel composition is rarely used due to the underlying concurrency problems that are similar to the ones faced in concurrent programming. In this paper, we define two synchronization techniques to tackle the concurrency problems in BTs compositions and we show how to exploit them to improve behavior predictability. Also, we introduce measures to assess execution performance and we show how design choices can affect them. To illustrate the proposed framework, we provide a set of experiments using the R1 robot and we gather statistically-significant data.
Michele Colledanchise, Lorenzo Natale
IROS2
2019 Conditional Behavior Trees: Definition, Executability, and Applications
abstract
Behavior Trees (BTs) are gaining acceptance in robotics to specify action policies at the deliberative level. Their advantages include modularity, ease of use and increasing tool support. In this paper, we define Conditional Behavior Trees (CBTs) as an extension of BTs wherein actions are decorated considering pre- and post-conditions. CBTs improve on basic BTs in that they enable monitoring the execution of single actions by checking pre- and post-conditions, respectively. Since there might exist action sequences wherein some preconditions are violated, CBT executability may depend on the success/failure of specific actions. We developed an encoding of CBT executability into satisfiability of propositional formulas to be checked off-line in a publicly-available tool that computes the encoding for generic CBTs. For the kind of application scenarios and related behavior specifications that we consider, we show that our approach is effective and yields formal guarantees about the executability of deliberative policies designed as CBTs.
Eleonora Giunchiglia, Michele Colledanchise, Lorenzo Natale, Armando Tacchella
SMC3
2018 Markerless Visual Servoing on Unknown Objects for Humanoid Robot Platforms
abstract
To precisely reach for an object with a humanoid robot, it is of central importance to have good knowledge of both end-effector, object pose and shape. In this work we propose a framework for markerless visual servoing on unknown objects, which is divided in four main parts: I) a leastsquares minimization problem is formulated to find the volume of the object graspable by the robot's hand using its stereo vision; II) a recursive Bayesian filtering technique, based on Sequential Monte Carlo (SMC) filtering, estimates the 6D pose (position and orientation) of the robot's end-effector without the use of markers; III) a nonlinear constrained optimization problem is formulated to compute the desired graspable pose about the object; IV) an image-based visual servo control commands the robot's end-effector toward the desired pose. We demonstrate effectiveness and robustness of our approach with extensive experiments on the iCub humanoid robot platform, achieving real-time computation, smooth trajectories and subpixel precisions.
Claudio Fantacci, Giulia Vezzani, Ugo Pattacini, Vadim Tikhanoff, Lorenzo Natale
ICRA5
2018 Improving Superquadric Modeling and Grasping with Prior on Object Shapes
abstract
This paper proposes an object modeling and grasping pipeline for humanoid robots. This work improves our previous approach based on superquadric functions. In particular, we speed up and refine the modeling process by using prior information on the object shape provided by an object classifier. We use our previous method for the computation of grasping pose to obtain pose candidates for both the robot hands and, then, we automatically choose the best candidate for grasping the object according to a given quality index. The performance of our pipeline has been assessed on a real robotic system, the iCub humanoid robot. The robot can grasp 18 objects of the YCB and iCub World datasets considerably different in terms of shape and dimensions with a high success rate.
Giulia Vezzani, Ugo Pattacini, Giulia Pasquale, Lorenzo Natale
ICRA4
2018 Improving the Parallel Execution of Behavior Trees
abstract
Behavior Trees (BTs) have become a popular framework for designing controllers of autonomous agents in the computer game and in the robotics industry. One of the key advantages of BTs lies in their modularity, where independent modules can be composed to create more complex ones. In the classical formulation of BTs, modules can be composed using one of the three operators: Sequence, Fallback, and Parallel. The Parallel operator is rarely used despite its strong potential against other control architectures such as Finite State Machines. This is due to the fact that concurrent actions may lead to unexpected problems similar to the ones experienced in concurrent programming. In this paper, we outline how to tackle the aforementioned problem by introducing Concurrent BTs (CBTs) as a generalization of BTs in which we include the notions of progress and resource usage. We show how CBTs allow safe concurrent executions of actions and we analyze the approach from a mathematical standpoint. To illustrate the use of CBTs, we provide a set of use cases in realistic robotics scenarios.
Michele Colledanchise, Lorenzo Natale
IROS2
2018 Speeding-Up Object Detection Training for Robotics with FALKON
abstract
Latest deep learning methods for object detection provide remarkable performance, but have limits when used in robotic applications. One of the most relevant issues is the long training time, which is due to the large size and imbalance of the associated training sets, characterized by few positive and a large number of negative examples (i.e. background). Proposed approaches are based on end-to-end learning by back-propagation [22] or kernel methods trained with Hard Negatives Mining on top of deep features [8]. These solutions are effective, but prohibitively slow for on-line applications. In this paper we propose a novel pipeline for object detection that overcomes this problem and provides comparable performance, with a 60x training speedup. Our pipeline combines (i) the Region Proposal Network and the deep feature extractor from [22] to efficiently select candidate RoIs and encode them into powerful representations, with (ii) the FALKON [23] algorithm, a novel kernel-based method that allows fast training on large scale problems (millions of points). We address the size and imbalance of training data by exploiting the stochastic subsampling intrinsic into the method and a novel, fast, bootstrapping approach. We assess the effectiveness of the approach on a standard Computer Vision dataset (PASCAL VOC 2007 [5]) and demonstrate its applicability to a real robotic scenario with the iCubWorld Transformations [18] dataset.
Elisa Maiettini, Giulia Pasquale, Lorenzo Rosasco, Lorenzo Natale
IROS4
2017 Incremental robot learning of new objects with fixed update time
abstract
We consider object recognition in the context of lifelong learning, where a robotic agent learns to discriminate between a growing number of object classes as it accumulates experience about the environment. We propose an incremental variant of the Regularized Least Squares for Classification (RLSC) algorithm, and exploit its structure to seamlessly add new classes to the learned model. The presented algorithm addresses the problem of having an unbalanced proportion of training examples per class, which occurs when new objects are presented to the system for the first time. We evaluate our algorithm on both a machine learning benchmark dataset and two challenging object recognition tasks in a robotic setting. Empirical evidence shows that our approach achieves comparable or higher classification performance than its batch counterpart when classes are unbalanced, while being significantly faster.
Raffaello Camoriano, Giulia Pasquale, Carlo Ciliberto, Lorenzo Natale, Lorenzo Rosasco, Giorgio Metta
ICRA4
2017 Self-supervised learning of tool affordances from 3D tool representation through parallel SOM mapping
abstract
Future humanoid robots will be expected to carry out a wide range of tasks for which they had not been originally equipped by learning new skills and adapting to their environment. A crucial requirement towards that goal is to be able to take advantage of external elements as tools to perform tasks for which their own manipulators are insufficient; the ability to autonomously learn how to use tools will render robots far more versatile and simpler to design. Motivated by this prospect, this paper proposes and evaluates an approach to allow robots to learn tool affordances based on their 3D geometry. To this end, we apply tool-pose descriptors to represent tools combined with the way in which they are grasped, and affordance vectors to represent the effect tool-poses achieve in function of the action performed. This way, tool affordance learning consists in determining the mapping between these 2 representations, which is achieved in 2 steps. First, the dimensionality of both representations is reduced by unsupervisedly mapping them onto respective Self-Organizing Maps (SOMs). Then, the mapping between the neurons in the tool-pose SOM and the neurons in the affordance SOM for pairs of tool-poses and their corresponding affordance vectors, respectively, is learned with a neural based regression model. This method enables the robot to accurately predict the effect of its actions using tools, and thus to select the best action for a given goal, even with tools not seen on the learning phase.
Tanis Mar, Vadim Tikhanoff, Giorgio Metta, Lorenzo Natale
ICRA4
2017 A grasping approach based on superquadric models
abstract
This paper addresses the problem of grasping unknown objects with a humanoid robot. Conventional approaches fail when the shape, dimension or pose of the objects are missing. We propose a novel approach in which the grasping problem is solved by modeling the object and the volume graspable by the hand with superquadric functions. The object model is computed in real-time using stereo vision. Pose computation is formulated as a nonlinear constrained optimization problem, which is solved in real-time using the Ipopt software package. Notably, our method finds solutions in which the fingers are located on portions of the object that are occluded by vision. The performance of our approach has been assessed on a real robotic system, the iCub humanoid robot. The experiments show that the proposed method computes proper poses, suitable for grasping even small objects, while avoiding hitting the table with the fingers.
Giulia Vezzani, Ugo Pattacini, Lorenzo Natale
ICRA3
2017 Event-driven encoding of off-the-shelf tactile sensors for compression and latency optimisation for robotic skin
abstract
We propose a method to compress the enormous amount of data originating from tactile sensors is presented that explicitly exploits the inherent sparseness over space and time, sending tactile “events” only when a contact is detected. The resulting modular architecture is based on FPGA modules that acquire data samples from off-the-shelf tactile sensors based on capacitive transducers and generate and transmit an event-driven readout. This architecture has been specifically implemented for integration on robots with a large number of tactile sensors, to reduce communication bandwidth, power and processing requirements. An asynchronous serial address-event representation protocol further optimises effective data transmission rate (efficiency of 94.1%) and latency (340 ns) with respect to more common transmission protocols (e.g., Ethernet, CAN). We propose two complementary algorithms for the translation of raw-data into events, optimising data rate and bandwidth, or exploiting the asynchronous nature of the event-driven encoding and the temporal information within the sensory signal. Data reduction capability can reach up to 20 % of the correspondent clock-based encoding, with limited information loss due to the compression.
Chiara Bartolozzi, Paolo Motto Ros, Francesco Diotalevi, Nawid Jamali, Lorenzo Natale, Marco Crepaldi, Danilo Demarchi
IROS5
2017 Visual end-effector tracking using a 3D model-aided particle filter for humanoid robot platforms
abstract
This paper addresses recursive markerless estimation of a robot's end-effector using visual observations from its cameras. The problem is formulated into the Bayesian framework and addressed using Sequential Monte Carlo (SMC) filtering. We use a 3D rendering engine and Computer Aided Design (CAD) schematics of the robot to virtually create images from the robot's camera viewpoints. These images are then used to extract information and estimate the pose of the end-effector. To this aim, we developed a particle filter for estimating the position and orientation of the robot's end-effector using the Histogram of Oriented Gradient (HOG) descriptors to capture robust characteristic features of shapes in both cameras and rendered images. We implemented the algorithm on the iCub humanoid robot and employed it in a closed-loop reaching scenario. We demonstrate that the tracking is robust to clutter, allows compensating for errors in the robot kinematics and servoing the arm in closed loop using vision.
Claudio Fantacci, Ugo Pattacini, Vadim Tikhanoff, Lorenzo Natale
IROS4
2017 A parallel kinematic mechanism for the torso of a humanoid robot: Design, construction and validation
abstract
The torso of a humanoid robot is a fundamental part of its kinematic structure because it defines the reachable workspace, supports the entire upper-body and can be used to control the position of the center of mass. The majority of the torso joints are designed exploiting serial or differential mechanisms, while parallel kinematic structures are less used mainly because of their greater design complexity. This paper describes the design and construction of a 4 degrees of freedom (DoF) torso for our new humanoid robot. Three degrees of freedom, namely roll, pitch and heave, have been implemented using a 3 DoF parallel kinematic structure, while the fourth DoF, namely yaw, has been implemented with a rotational joint on top of the parallel structure. The design has been optimized to reduce the cost and the volume of the system. A first prototype of the torso has been constructed and validated with respect to our design requirements. Eventually, experimental tests have been conducted to assess the functionality of the proposed system.
Luca Fiorio, Alessandro Scalzo, Lorenzo Natale, Giorgio Metta, Alberto Parmiggiani
IROS3
2017 The design and validation of the R1 personal humanoid
abstract
In recent years the robotics field has witnessed an interesting new trend. Several companies started the production of service robots whose aim is to cooperate with humans. The robots developed so far are either rather expensive or unsuitable for manipulation tasks. This article presents the result of a project which wishes to demonstrate the feasibility of an affordable humanoid robot. R1 is able to navigate, and interact with the environment (grasping and carrying objects, operating switches, opening doors etc). The robot is also equipped with a speaker, microphones and it mounts a display in the head to support interaction using natural channels like speech or (simulated) eye movements. The final cost of the robot is expected to range around that of a family car, possibly, when produced in large quantities, even significantly lower. This goal was tackled along three synergistic directions: use of polymeric materials, light-weight design and implementation of novel actuation solutions. These lines, as well as the robot with its main features, are described hereafter.
Alberto Parmiggiani, Luca Fiorio, Alessandro Scalzo, Anand Vazhapilli Sureshbabu, Marco Randazzo, Marco Maggiali, Ugo Pattacini, Hagen Lehmann, Vadim Tikhanoff, Daniele Domenichelli, Alberto Cardellino, Pierpaolo Congiu, Andrea Pagnin, Roberto Cingolani, Lorenzo Natale, Giorgio Metta
IROS15
2017 Memory Unscented Particle Filter for 6-DOF Tactile Localization
abstract
This paper addresses 6-DOF (degree-of-freedom) tactile localization, i.e., the pose estimation of tridimensional objects using tactile measurements. This estimation problem is fundamental for the operation of autonomous robots that are often required to manipulate and grasp objects whose pose is a priori unknown. The nature of tactile measurements, the strict time requirements for real-time operation, and the multimodality of the involved probability distributions pose remarkable challenges and call for advanced nonlinear filtering techniques. Following a Bayesian approach, this paper proposes a novel and effective algorithm, named memory unscented particle filter (MUPF), which solves 6-DOF localization recursively in real time by only exploiting contact point measurements. The MUPF combines a modified particle filter that incorporates a sliding memory of past measurements to better handle multimodal distributions, along with the unscented Kalman filter that moves the particles toward regions of the search space that are more likely with the measurements. The performance of the proposed MUPF algorithm has been assessed both in simulation and on a real robotic system equipped with tactile sensors (i.e., the iCub humanoid robot). The experiments show that the algorithm provides accurate and reliable localization even with a low number of particles and, hence, is compatible with real-time requirements.
Giulia Vezzani, Ugo Pattacini, Giorgio Battistelli, Luigi Chisci, Lorenzo Natale
IEEE Trans. Robotics5
2016 Robustness in view-graph SLAM
Tariq Abuhashim, Lorenzo Natale
FUSION2
2016 Towards automated system and experiment reproduction in robotics
abstract
Even though research on autonomous robots and human-robot interaction accomplished great progress in recent years, and reusable soft- and hardware components are available, many of the reported findings are only hardly reproducible by fellow scientists. Usually, reproducibility is impeded because required information, such as the specification of software versions and their configuration, required data sets, and experiment protocols are not mentioned or referenced in most publications. In order to address these issues, we recently introduced an integrated tool chain and its underlying development process to facilitate reproducibility in robotics. In this contribution we instantiate the complete tool chain in a unique user study in order to assess its applicability and usability. To this end, we chose three different robotic systems from independent institutions and modeled them in our tool chain, including three exemplary experiments. Subsequently, we asked twelve researchers to reproduce one of the formerly unknown systems and the associated experiment. We show that all twelve scientists were able to replicate a formerly unknown robotics experiment using our tool chain.
Florian Lier, Marc Hanheide, Lorenzo Natale, Simon Schulz, Jonathan Weisz, Sven Wachsmuth, Sebastian Wrede 0001
IROS3
2016 Object identification from few examples by improving the invariance of a Deep Convolutional Neural Network
abstract
The development of reliable and robust visual recognition systems is a main challenge towards the deployment of autonomous robotic agents in unconstrained environments. Learning to recognize objects requires image representations that are discriminative to relevant information while being invariant to nuisances, such as scaling, rotations, light and background changes, and so forth. Deep Convolutional Neural Networks can learn such representations from large web-collected image datasets and a natural question is how these systems can be best adapted to the robotics context where little supervision is often available. In this work, we investigate different training strategies for deep architectures on a new dataset collected in a real-world robotic setting. In particular we show how deep networks can be tuned to improve invariance and discriminability properties and perform object identification tasks with minimal supervision.
Giulia Pasquale, Carlo Ciliberto, Lorenzo Rosasco, Lorenzo Natale
IROS4
2015 Learning symbolic representations of actions from human demonstrations
abstract
In this paper, a robot learning approach is pro- posed which integrates Visuospatial Skill Learning, Imitation Learning, and conventional planning methods. In our approach, the sensorimotor skills (i.e., actions) are learned through a learning from demonstration strategy. The sequence of per- formed actions is learned through demonstrations using Visu- ospatial Skill Learning. A standard action-level planner is used to represent a symbolic description of the skill, which allows the system to represent the skill in a discrete, symbolic form. The Visuospatial Skill Learning module identifies the underlying constraints of the task and extracts symbolic predicates (i.e., action preconditions and effects), thereby updating the planner representation while the skills are being learned. Therefore the planner maintains a generalized representation of each skill as a reusable action, which can be planned and performed inde- pendently during the learning phase. Preliminary experimental results on the iCub robot are presented.
Seyed Reza Ahmadzadeh, Ali Paikan, Fulvio Mastrogiovanni, Lorenzo Natale, Petar Kormushev, Darwin G. Caldwell
ICRA4
2015 Self-supervised learning of grasp dependent tool affordances on the iCub Humanoid robot
abstract
The ability to learn about and efficiently use tools constitutes a desirable property for general purpose humanoid robots, as it allows them to extend their capabilities beyond the limitations of their own body. Yet, it is a topic that has only recently been tackled from the robotics community. Most of the studies published so far make use of tool representations that allow their models to generalize the knowledge among similar tools in a very limited way. Moreover, most studies assume that the tool is always grasped in its common or canonical grasp position, thus not considering the influence of the grasp configuration in the outcome of the actions performed with them. In the current paper we present a method that tackles both issues simultaneously by using an extended set of functional features and a novel representation of the effect of the tool use. Together, they implicitly account for the grasping configuration and allow the iCub to generalize among tools based on their geometry. Moreover, learning happens in a self-supervised manner: First, the robot autonomously discovers the affordance categories of the tools by clustering the effect of their usage. These categories are subsequently used as a teaching signal to associate visually obtained functional features to the expected tool's affordance. In the experiments, we show how this technique can be effectively used to select, given a tool, the best action to achieve a desired effect.
Tanis Mar, Vadim Tikhanoff, Giorgio Metta, Lorenzo Natale
ICRA4
2015 A new design of a fingertip for the iCub hand
abstract
Tactile sensing is of fundamental importance for object manipulation and perception. Several sensors for hands have been proposed in the literature, however, only a few of them can be fully integrated with robotic hands. Typical problems preventing integration include the need for deformable sensors that can be deployed on curved surfaces, and wiring complexity. In this paper we describe a fingertip for the hands of the iCub robot, each fingertip consists of 12 sensors. Our approach builds on previous work on the iCub tactile system. The sensing elements of the fingertip are capacitive sensors made from a flexible PCB, and a multi-layer fabric that includes the dielectric material and the conductive layer. The novelty the proposed sensor lies in incorporating the multi-layer fabric technology into a small fingertip sensor that can be attached to the hands of a humanoid robot. The new sensors are more robust. The manufacturing is easier and relies on industrial techniques for the fabrication of the components, which results in higher repeatability. We performed experimental characterization of the sensor. We show that the sensor is able to detect forces as low as 0.05 N with no cross-talk between the taxels. We identified some hysteresis in the response of the sensor which must be taken into account if the robot exerts large forces for a long period of time. The taxels have spatially overlapping receptive fields, this has been demonstrated to be a useful property that allows hyperacuity.
Nawid Jamali, Marco Maggiali, Francesco Giovannini, Giorgio Metta, Lorenzo Natale
IROS5
2015 A best-effort approach for run-time channel prioritization in real-time robotic application
abstract
Application domains of robotic systems are growing in complexity. It seems therefore plausible that robotic software will continue to be designed to be executed on distributed computer architectures interconnected through a network. It is a common practice today to rely on best-effort performance and assume that the latter are adequate given enough computational and networking resources. This approach however does not make best use of the available resources and, maybe more importantly, does not guarantee that performance remain constant over time. Real-time and Quality of Service become therefore important aspects in the software architecture of a robot. This article describes an approach for introducing these concepts in a publish-subscribe software middleware. The key contribution of our approach is that it leverages on the services provided by the operating system (scheduling priority and packet QoS) and abstracts them in a set of levels of priority that can be assigned dynamically, and with the granularity of individual communication channels. We implemented our approach on the YARP middleware and performed an experimental evaluation that demonstrates its benefit for increasing determinism and reducing latency in data communication. We further demonstrate this in a real-robot experiment that shows increased performance in a closed-loop scenario.
Ali Paikan, Ugo Pattacini, Daniele Domenichelli, Marco Randazzo, Giorgio Metta, Lorenzo Natale
IROS6
2015 Tactile Superresolution and Biomimetic Hyperacuity
abstract
Motivated by the impact of superresolution methods for imaging, we undertake a detailed and systematic analysis of localization acuity for a biomimetic fingertip and a flat region of tactile skin. We identify three key factors underlying superresolution that enable the perceptual acuity to surpass the sensor resolution: 1) the sensor is constructed with multiple overlapping, broad but sensitive receptive fields; 2) the tactile perception method interpolates between receptors (taxels) to attain subtaxel acuity; and 3) active perception ensures robustness to unknown initial contact location. All factors follow from active Bayesian perception applied to biomimetic tactile sensors with an elastomeric covering that spreads the contact over multiple taxels. In consequence, we attain extreme superresolution with a 35-fold improvement of localization acuity (0.12 mm) over sensor resolution (4 mm). We envisage that these principles will enable cheap high-acuity tactile sensors that are highly customizable to suit their robotic use. Practical applications encompass any scenario where an end-effector must be placed accurately via the sense of touch.
Nathan F. Lepora, Uriel Martinez-Hernandez, Mathew H. Evans, Lorenzo Natale, Giorgio Metta, Tony J. Prescott
IEEE Trans. Robotics4
2014 Prioritized optimal control
abstract
This paper presents a new technique to control highly redundant mechanical systems, such as humanoid robots. We take inspiration from two approaches. Prioritized control is a widespread multi-task technique in robotics and animation: tasks have strict priorities and they are satisfied only as long as they do not conflict with any higher-priority task. Optimal control instead formulates an optimization problem whose solution is either a feedback control policy or a feedforward trajectory of control inputs. We introduce strict priorities in multi-task optimal control problems, as an alternative to weighting task errors proportionally to their importance. This ensures the respect of the specified priorities, while avoiding numerical conditioning issues. We compared our approach with both prioritized control and optimal control with tests on a simulated robot with 11 degrees of freedom.
Andrea Del Prete, Francesco Romano, Lorenzo Natale, Giorgio Metta, Giulio Sandini, Francesco Nori
ICRA3
2014 Exploiting global force torque measurements for local compliance estimation in tactile arrays
abstract
In this paper we tackle the problem of estimating the local compliance of tactile arrays exploiting global measurements from a single force and torque sensor. The proposed procedure exploits a transformation matrix (describing the relative position between the local tactile elements and the global force/torque measurements) to define a linear regression problem on the unknown local stiffness. Experiments have been conducted on the foot of the iCub robot, sensorized with a single force/torque sensor and a tactile array of 250 tactile elements (taxels) on the foot sole. Results show that a simple calibration procedure can be employed to estimate the stiffness parameters of virtual springs over a tactile array and to use these model to predict normal forces exerted on the array based only on the tactile feedback. Leveraging on previous works [1] the proposed procedure does not necessarily need a-priori information on the transformation matrix of the taxels which can be directly estimated from available measurements.
Carlo Ciliberto, Luca Fiorio, Marco Maggiali, Lorenzo Natale, Lorenzo Rosasco, Giorgio Metta, Giulio Sandini, Francesco Nori
IROS4
2014 Enhancing software module reusability using port plug-ins: An experiment with the iCub robot
abstract
Systematically developing high-quality reusable software components is a difficult task and requires careful design to find a proper balance between potential reuse, functionalities and ease of implementation. Extendibility is an important property for software which helps to reduce cost of development and significantly boosts its reusability. This work introduces an approach to enhance components reusability by extending their functionalities using plug-ins at the level of the connection points (ports). Application-dependent functionalities such as data monitoring and arbitration can be implemented using a conventional scripting language and plugged into the ports of components. The main advantage of our approach is that it avoids to introduce application-dependent modifications to existing components, thus reducing development time and fostering the development of simpler and therefore more reusable components. Another advantage of our approach is that it reduces communication and deployment overheads as extra functionalities can be added without introducing additional modules. The details of the plug-in system is described in the paper and its advantages for the development of robotics applications are demonstrated by developing a step-by-step example on the iCub humanoid robot.
Ali Paikan, Vadim Tikhanoff, Giorgio Metta, Lorenzo Natale
IROS4
2014 An alternative approach to robot safety
abstract
Robotic technology has made significant progresses in the past years. Robots are now common in large manufacturing plants and other industrial settings, safely confined in closed work cells. But to be even more helpful, robots need the capability of interacting physically with humans, and with unstructured environments. This poses new challenges in the design of safe robotic systems. In this article we addressed this problem by proposing a novel design for the joints of the iCub robot. The new design provides the robot with an overload protection mechanism. The overload protection acts as a “passive” torque saturator, which is intrinsically safe. We constructed a prototype of a robotic joint that implements this approach. We first show that our solution is effective in a typical impact scenario. We then evaluate the possible problems arising when the device is controlled with a position control loop. We show that a conventional feedback control loop can trigger positive feedback and instability. Operating the actuator in these conditions is dangerous and can lead to severe failures. We therefore propose the implementation of a relatively simple control strategy that allows to avoid this situation by monitoring slippage, without additional sensors. The quantitative evaluations in the paper demonstrate that our approach is effective and can improve the robustness and safety of complex robotic systems. Indeed these aspects are particularly critical in the case of humaniod robots that are systems prone to severe whole-body impacts in unstructured environments (e.g. falling).
Alberto Parmiggiani, Marco Randazzo, Lorenzo Natale, Giorgio Metta
IROS3
2014 Partial force control of constrained floating-base robots
abstract
Legged robots are typically in rigid contact with the environment at multiple locations, which add a degree of complexity to their control. We present a method to control the motion and a subset of the contact forces of a floating-base robot. We derive a new formulation of the lexicographic optimization problem typically arising in multi-task motion/force control frameworks. The structure of the constraints of the problem (i.e. the dynamics of the robot) allows us to find a sparse analytical solution. This leads to an equivalent optimization with reduced computational complexity, comparable to inverse-dynamics based approaches. At the same time, our method preserves the flexibility of optimization based control frameworks. Simulations were carried out to achieve different multi-contact behaviors on a 23-degree-of-freedom humanoid robot, validating the presented approach. A comparison with another state-of-the-art control technique with similar computational complexity shows the benefits of our controller, which can eliminate force/torque discontinuities.
Andrea Del Prete, Nicolas Mansard, Francesco Nori, Giorgio Metta, Lorenzo Natale
IROS5
2014 Developmental Perception of the Self and Action
abstract
This paper describes a developmental framework for action-driven perception in anthropomorphic robots. The key idea of the framework is that action generation develops the agent's perception of its own body and actions. Action-driven development is critical for identifying changing body parts and understanding the effects of actions in unknown or nonstationary environments. We embedded minimal knowledge into the robot's cognitive system in the form of motor synergies and actions to allow motor exploration. The robot voluntarily generates actions and develops the ability to perceive its own body and the effect that it generates on the environment. The robot, in addition, can compose this kind of learned primitives to perform complex actions and characterize them in terms of their sensory effects. After learning, the robot can recognize manipulative human behaviors with cross-modal anticipation for recovery of unavailable sensory modality, and reproduce the recognized actions afterward. We evaluated the proposed framework in the experiments with a real robot. In the experiments, we achieved autonomous body identification, learning of fixation, reaching and grasping actions, and developmental recognition of human actions as well as their reproduction.
Ryo Saegusa, Giorgio Metta, Giulio Sandini, Lorenzo Natale
IEEE Trans. Neural Networks Learn. Syst.4
2013 Active contour following to explore object shape with robot touch
abstract
In this work, we present an active tactile perception approach for contour following based on a probabilistic framework. Tactile data were collected using a biomimetic fingertip sensor. We propose a control architecture that implements a perception-action cycle for the exploratory procedure, which allows the fingertip to react to tactile contact whilst regulating the applied contact force. In addition' the fingertip is actively repositioned to an optimal position to ensure accurate perception. The method is trained off-line and then the testing performed on-line based on contour following around several different test shapes. We then implement object recognition based on the extracted shapes. Our active approach is compared with a passive approach, demonstrating that active perception is necessary for successful contour following and hence shape recognition.
Uriel Martinez-Hernandez, Giorgio Metta, Tony J. Dodd, Tony J. Prescott, Lorenzo Natale, Nathan F. Lepora
World Haptics5
2013 Perception during interaction is not based on statistical context
Alessandra Sciutti, Andrea Del Prete, Lorenzo Natale, David Burr, Giulio Sandini, Monica Gori
HRI3
2013 Weakly supervised strategies for natural object recognition in robotics
abstract
The paper aims at building a computer vision system for automatic image labeling in robotics scenarios. We show that the weak supervision provided by a human demonstrator, through the exploitation of the independent motion, enables a realistic Human-Robot Interaction (HRI) and achieves an automatic image labeling. We start by reviewing the underlying principles of our previous method for egomotion compensation [1], then we extend our approach removing the dependency on a known kinematics in order to provide a general method for a wide range of devices. From sparse salient features we predict the egomotion of the camera through a heteroscedastic learning method. Subsequently we use an object recognition framework for testing the automatic image labeling process: we rely on the State of the Art method from Yang et al. [2], employing local features augmented through a sparse coding stage and classified with linear SVMs. The application has been implemented and validated on the iCub humanoid robot and experiments are presented to show the effectiveness of the proposed approach. The contribution of the paper is twofold: first we overcome the dependency on the kinematics in the independent motion detection method, secondly we present a practical application for automatic image labeling through a natural HRI.
Sean Ryan Fanello, Carlo Ciliberto, Lorenzo Natale, Giorgio Metta
ICRA3
2013 Developmental action perception for manipulative interaction
abstract
The paper describes a developmental framework of action-driven perception in anthropomorphic robots. The key idea of the framework is that action develops the agent's perception of the own body and its action. In this framework, a robot voluntarily generates movements, and then develops the ability to perceive its own body and the effects of action primitives. The robot, moreover, demonstrates manipulative actions composed of the learned primitives, and characterizes the actions based on their sensory effects. After learning, the robot can predictively recognize humans' manipulative actions with cross-modal recovery of unavailable sensory information and reproduce the recognized actions. We evaluated the proposed framework in experiments with a real robot. In the experiments, we achieved developmental recognition of human actions as well as their reproduction.
Ryo Saegusa, Giorgio Metta, Giulio Sandini, Lorenzo Natale
ICRA4
2013 On the impact of learning hierarchical representations for visual recognition in robotics
abstract
Recent developments in learning sophisticated, hierarchical image representations have led to remarkable progress in the context of visual recognition. While these methods are becoming standard in modern computer vision systems, they are rarely adopted in robotics. The question arises of whether solutions, which have been primarily developed for image retrieval, can perform well in more dynamic and unstructured scenarios. In this paper we tackle this question performing an extensive evaluation of state of the art methods for visual recognition on a iCub robot. We consider the problem of classifying 15 different objects shown by a human demonstrator in a challenging Human-Robot Interaction scenario. The classification performance of hierarchical learning approaches are shown to outperform benchmark solutions based on local descriptors and template matching. Our results show that hierarchical learning systems are computationally efficient and can be used for real-time training and recognition of objects.
Carlo Ciliberto, Sean Ryan Fanello, Matteo Santoro, Lorenzo Natale, Giorgio Metta, Lorenzo Rosasco
IROS4
2012 Advances in tactile sensing and touch based human-robot interaction
abstract
The problem of "providing robots with the sense of touch" is fundamental in order to develop the next generations of robots capable of interacting with humans in different contexts: in daily housekeeping activities, as working partners or as caregivers, just to name a few.
Giorgio Cannata, Fulvio Mastrogiovanni, Giorgio Metta, Lorenzo Natale
HRI4
2012 Imitation learning of non-linear point-to-point robot motions using dirichlet processes
abstract
In this paper we discuss the use of the infinite Gaussian mixture model and Dirichlet processes for learning robot movements from demonstrations. Starting point of this work is an earlier paper where the authors learn a non-linear dynamic robot movement model from a small number of observations. The model in that work is learned using a classical finite Gaussian mixture model (FGMM) where the Gaussian mixtures are appropriately constrained. The problem with this approach is that one needs to make a good guess for how many mixtures the FGMM should use. In this work, we generalize this approach to use an infinite Gaussian mixture model (IGMM) which does not have this limitation. Instead, the IGMM automatically finds the number of mixtures that are necessary to reflect the data complexity. For use in the context of a non-linear dynamic model, we develop a Constrained IGMM (CIGMM). We validate our algorithm on the same data that was used in [5], where the authors use motion capture devices to record the demonstrations. As further validation we test our approach on novel data acquired on our iCub in a different demonstration scenario in which the robot is physically driven by the human demonstrator.
Volker Krüger, Vadim Tikhanoff, Lorenzo Natale, Giulio Sandini
ICRA3
2012 A heteroscedastic approach to independent motion detection for actuated visual sensors
abstract
We present an original method for independent motion detection in dynamic scenes. The algorithm is designed for robotics real-time applications and it overcomes the short-comings of current approaches for the egomotion estimation in presence of many outliers, occlusions and cluttered background. The method relies on a stereo system which performs the reprojection of a sparse set of features following the camera displacement. We assume that noisy prior knowledge of the motion is available (i.e. a robot's kinematic model). Since this estimation leads to a heteroscedastic regression problem due to input-dependent noise, we employ a simple, but computationally efficient approach in order to accurately determine the latent egomotion subspace spanned by the Degrees of Freedom (DOFs) of the robot. The algorithm has been implemented and validated on the iCub humanoid robot. Qualitative and quantitative experiments are presented to show the effectiveness of the proposed approach. The contribution of the paper is a modular framework for independent motion detection naturally extendable to any architecture featuring a visual sensor that can be directly controllable.
Carlo Ciliberto, Sean Ryan Fanello, Lorenzo Natale, Giorgio Metta
IROS3
2012 Interactive online learning of the kinematic workspace of a humanoid robot
abstract
We describe an interactive learning strategy that enables a humanoid robot to build a representation of its workspace: we call it a Reachable Space Map. The robot learns this map autonomously and online during the execution of goal-directed reaching movements; reaching control is based on kinematic models that are learned online as well. The map can be used to estimate the reachability of a fixated object and to plan preparatory movements (e.g. bending or rotating the waist) that improve the effectiveness of the subsequent reaching action. Three main concepts make our solution innovative with respect to previous works: the use of a gaze-centered motor representation to describe the robot workspace, the primary role of action in building and representing knowledge (i.e. interactive learning), the realization of autonomous online learning. We evaluate our strategy by learning the workspace of a simulated humanoid robot and we show how this knowledge can be exploited to plan and execute complex actions, like whole-body bimanual reaching.
Lorenzo Jamone, Lorenzo Natale, Giulio Sandini, Atsuo Takanishi
IROS2
2012 Control of contact forces: The role of tactile feedback for contact localization
abstract
This paper investigates the role of precise estimation of contact points in force control. This analysis is motivated by scenarios in which robots make contacts, either voluntarily or accidentally, with different parts of their body. Control paradigms that are usually implemented in robots with no tactile system, make the hypothesis that contacts occur at the end-effectors only. In this paper we try to investigate what happens when this assumption is not verified. First we consider a simple feedforward force control law, and then we extend it by introducing a proportional feedback term. For both controllers we find the error in the resulting contact force, that is induced by a hypothetic error in the estimation of the contact point. We show that, depending on the geometry of the contact, incorrect estimation of contact points can induce undesired joint accelerations. We validate the presented analysis with tests on a simulated robot arm. Moreover we consider a complex real world scenario, where most of the assumptions that we make in our analytical derivation do not hold. Through tests on the iCub humanoid robot we see how errors in contact localization affect the performance of a parallel force/position controller. In order to estimate contact points and contact forces on the forearm of the iCub we do not use any model of the environment, but we exploit its 6-axis force/torque sensor and its sensorized skin.
Andrea Del Prete, Francesco Nori, Giorgio Metta, Lorenzo Natale
IROS4
2011 Active perception for action mirroring
abstract
The paper describes a constructive approach on active perception for anthropomorphic robots. The key idea is that a robot tries to identify a human's action as an own action based on the observation of action effects for objects. In the proposed framework, the active perception is decomposed into the three phases; First, a robot voluntarily generates actions to discover the own body and objects. Second, the robot characterizes its own action with the effect for the objects. Third, the robot identifies the human action with the own action. The mirrored perception of the own action and the human's action allows the robot to share the goal-directed behavior with humans. The proposed framework of active perception was experimentally validated with the integrated sensory modalities of vision, proprioception and touch.
Ryo Saegusa, Lorenzo Natale, Giorgio Metta, Giulio Sandini
IJCNN2
2011 Reexamining Lucas-Kanade method for real-time independent motion detection: Application to the iCub humanoid robot
abstract
Visual motion is a simple yet powerful cue widely used by biological systems to improve their perception and adaptation to the environment. Examples of tasks that greatly benefit from the ability to detect movement are object segmentation, 3D scene reconstruction and control of attention. In computer vision several algorithms for computing visual motion and optic flow exist. However their application in robotics is not straightforward as in these platforms visual motion is often dominated by (self) motion produced by the movement of the robot (egomotion) making it difficult to disambiguate between motion induced by the scene dynamics or by the own actions of the robot. Independent motion detection is an active field in computer vision and robotics, however approaches in this area typically require that some models of both the environment and the robot visual system are available and are hardly suitable for real-time control. In this paper we describe the motionCUT, a derivation of the Lucas-Kanade optical flow algorithm that allows detecting moving objects, irrespectively of the egomotion produced by the robot. Our method is purely visual and does not require information other than the images coming from the cameras. As such it can be easily adapted to any robotic platform. The system was tested on a stereo tracking task on the iCub humanoid robot, demonstrating that the algorithm performs well and can easily execute in real-time.
Carlo Ciliberto, Ugo Pattacini, Lorenzo Natale, Francesco Nori, Giorgio Metta
IROS3
2011 Online multiple instance learning applied to hand detection in a humanoid robot
abstract
We propose an algorithm for the visual detection and localisation of the hand of a humanoid robot. This algorithm imposes low requirements on the type of supervision required to achieve good performance. In particular the system performs feature selection and adaptation using images that are only labelled as containing the hand or not, without any explicit segmentation. Our algorithm is an online variant of Multiple Instance Learning based on boosting. Experiments in real-world conditions on the iCub humanoid robot confirm that the algorithm can learn the visual appearance of the hand, reaching an accuracy comparable with its off-line version. This remains true when supervision is generated by the robot itself in a completely autonomous fashion. Algorithms with weak supervision requirements like the one we describe are useful for autonomous robots that learn and adapt online to a changing environment. The algorithm is not hand-specific and could be easily applied to wide range of problems involving visual recognition of generic objects.
Carlo Ciliberto, Fabrizio Smeraldi, Lorenzo Natale, Giorgio Metta
IROS3
2011 Towards a platform-independent cooperative human-robot interaction system: II. Perception, execution and imitation of goal directed actions
abstract
If robots are to cooperate with humans in an increasingly human-like manner, then significant progress must be made in their abilities to observe and learn to perform novel goal directed actions in a flexible and adaptive manner. The current research addresses this challenge. In CHRIS.I [1], we developed a platform-independent perceptual system that learns from observation to recognize human actions in a way which abstracted from the specifics of the robotic platform, learning actions including “put X on Y” and “take X”. In the current research, we extend this system from action perception to execution, consistent with current developmental research in human understanding of goal directed action and teleological reasoning. We demonstrate the platform independence with experiments on three different robots. In Experiments 1 and 2 we complete our previous study of perception of actions “put” and “take” demonstrating how the system learns to execute these same actions, along with new related actions “cover” and “uncover” based on the composition of action primitives “grasp X” and “release X at Y”. Significantly, these compositional action execution specifications learned on one iCub robot are then executed on another, based on the abstraction layer of motor primitives. Experiment 3 further validates the platform-independence of the system, as a new action that is learned on the iCub in Lyon is then executed on the Jido robot in Toulouse. In Experiment 4 we extended the definition of action perception to include the notion of agency, again inspired by developmental studies of agency attribution, exploiting the Kinect motion capture system for tracking human motion. Finally in Experiment 5 we demonstrate how the combined representation of action in terms of perception and execution provides the basis for imitation. This provides the basis for an open ended cooperation capability where new actions can be learned and integrated into shared plans for cooperation. Part of the novelty of this research is the robots' use of spoken language understanding and visual perception to generate action representations in a platform independent manner based on physical state changes. This provides a flexible capability for goal-directed action imitation.
Stéphane Lallée, Ugo Pattacini, Jean-David Boucher, Séverin Lemaignan, Alexander Lenz, Chris Melhuish, Lorenzo Natale, Sergey Skachek, Katharina Hamann, Jasmin Steinwender, Akin Sisbot, Giorgio Metta, Rachid Alami 0001, Matthieu Warnier, Julien Guitton, Felix Warneken, Peter Ford Dominey
IROS7
2011 Skin spatial calibration using force/torque measurements
abstract
This paper deals with the problem of estimating the position of tactile elements (i.e. taxels) that are mounted on a robot body part. This problem arises with the adoption of tactile systems with a large number of sensors, and it is particularly critical in those cases in which the system is made of flexible material that is deployed on a curved surface. In this scenario the location of each taxel is partially unknown and difficult to determine manually. Placing the device is in fact an inaccurate procedure that is affected by displacements in both position and orientation. Our approach is based on the idea that it is possible to automatically infer the position of the taxels by measuring the interaction forces exchanged between the sensorized part and the environment. The location of the contact is estimated through force/torque (F/T) measures gathered by a sensor mounted on the kinematic chain of the robot. Our method requires few hypotheses and can be effectively implemented on a real platform, as demonstrated by the experiments with the iCub humanoid robot.
Andrea Del Prete, Simone Denei, Lorenzo Natale, Fulvio Mastrogiovanni, Francesco Nori, Giorgio Cannata, Giorgio Metta
IROS3
2011 A comparison between joint level torque sensing and proximal F/T sensor torque estimation: Implementation on the iCub
abstract
When a robot is required to safely interact with a physical environment, two approaches are typically reported in literature: using a force/torque sensor to regulate the interaction forces at the end effector, or integrating sensors in each robot joint to regulate their torques. In this paper we want to discuss the benefits and the disadvantages of the two approaches, showing a direct comparison between the information which can be obtained from the two categories of sensors. Results obtained on the new iCub arm, which integrates torque sensing capabilities at joint level will be presented and discussed.
Marco Randazzo, Matteo Fumagalli 0001, Francesco Nori, Lorenzo Natale, Giorgio Metta, Giulio Sandini
IROS4
2011 Force Control and Reaching Movements on the iCub Humanoid Robot
Giorgio Metta, Lorenzo Natale, Francesco Nori, Giulio Sandini
ISRR2
2011 Methods and Technologies for the Implementation of Large-Scale Robot Tactile Sensors
abstract
Even though the sense of touch is crucial for humans, most humanoid robots lack tactile sensing. While a large number of sensing technologies exist, it is not trivial to incorporate them into a robot. We have developed a compliant “skin” for humanoids that integrates a distributed pressure sensor based on capacitive technology. The skin is modular and can be deployed on nonflat surfaces. Each module scans locally a limited number of tactile-sensing elements and sends the data through a serial bus. This is a critical advantage as it reduces the number of wires. The resulting system is compact and has been successfully integrated into three different humanoid robots. We have performed tests that show that the sensor has favorable characteristics and implemented algorithms to compensate the hysteresis and drift of the sensor. Experiments with the humanoid robot iCub prove that the sensors can be used to grasp unmodeled, fragile objects.
Alexander Schmitz, Perla Maiolino, Marco Maggiali, Lorenzo Natale, Giorgio Cannata, Giorgio Metta
IEEE Trans. Robotics4
2010 Machine-learning based control of a human-like tendon-driven neck
abstract
This paper describes the control of a human-like robotic neck actuated with tendons. The controller regulates the length of the tendons to achieve a desired orientation of the neck and at the same time it maintains the tension of the tendons within certain limits. The solution we propose does not use any model of the system, but it relies on online learning of the different Jacobian mappings required by the controller. Learning, data acquisition and control are simultaneous; thus learning is completely autonomous, and purely online. We show that after enough iterations the controller produces straight trajectories in the task space and is able to maintain the tension of the tendons within safe limits.
Lorenzo Jamone, Matteo Fumagalli 0001, Giorgio Metta, Lorenzo Natale, Francesco Nori, Giulio Sandini
ICRA4
2010 Safe and effective learning: A case study
abstract
In this paper we consider the problem of ensuring that a multi-agent robot control system is both safe and effective in the presence of learning components. Safety, i.e., proving that a potentially dangerous configuration is never reached in the control system, usually competes with effectiveness, i.e., ensuring that tasks are performed at an acceptable level of quality. In particular, we focus on a robot playing the air hockey game against a human opponent, where the robot has to learn how to minimize opponent's goals (defense play). This setup is paradigmatic since the robot must see, decide and move fastly, but, at the same time, it must learn and guarantee that the control system is safe throughout the process. We attack this problem using automata-theoretic formalisms and associated verification tools, showing experimentally that our approach can yield safety without heavily compromising effectiveness.
Giorgio Metta, Lorenzo Natale, Shashank Pathak, Luca Pulina, Armando Tacchella
ICRA2
2010 Safe Learning with Real-Time Constraints: A Case Study
Giorgio Metta, Lorenzo Natale, Shashank Pathak, Luca Pulina, Armando Tacchella
IEA/AIE (1)2
2010 Exploiting proximal F/T measurements for the iCub active compliance
abstract
During the last decades, interaction (with humans and with the environment) has become an increasingly interesting topic of research within the field of robotics. At the basis of interaction, a fundamental role is played by the ability to actively regulate the interaction forces. In this paper we propose a technique for controlling the interaction forces exploiting a proximal six axes force/torque sensor. The major assumption is the knowledge of the point where external forces are applied. The proposed approach is tested and validated on the four limbs of the iCub, a humanoid robot designed for research in embodied cognition. Remarkably, the proposed approach can be used to implement active compliance in other non passively back-drivable manipulators by simply inserting one or more force/torque sensor anywhere along the kinematic chain.
Matteo Fumagalli 0001, Marco Randazzo, Francesco Nori, Lorenzo Natale, Giorgio Metta, Giulio Sandini
IROS4
2010 Towards a platform-independent cooperative human-robot interaction system: I. Perception
abstract
One of the long term objectives of robotics and artificial cognitive systems is that robots will increasingly be capable of interacting in a cooperative and adaptive manner with their human counterparts in open-ended tasks that can change in real-time. In such situations, an important aspect of the robot behavior will be the ability to acquire new knowledge of the cooperative tasks by observing humans. At least two significant challenges can be identified in this context. The first challenge concerns development of methods to allow the characterization of human actions such that robotic systems can observe and learn new actions, and more complex behaviors made up of those actions. The second challenge is associated with the immense heterogeneity and diversity of robots and their perceptual and motor systems. The associated question is whether the identified methods for action perception can be generalized across the different perceptual systems inherent to distinct robot platforms. The current research addresses these two challenges. We present results from a cooperative human-robot interaction system that has been specifically developed for portability between different humanoid platforms. Within this architecture, the physical details of the perceptual system (e.g. video camera vs IR video with reflecting markers) are encapsulated at the lowest level. Actions are then automatically characterized in terms of perceptual primitives related to motion, contact and visibility. The resulting system is demonstrated to perform robust object and action learning and recognition on two distinct robotic platforms. Perhaps most interestingly, we demonstrate that knowledge acquired about action recognition with one robot can be directly imported and successfully used on a second distinct robot platform for action recognition. This will have interesting implications for the accumulation of shared knowledge between distinct heterogeneous robotic systems.
Stéphane Lallée, Séverin Lemaignan, Alexander Lenz, Chris Melhuish, Lorenzo Natale, Sergey Skachek, Tijn van der Zant, Felix Warneken, Peter Ford Dominey
IROS5
2010 An experimental evaluation of a novel minimum-jerk cartesian controller for humanoid robots
abstract
In this paper we describe the design of a Cartesian Controller for a generic robot manipulator. We address some of the challenges that are typically encountered in the field of humanoid robotics. The solution we propose deals with a large number of degrees of freedom, produce smooth, human-like motion and is able to compute the trajectory on-line. In this paper we support the idea that to produce significant advancements in the field of robotics it is important to compare different approaches not only at the theoretical level but also at the implementation level. For this reason we test our software on the iCub platform and compare its performance against other available solutions.
Ugo Pattacini, Francesco Nori, Lorenzo Natale, Giorgio Metta, Giulio Sandini
IROS3
2010 A tactile sensor for the fingertips of the humanoid robot iCub
abstract
In order to successfully perform object manipulation, humanoid robots must be equipped with tactile sensors. However, the limited space that is available in robotic fingers imposes severe design constraints. In [1] we presented a small prototype fingertip which incorporates a capacitive pressure system. This paper shows an improved version, which has been integrated on the hand of the humanoid robot iCub. The fingertip is 14.5 mm long and 13 mm wide. The capacitive pressure sensor system has 12 sensitive zones and includes the electronics to send the 12 measurements over a serial bus with only 4 wires. Each synthetic fingertip is shaped approximately like a human fingertip. Furthermore, an integral part of the capacitive sensor is soft silicone foam, and therefore the fingertip is compliant. We describe the structure of the fingertip, their integration on the humanoid robot iCub and present test results to show the characteristics of the sensor.
Alexander Schmitz, Marco Maggiali, Lorenzo Natale, Bruno Bonino, Giorgio Metta
IROS3
2010 Touch sensors for humanoid hands
abstract
The sense of touch is of major importance for object handling. Nevertheless, adequate cutaneous sensors for humanoid robot hands are still missing. Designing such sensors is challenging, because they should not only give reliable measurements and integrate many sensing points into little space, but they should also be compliant and should not obstruct the other functions of the robot. This paper presents a capacitive pressure sensor system with 108 sensitive zones for the hands of the humanoid robot iCub. In particular, the palm has 48 taxels and each of the five fingertips has 12 taxels. The size and the shape of the hand are similar to that of a human child. When designing the sensors, we paid special attention to the integration on the robot. Also the ease and speed of production was an important design factor. Furthermore, the sensor incorporates silicone foam and is therefore compliant. We show the working principle of the sensor, how it has been integrated into the hands, and describe experiments that have been performed to show the characteristics of the sensor.
Alexander Schmitz, Marco Maggiali, Lorenzo Natale, Giorgio Metta
RO-MAN3
2010 The iCub humanoid robot: An open-systems platform for research in cognitive development
Giorgio Metta, Lorenzo Natale, Francesco Nori, Giulio Sandini, David Vernon, Luciano Fadiga, Claes von Hofsten, Kerstin Rosander, Manuel Lopes 0001, José Santos-Victor, Alexandre Bernardino, Luis Montesano
Neural Networks2
2007 Autonomous learning of 3D reaching in a humanoid robot
abstract
In this paper, we describe the implementation of a precise reaching controller on an upper-torso humanoid robot. The solution we propose does not rely on prior models of the kinematic structure of either the arm or the head. A learning strategy enables the robot to acquire the required sensory-motor transformations. After learning the robot is able to precisely reach for a visually identified object in the 3-dimensional space. In this technique we use the fixation point (represented in the head joints motor space) as a reference frame to code the position of the object and to represent the eye-to-hand Jacobian matrix. This strategy successfully deals with the kinematic redundancy of the structure and constraints the dimensionality of the problem.
Francesco Nori, Lorenzo Natale, Giulio Sandini, Giorgio Metta
IROS2
2003 Learning about objects through action -initial steps towards artificial cognition
abstract
Within the field of Neuro Robotics we are driven primarily by the desire to understand how humans and animals live and grow and solve every day's problems. To this aim we adopted a "learn by doing" approach by building artificial systems, e.g. robots that not only look like human beings but also represent a model of some brain process. They should, ideally, behave and interact like human beings (being situated). The main emphasis in robotics has been on systems that act as a reaction to an external stimulus (e.g. tracking, reaching), rather than as a result of an internal drive to explore or "understand" the environment. We think it is now appropriate to try to move from acting, in the sense explained above, to "understanding". As a starting point we addressed the problem of learning about the effects and consequences of self-generated actions. How does the robot learn how to pull an object toward itself or to push it away? How does the robot learn that spherical objects roll while a cube only slides if pushed? Interacting with objects is important because it implicitly explores object representation, event understanding, and can provide definition of object-hood that could not be grasped with a mere passive observation of the world. Further, learning to understand what one's own body can do is an essential step toward learning by imitation. In this view two actions are similar not only if their kinematics and dynamics are similar but rather if the effects on the external world are the same. Along this line of research we discuss some recent experiments performed at the AI-Lab at MIT and at the LIRA-Lab at the University of Genova on COG and Babybot respectively. We show how the humanoid robots can learn how to poke and prod objects to obtain a consistently repeatable effect (e.g. sliding in a given direction), to help visual segmentation, and to interpret a poking action performed by a human manipulator.
Paul M. Fitzpatrick, Giorgio Metta, Lorenzo Natale, Ajit Rao, Giulio Sandini
ICRA3