VLDB 2026 Research / reviewers in the wild / expert
Alberto Sanfeliu
dblp:10/4340
· DBLP profile ↗
126ranked-venue papers
12as first author
21since 2021 · last 2025
0000-0003-3868-9678ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 110 · 9 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 42 · 6 first-author · 1 since 2021Systems, architecture and hardware · 41 · 8 since 2021Human-computer interaction and ubiquitous computing · 15 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 10 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Personalised Explainable Robots Using LLMsabstractIn the field of Human-Robot Interaction (HRI), a key challenge lies in enabling humans to comprehend the decisions and behaviours of robots. One promising approach involves leveraging Theory of Mind (ToM) frameworks, wherein a robot estimates the mental model that a user holds about its functioning and compares this with the representation of its internal mental model. This comparison allows the robot to identify potential mismatches and generate communicative actions to bridge such gaps. Effective communication requires the robot to maintain unique mental models for each user and personalise explanations based on past interactions. To address this, we propose an architecture grounded in Large Language Models (LLMs) that operationalises this theoretical framework. We demonstrate the feasibility of this approach through qualitative examples, showcasing responses provided by a robot patrolling a geriatric hospital. Ferran Gebellí, Lavinia Hriscu, Raquel Ros, Séverin Lemaignan, Alberto Sanfeliu, Anais Garrell |
HRI | 5 |
| 2025 | Enhancing Context-Aware Human Motion Prediction for Efficient Robot HandoversabstractAccurate human motion prediction (HMP) is critical for seamless human-robot collaboration, particularly in handover tasks that require real-time adaptability. Despite the high accuracy of state-of-the-art models, their computational complexity limits practical deployment in real-world robotic applications. In this work, we enhance human motion forecasting for handover tasks by leveraging siMLPe [1], a lightweight yet powerful architecture, and introducing key improvements. Our approach, named IntentMotion incorporates intention-aware conditioning, task-specific loss functions, and a novel intention classifier, significantly improving motion prediction accuracy while maintaining efficiency. Experimental results demonstrate that our method reduces body loss error by over 50%, achieves 200× faster inference, and requires only 3% of the parameters compared to existing state-of-the-art HMP models in robotics. These advancements establish our framework as a highly efficient and scalable solution for real-time human-robot interaction. Gerard Gómez-Izquierdo, Javier Laplaza, Alberto Sanfeliu, Anais Garrell |
IROS | 3 |
| 2025 | Negotiation of Assignation Plans in Human-Robot Team Task SchedulingabstractIn recent years, considerable attention has been given to improving human-robot collaboration. Despite advances in robotic capabilities and interaction techniques, achieving a fair distribution of tasks remains challenging due to the dynamic nature of human preferences and situational constraints. This paper presents a novel negotiation framework that enables robots to effectively communicate with humans to facilitate fair and adaptive task allocation. Our approach leverages automated planning techniques with the Planning Domain Definition Language (PDDL), explicitly encoding tasks, constraints, and preferences from both human and robotic perspectives. Task allocation is optimized based on three key criteria: the robot’s effort, the human’s effort, and overall task success. Additionally, we integrate a Natural Language Processing (NLP) model that interprets human preferences and informs the negotiation process, ensuring that the robot generates task proposals aligned with human input. The negotiation follows an alternating-offer protocol, with the robot employing a sigmoid conceder strategy to iteratively refine task allocation, leading to balanced and mutually acceptable plans. To evaluate our approach, we conduct a comprehensive user study with non-trained volunteers interacting with the robot, assessing the effectiveness, fairness, and adaptability of the proposed system in real-world scenarios. Llum Fuster-Palà, Marc Dalmasso, Artur Aubach-Altes, Silvia Izquierdo-Badiola, Alberto Sanfeliu, Anais Garrell |
RO-MAN | 5 |
| 2025 | AI or Human? Understanding Perceptions of Embodied Robots with LLMsabstractThe pursuit of artificial intelligence has long been associated to the the challenge of effectively measuring intelligence. Even if the Turing Test was introduced as a means of assessing a system’s intelligence, its relevance and application within the field of human-robot interaction remain largely underexplored. This study investigates the perception of intelligence in embodied robots by performing a Turing Test within a robotic platform. A total of 34 participants were tasked with distinguishing between AI- and human-operated robots while engaging in two interactive tasks: an information retrieval and a package handover. These tasks assessed the robot’s perception and navigation abilities under both static and dynamic conditions. Results indicate that participants were unable to reliably differentiate between AI- and human-controlled robots beyond chance levels. Furthermore, analysis of participant responses reveals key factors influencing the perception of artificial versus human intelligence in embodied robotic systems. These findings provide insights into the design of future interactive robots and contribute to the ongoing discourse on intelligence assessment in AI-driven systems. Lavinia Hriscu, Alberto Sanfeliu, Anais Garrell |
RO-MAN | 2 |
| 2024 | Learning Priors of Human Motion With Vision TransformersabstractA clear understanding of where humans move in a scenario, their usual paths and speeds, and where they stop, is very important for different applications, such as mobility studies in urban areas or robot navigation tasks within human-populated environments. We propose in this article, a neural architecture based on Vision Transformers (ViTs) to provide this information. This solution can arguably capture spatial correlations more effectively than Convolutional Neural Networks (CNNs). In the paper, we describe the methodology and proposed neural architecture and show the experiments' results with a standard dataset. We show that the proposed ViT architecture improves the metrics compared to a method based on a CNN. Placido Falqueto, Alberto Sanfeliu, Luigi Palopoli 0002, Daniele Fontanelli |
COMPSAC | 2 |
| 2024 | Perception for Collaborative Robots in Pruning OperationsabstractIn this work a set of novel approaches based on well-known computer vision techniques is proposed to deal with the autonomous perception part of an HRI robotic system. In the considered scenario, a human-robot duo interact to plan maintenance and pruning operations in a vineyard. Once the human expert has selected a set of branches to cut in the approximate area, the robot stores this information in order to be able to return later, identify the branches and cut them at an adequate point. This is achieved by producing models of the plant by segmenting the view through a combination of intensity and depth data, using RGB-D camera sensors. This segmentation is fed back into the pipeline, used as a base mask to identify parts of the plant through watershed segmentation, and identifying the branches and gems in the segmented images. The proposed method was develop based on real data, and tested in experimental scenarios with the real robot. Marco Giacchetti, Edmundo Guerra, Francisco Cristóbal García, Yolanda Bolea, Antoni Grau-Saldes, Alberto Sanfeliu |
ETFA | 6 |
| 2024 | Exploring Transformers and Visual Transformers for Force Prediction in Human-Robot Collaborative Transportation TasksabstractIn this paper, we analyze the possibilities offered by Deep Learning State-of-the-Art architectures such as Transformers and Visual Transformers in generating a prediction of the human’s force in a Human-Robot collaborative object transportation task at a middle distance. We outperform our previous predictor by achieving a success rate of 93.8% in testset and 90.9% in real experiments with 21 volunteers predicting in both cases the force that the human will exert during the next 1 s. A modification in the architecture allows us to obtain a second output from the model with a velocity prediction, which allows us to improve the capabilities of our predictor if it is used to estimate the trajectory that the human-robot pair will follow. An ablation test is also performed to verify the relative contribution to performance of each input. Jose Enrique Domínguez-Vidal, Alberto Sanfeliu |
ICRA | 2 |
| 2024 | Force and Velocity Prediction in Human-Robot Collaborative Transportation Tasks through Video Retentive NetworksabstractIn this article, we propose a generalization of a Deep Learning State-of-the-Art architecture such as Retentive Networks so that it can accept video sequences as input. With this generalization, we design a force/velocity predictor applied to the medium-distance Human-Robot collaborative object transportation task. We achieve better results than with our previous predictor by reaching success rates in testset of up to 93.7% in predicting the force to be exerted by the human and up to 96.5% in the velocity of the human-robot pair during the next 1 s, and up to 91.0% and 95.0% respectively in real experiments. This new architecture also manages to improve inference times by up to 32.8% with different graphics cards. Finally, an ablation test allows us to detect that one of the input variables used so far, such as the position of the task goal, could be discarded allowing this goal to be chosen dynamically by the human instead of being pre-set. Jose Enrique Domínguez-Vidal, Alberto Sanfeliu |
IROS | 2 |
| 2024 | Perception-Driven Shared Control Architecture for Agricultural Robots Performing Harvesting TasksabstractThis paper introduces a shared control framework designed specifically for agricultural mobile manipulators engaged in harvesting operations. The shared control strategy allows for achieving such operations by dynamically exchanging the control between the robotic system and a human operator depending on the uncertainty in the environment perception. For this purpose, the robot’s behavior is dynamically adapted to switch between two control modes with a different level of autonomy of the robot. The level of autonomy is encoded in two different admittance behaviors which are included in a first-order Hierarchical Quadratic Programming (HQP) control framework, that allows the robot to simultaneously address other control objectives at the same time. Experimental results with a dual-arm mobile robot, developed as part of the EU-funded CANOPIES project, demonstrate the effectiveness of the proposed method in real conditions. Jozsef Palmieri, Paolo Di Lillo, Alberto Sanfeliu, Alessandro Marino |
IROS | 3 |
| 2024 | Anticipation and Proactivity. Unraveling Both Concepts in Human-Robot Interaction through a Handover ExampleabstractWhile robots have advanced in understanding their environments, collaborative tasks demand a deeper comprehension of human intentions to mitigate uncertainty. Anticipatory and proactive behaviours are pivotal in enhancing Human-Robot Interactions (HRI), yet literature often conflates these terms. This study elucidates the distinction between anticipation and proactivity, offering clear definitions and exemplifying their implications through a handover scenario. Through a user study with 24 volunteers performing a total of 72 experiments, we have found that humans are able to distinguish both behaviours and that there is a statistically significant increase in the anthropomorphism of the robot when it behaves proactively. Additionally, both anticipation and proactivity show statistically significant increases in multiple aspects of effective HRI (fluency, comfort, performance, etc.). However, no clear preference for either has been detected. Jose Enrique Domínguez-Vidal, Alberto Sanfeliu |
RO-MAN | 2 |
| 2023 | Improving Human-Robot Interaction Effectiveness in Human-Robot Collaborative Object Transportation Using Force PredictionabstractIn this work, we analyse the use of a prediction of the human's force in a Human-Robot collaborative object transportation task at a middle distance. We check that this force prediction can improve multiple parameters associated with effective Human-Robot Interaction (HRI) such as perception of the robot's contribution to the task, comfort or trust in the robot in a physical Human Robot Interaction (pHRI). We present a Deep Learning model that allows to predict the force that a human will exert in the next 1$s$using as inputs the force previously exerted by the human, the robot's velocity and environment information obtained from the robot's LiDAR. Its success rate is up to 92.3% in testset and up to 89.1 % in real experiments. We demonstrate that this force prediction, in addition to being able to be used directly to detect changes in the human's intention, can be processed to obtain an estimate of the human's desired trajectory. We have validated this approach with a user study involving 18 volunteers. Jose Enrique Domínguez-Vidal, Alberto Sanfeliu |
IROS | 2 |
| 2023 | Inference VS. Explicitness. Do We Really Need the Perfect Predictor? The Human-Robot Collaborative Object Transportation CaseabstractWhen robots interact with humans, limitations in their internal models arise due to the uncertainty and even randomness of human behavior. This has led to attempts to predict human future actions and infer their intent. However, some authors argue for combining inference engines with communication systems that explicitly elicit human intention. This work builds on our Perception-Intention-Action (PIA) cycle, a framework that considers human intention at the same level as perception of the environment. The PIA cycle is used in a collaborative task to compare the effect on different human-robot interaction aspects of using a force predictor that infers human implicit intention versus a communication system that explicitly elicits human intention. A study with 18 volunteers shows that allowing humans to directly express themselves can achieve the same improvement as an intention predictor. Jose Enrique Domínguez-Vidal, Alberto Sanfeliu |
RO-MAN | 2 |
| 2023 | Real-Life Experiment Metrics for Evaluating Human-Robot Collaborative Navigation TasksabstractAs robots move from laboratories and industries to the real world, they must develop new abilities to collaborate with humans in various aspects, including human-robot collaborative navigation (HRCN) tasks. Then, it is required to develop general methodologies to evaluate these robots’ behaviors. These methodologies should incorporate objective and subjective measurements. Objective measurements for evaluating a robot’s behavior while navigating with others can be accomplished using social distances in conjunction with task characteristics, people-robot relationships, and physical space. Additionally, the objective evaluation of the task must consider human behavior, which is influenced by changes and the structure of their environment. Subjective evaluations of robot’s behaviors can be conducted using surveys that address various aspects of robot usability. This includes people’s perceptions of their interaction during their collaborative task with the robot, focusing on aspects such as sociability, comfort, and task-intelligence. Moreover, the communicative interaction between the agents (people and robots) involved in the collaborative task should also be evaluated. Therefore, this paper presents a comprehensive methodology for objectively and subjectively evaluating HRCN tasks. Ely Repiso-Polo, Anais Garrell, Alberto Sanfeliu |
RO-MAN | 3 |
| 2022 | IVO Robot: A New Social Robot for Human-Robot CollaborationabstractWe present a new social robot named IVO, a robot capable of collaborating with humans and solving different tasks. The robot is intended to cooperate and work with humans in a useful and socially acceptable manner to serve as a research platform for long-term Social Human-Robot Interaction. In this paper, we proceed to describe this new platform, its communication skills and the current capabilities the robot possesses, such as, handing over an object to or from a person or performing guiding tasks with a human through physical contact. We describe the social abilities of the IVO robot, furthermore, we present the experiments performed for each robot's capacity using its current version. Javier Laplaza, Jose Enrique Domínguez-Vidal, Fernando Herrero, Alberto Sanfeliu, Anais Garrell |
HRI | 7 |
| 2022 | Context and Intention aware 3D Human Body Motion Prediction using an Attention Deep Learning model in Handover TasksabstractThis work explores how contextual information and human intention affect the motion prediction of humans during a handover operation with a social robot. By classifying human intention in four different classes, we developed a model able to generate a different motion for each intention class. Furthermore, the model uses a multi-headed attention architecture to add contextual information to the pipeline, such as the position of the robot end effector (REE) or the position of obstacles in the interaction scene. We generate predictions up to two and half seconds in the future given an input sequence of one second containing the previous motion of the human. The results show an improvement of the prediction accuracy, both for the full skeleton prediction and the human hand used for the delivery. The model also allows to generate different sequences with the desired human intention. Javier Laplaza, Francesc Moreno-Noguer, Alberto Sanfeliu |
IROS | 3 |
| 2022 | Context and Intention for 3D Human Motion Prediction: Experimentation and User study in Handover TasksabstractIn this work we present a novel attention deep learning model that uses context and human intention for 3D human body motion prediction in handover human-robot tasks. This model uses a multi-head attention architecture which incorporates as inputs the human motion, the robot end effector and the position of the obstacles. The outputs of the model are the predicted motion of the human body and the predicted human intention. We use this model to analyze a handover collaborative task with a robot where the robot is able to predict the future motion of the human and use this information in it’s planner. Several experiments are performed where human volunteers fill a standard poll to rate different features, taking into account when the robot uses the prediction versus when the robot doesn’t use the prediction. Javier Laplaza, Anais Garrell, Francesc Moreno-Noguer, Alberto Sanfeliu |
RO-MAN | 4 |
| 2021 | Body Size and Depth Disambiguation in Multi-Person Reconstruction from Single ImagesabstractWe address the problem of multi-person 3D body pose and shape estimation from a single image. While this problem can be addressed by applying single-person approaches multiple times for the same scene, recent works have shown the advantages of building upon deep architectures that simultaneously reason about all people in the scene in a holistic manner by enforcing, e.g., depth order constraints or minimizing interpenetration among reconstructed bodies. However, existing approaches are still unable to capture the size variability of people caused by the inherent body scale and depth ambiguity. In this work, we tackle this challenge by devising a novel optimization scheme that learns the appropriate body scale and relative camera pose, by enforcing the feet of all people to remain on the ground floor. A thorough evaluation on MuPoTS- 3D and 3DPW datasets demonstrates that our approach is able to robustly estimate the body translation and shape of multiple people while retrieving their spatial arrangement, consistently improving current state-of-the-art, especially in scenes with people of very different heights. Code can be found at: https://github.com/nicolasugrinovic/size_depth_disambiguation Nicolas Ugrinovic, Adria Ruiz, Antonio Agudo, Alberto Sanfeliu, Francesc Moreno-Noguer |
3DV | 4 |
| 2021 | Human-Robot Collaborative Multi-Agent Path Planning using Monte Carlo Tree Search and Social Reward SourcesabstractThe collaboration between humans and robots in an object search task requires the achievement of shared plans obtained from communicating and negotiating. In this work, we assume that the robot computes, as a first step, a multi-agent plan for both itself and the human. Then, both plans are submitted to human scrutiny, who either agrees or modifies it forcing the robot to adapt its own restrictions or preferences. This process is repeated along the search task as many times as required by the human. Our planner is based on a decentralized variant of Monte Carlo Tree Search (MCTS), with one robot and one human as agents. Moreover, our algorithm allows the robot and the human to optimize their own actions by maintaining a probability distribution over the plans in a joint-action space. The method allows an objective function definition over action sequences, it assumes intermittent communication, it is anytime and suitable for on-line replanning. To test it, we have developed a human-robot communication mobile phone interface. Validation is provided by real-life search experiments of a Parcheesi token in an urban space, including also an acceptability study. Marc Dalmasso, Anais Garrell, José Enrique Domínguez, Pablo Jiménez, Alberto Sanfeliu |
ICRA | 5 |
| 2021 | User-Friendly Smartphone Interface to Share Knowledge in Human-Robot Collaborative Search TasksabstractLong-distance human-robot collaborative tasks require robust forms of knowledge-sharing among agents in order to optimize the performance of the task. In this paper, we propose to take advantage of the proliferation of mobile phones to use them as a reliable low-cost communication interface, as opposed to the use of specific gadgets or speech and gesture recognition techniques that are highly prone to failure in the presence of noise or occlusions. Our interface is focused on search tasks, and it allows the user to share with other agents real-time information such as their position, their intention or even what they would like the other agents to do. To test its acceptability, a user study was conducted with 20 volunteers in a human-human scenario. A second round of experiments with other 30 volunteers was conducted to test different ways to encourage user interaction with our interface. Finally, real-life experiments were also conducted with a robot to apply learned knowledge to the desired scenario. We found a statistically significant improvement in the amount of information exchanged between agents. Jose Enrique Domínguez-Vidal, Iván J. Torres-Rodríguez, Anais Garrell, Alberto Sanfeliu |
RO-MAN | 4 |
| 2021 | Attention deep learning based model for predicting the 3D Human Body Pose using the Robot Human Handover PhasesabstractThis work proposes a human motion prediction model for handover operations. We use in this work, the different phases of the handover operation to improve the human motion predictions. Our attention deep learning based model takes into account the position of the robot’s End Effector (REE) and the phase in the handover operation to predict future human poses. Our model outputs a distribution of possible positions rather than one deterministic position, a key feature in order to allow robots to collaborate with humans. We provide results of the human upper body and the human right hand, also referred as Human End Effector (HEE).The attention deep learning based model has been trained and evaluated with a dataset created using human volunteers and an anthropomorphic robot, simulating handover operations where the robot is the giver and the human the receiver. For each operation, the human skeleton is obtained with an Intel RealSense D435i camera attached inside the robot’s head. The results shown a great improvement of the human’s right hand prediction and 3D body compared with other methods. Javier Laplaza, Albert Pumarola, Francesc Moreno-Noguer, Alberto Sanfeliu |
RO-MAN | 4 |
| 2021 | Dual-Branch CNNs for Vehicle Detection and Tracking on LiDAR DataabstractWe present a novel vehicle detection and tracking system that works solely on 3D LiDAR information. Our approach segments vehicles using a dual-view representation of the 3D LiDAR point cloud on two independently trained convolutional neural networks, one for each view. A bounding box growing algorithm is applied to the fused output of the networks to properly enclose the segmented vehicles. Bounding boxes are grown using a probabilistic method that takes into account also occluded areas. The final vehicle bounding boxes act as observations for a multi-hypothesis tracking system which allows to estimate the position and velocity of the observed vehicles. We thoroughly evaluate our system on the KITTI benchmarks both for detection and tracking separately and show that our dual-branch classifier consistently outperforms previous single-branch approaches, improving or directly competing to other state of the art LiDAR-based methods. Victor Vaquero, Iván del Pino, Francesc Moreno-Noguer, Joan Solà, Alberto Sanfeliu, Juan Andrade-Cetto |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2020 | GANimation: One-Shot Anatomically Consistent Facial Animation
Albert Pumarola, Antonio Agudo, Aleix Martinez, Alberto Sanfeliu, Francesc Moreno-Noguer |
Int. J. Comput. Vis. | 4 |
| 2019 | 3DPeople: Modeling the Geometry of Dressed HumansabstractRecent advances in 3D human shape estimation build upon parametric representations that model very well the shape of the naked body, but are not appropriate to represent the clothing geometry. In this paper, we present an approach to model dressed humans and predict their geometry from single images. We contribute in three fundamental aspects of the problem, namely, a new dataset, a novel shape parameterization algorithm and an end-to-end deep generative network for predicting shape. First, we present 3D-People, a large-scale synthetic dataset with 2 Million photo-realistic images of 80 subjects performing 70 activities and wearing diverse outfits. Besides providing textured 3D meshes for clothes and body we annotated the dataset with segmentation masks, skeletons, depth, normal maps and optical flow. All this together makes 3D-People suitable for a plethora of tasks. We then represent the 3D shapes using 2D geometry images. To build these images we propose a novel spherical area-preserving parameterization algorithm based on the optimal mass transportation method. We show this approach to improve existing spherical maps which tend to shrink the elongated parts of the full body models such as the arms and legs, making the geometry images incomplete. Finally, we design a multi-resolution deep generative network that, given an input image of a dressed human, predicts his/her geometry image (and thus the clothed body shape) in an end-to-end manner. We obtain very promising results in jointly capturing body pose and clothing shape, both for synthetic validation and on the wild images. Albert Pumarola, Jordi Sanchez, Gary Pui-Tung Choi, Alberto Sanfeliu, Francesc Moreno-Noguer |
ICCV | 4 |
| 2019 | Teaching a Drone to Accompany a Person from Demonstrations using Non-Linear ASFMabstractIn this paper, we present a new method based on the Aerial Social Force Model (ASFM) to allow human-drone side-by-side social navigation in real environments. To tackle this problem, the present work proposes a new nonlinear-based approach using Neural Networks. To learn and test the rightness of the new approach, we built a new dataset with simulated environments and we recorded motion controls provided by a human expert tele-operating the drone. The recorded data is then used to train a neural network which maps interaction forces to acceleration commands. The system is also reinforced with a human path prediction module to improve the drone's navigation, as well as, a collision detection module to completely avoid possible impacts. Moreover, a performance metric is defined which allows us to numerically evaluate and compare the fulfillment of the different learned policies. The method was validated by a large set of simulations; we also conducted real-life experiments with an autonomous drone to verify the framework described for the navigation process. In addition, a user study has been realized to reveal the social acceptability of the method. Anais Garrell, Carles Coll, René Alquézar, Alberto Sanfeliu |
IROS | 4 |
| 2019 | People's V-Formation and Side-by-Side Model Adapted to Accompany Groups of People by Social RobotsabstractThis paper presents a new method to allow robots to accompany a person or a group of people imitating pedestrians behavior. Two-people groups usually walk in a side-by-side formation and three-people groups walk in a V-formation so that they can see each other. For this reason, the proposed method combines a Side-by-side and V-formation pedestrian model with the Anticipative Kinodynamic Planner (AKP). Combining these methods, the robot is able to do an anticipatory accompaniment of groups of humans, as well as to avoid static and dynamic obstacles in advance, while keeping the prescribed formations. The proposed framework allows also a dynamical re-positioning of the robot, if the physical position of the partners change in the group formation. Furthermore, people have a randomness factor that the robot has to manage, for that reason, the system was adapted to deal with changes in people's velocity, orientation and occlusions. Finally, the method has been validated using synthetic experiments and real-life experiments with our Tibi robot. In addition, a user study has been realized to reveal the social acceptability of the method. Ely Repiso-Polo, Francesco Zanlungo, Takayuki Kanda 0001, Anais Garrell, Alberto Sanfeliu |
IROS | 5 |
| 2019 | Online learning and detection of faces with low human supervision
Michael Villamizar, Alberto Sanfeliu, Francesc Moreno-Noguer |
Vis. Comput. | 2 |
| 2018 | Geometry-Aware Network for Non-Rigid Shape Prediction From a Single ViewabstractWe propose a method for predicting the 3D shape of a deformable surface from a single view. By contrast with previous approaches, we do not need a pre-registered template of the surface, and our method is robust to the lack of texture and partial occlusions. At the core of our approach is a geometry-aware deep architecture that tackles the problem as usually done in analytic solutions: first perform 2D detection of the mesh and then estimate a 3D shape that is geometrically consistent with the image. We train this architecture in an end-to-end manner using a large dataset of synthetic renderings of shapes under different levels of deformation, material properties, textures and lighting conditions. We evaluate our approach on a test split of this dataset and available real benchmarks, consistently improving state-of-the-art solutions with a significantly lower computational time. Albert Pumarola, Antonio Agudo, Lorenzo Porzi, Alberto Sanfeliu, Vincent Lepetit, Francesc Moreno-Noguer |
CVPR | 4 |
| 2018 | Unsupervised Person Image Synthesis in Arbitrary PosesabstractWe present a novel approach for synthesizing photorealistic images of people in arbitrary poses using generative adversarial learning. Given an input image of a person and a desired pose represented by a 2D skeleton, our model renders the image of the same person under the new pose, synthesizing novel views of the parts visible in the input image and hallucinating those that are not seen. This problem has recently been addressed in a supervised manner [16, 35], i.e., during training the ground truth images under the new poses are given to the network. We go beyond these approaches by proposing a fully unsupervised strategy. We tackle this challenging scenario by splitting the problem into two principal subtasks. First, we consider a pose conditioned bidirectional generator that maps back the initially rendered image to the original pose, hence being directly comparable to the input image without the need to resort to any training image. Second, we devise a novel loss function that incorporates content and style terms, and aims at producing images of high perceptual quality. Extensive experiments conducted on the DeepFashion dataset demonstrate that the images rendered by our model are very close in appearance to those obtained by fully supervised approaches. Albert Pumarola, Antonio Agudo, Alberto Sanfeliu, Francesc Moreno-Noguer |
CVPR | 3 |
| 2018 | GANimation: Anatomically-Aware Facial Animation from a Single Image
Albert Pumarola, Antonio Agudo, Aleix Martinez, Alberto Sanfeliu, Francesc Moreno-Noguer |
ECCV (10) | 4 |
| 2018 | Sustainable Robotics Solutions in Smart Cities. The Challenge of the ECHORD++ ProjectabstractThe objective of this paper is to explain novel sustainable robotics solutions for cities. Those new proposals appear under the ECHORD++ project which is a good tool to meet academia and industry with the objective of providing innovative technological solutions. In this paper, authors explain the tool as well as the methodology to promote robotics research in urban environments, and the on-going experience will demonstrate that huge advances are made in this field. Antoni Grau-Saldes, Yolanda Bolea, Ana Puig-Pey, Alberto Sanfeliu, Josep Casanovas |
ETFA | 4 |
| 2018 | Hallucinating Dense Optical Flow from Sparse Lidar for Autonomous VehiclesabstractIn this paper we propose a novel approach to estimate dense optical flow from sparse lidar data acquired on an autonomous vehicle. This is intended to be used as a drop-in replacement of any image-based optical flow system when images are not reliable due to e.g. adverse weather conditions or at night. In order to infer high resolution 2D flows from discrete range data we devise a three-block architecture of multiscale filters that combines multiple intermediate objectives, both in the lidar and image domain. To train this network we introduce a dataset with approximately 20K lidar samples of the Kitti dataset which we have augmented with a pseudo ground-truth image-based optical flow computed using FlowNet2. We demonstrate the effectiveness of our approach on Kitti, and show that despite using the low-resolution and sparse measurements of the lidar, we can regress dense optical flow maps which are at par with those estimated with image-based methods. Victor Vaquero, Alberto Sanfeliu, Francesc Moreno-Noguer |
ICPR | 2 |
| 2018 | Deep Lidar CNN to Understand the Dynamics of Moving VehiclesabstractPerception technologies in Autonomous Driving are experiencing their golden age due to the advances in Deep Learning. Yet, most of these systems rely on the semantically rich information of RGB images. Deep Learning solutions applied to the data of other sensors typically mounted on autonomous cars (e.g. lidars or radars) are not explored much. In this paper we propose a novel solution to understand the dynamics of moving vehicles of the scene from only lidar information. The main challenge of this problem stems from the fact that we need to disambiguate the proprio-motion of the “observer” vehicle from that of the external “observed” vehicles. For this purpose, we devise a CNN architecture which at testing time is fed with pairs of consecutive lidar scans. However, in order to properly learn the parameters of this network, during training we introduce a series of so-called pretext tasks which also leverage on image data. These tasks include semantic information about vehicleness and a novel lidar-flow feature which combines standard image-based optical flow with lidar scans. We obtain very promising results and show that including distilled image information only during training, allows improving the inference results of the network at test time, even when image data is no longer used. Victor Vaquero, Alberto Sanfeliu, Francesc Moreno-Noguer |
ICRA | 2 |
| 2018 | Robot Approaching and Engaging People in a Human-Robot Companion FrameworkabstractThis paper presents a new model to make robots capable of approaching and engaging people with a human-like behavior, while they are walking in a side-by-side formation with a person. This method extends our previous work [1], which allows the robot to adapt its navigation behaviour according to the person being accompanied and the dynamic environment. In the current work, the robot is able to predict the best encounter point between the human-robot group and the approached person. Then, in the encounter point the robot modifies its position to achieve an engagement with both people. The encounter point is computed using a gradient descent method that takes into account all people predictions. Moreover, we make use of the Extended Social Force Model (ESFM), and it is modified to include the dynamic goal. The method has been validated over several situations and in real-life experiments, in addition, a user study has been realized to reveal the social acceptability of the robot in this task. Ely Repiso-Polo, Anais Garrell, Alberto Sanfeliu |
IROS | 3 |
| 2018 | Boosted Random Ferns for Object DetectionabstractIn this paper we introduce the Boosted Random Ferns (BRFs) to rapidly build discriminative classifiers for learning and detecting object categories. At the core of our approach we use standard random ferns, but we introduce four main innovations that let us bring ferns from an instance to a category level, and still retain efficiency. First, we define binary features on the histogram of oriented gradients-domain (as opposed to intensity-), allowing for a better representation of intra-class variability. Second, both the positions where ferns are evaluated within the sliding window, and the location of the binary features for each fern are not chosen completely at random, but instead we use a boosting strategy to pick the most discriminative combination of them. This is further enhanced by our third contribution, that is to adapt the boosting strategy to enable sharing of binary features among different ferns, yielding high recognition rates at a low computational cost. And finally, we show that training can be performed online, for sequentially arriving images. Overall, the resulting classifier can be very efficiently trained, densely evaluated for all image locations in about 0.1 seconds, and provides detection rates similar to competing approaches that require expensive and significantly slower processing times. We demonstrate the effectiveness of our approach by thorough experimentation in publicly available datasets in which we compare against state-of-the-art, and for tasks of both 2D detection and 3D multi-view estimation. Michael Villamizar, Juan Andrade-Cetto, Alberto Sanfeliu, Francesc Moreno-Noguer |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Joint coarse-and-fine reasoning for deep optical flowabstractWe propose a novel representation for dense pixel-wise estimation tasks using CNNs that boosts accuracy and reduces training time, by explicitly exploiting joint coarse-and-fine reasoning. The coarse reasoning is performed over a discrete classification space to obtain a general rough solution, while the fine details of the solution are obtained over a continuous regression space. In our approach both components are jointly estimated, which proved to be beneficial for improving estimation accuracy. Additionally, we propose a new network architecture, which combines coarse and fine components by treating the fine estimation as a refinement built on top of the coarse solution, and therefore adding details to the general prediction. We apply our approach to the challenging problem of optical flow estimation and empirically validate it against state-of-the-art CNN-based solutions trained from scratch and tested on large optical flow datasets. Victor Vaquero, Germán Ros 0001, Francesc Moreno-Noguer, Antonio M. López 0001, Alberto Sanfeliu |
ICIP | 5 |
| 2017 | PL-SLAM: Real-time monocular visual SLAM with points and linesabstractLow textured scenes are well known to be one of the main Achilles heels of geometric computer vision algorithms relying on point correspondences, and in particular for visual SLAM. Yet, there are many environments in which, despite being low textured, one can still reliably estimate line-based geometric primitives, for instance in city and indoor scenes, or in the so-called “Manhattan worlds”, where structured edges are predominant. In this paper we propose a solution to handle these situations. Specifically, we build upon ORB-SLAM, presumably the current state-of-the-art solution both in terms of accuracy as efficiency, and extend its formulation to simultaneously handle both point and line correspondences. We propose a solution that can even work when most of the points are vanished out from the input images, and, interestingly it can be initialized from solely the detection of line correspondences in three consecutive frames. We thoroughly evaluate our approach and the new initialization strategy on the TUM RGB-D benchmark and demonstrate that the use of lines does not only improve the performance of the original ORB-SLAM solution in poorly textured frames, but also systematically improves it in sequence frames combining points and lines, without compromising the efficiency. Albert Pumarola, Alexander Vakhitov, Antonio Agudo, Alberto Sanfeliu, Francesc Moreno-Noguer |
ICRA | 4 |
| 2017 | Aerial social force model: A new framework to accompany people using autonomous flying robotsabstractIn this paper, we propose a novel Aerial Social Force Model (ASFM) that allows autonomous flying robots to accompany humans in urban environments in a safe and comfortable manner. To date, we are not aware of other state-of-the-art method that accomplish this task. The proposed approach is a 3D version of the Social Force Model (SFM) for the field of aerial robots which includes an interactive human-robot navigation scheme capable of predicting human motions and intentions so as to safely accompany them to their final destination. ASFM also introduces a new metric to fine-tune the parameters of the force model, and to evaluate the performance of the aerial robot companion based on comfort and distance between the robot and humans. The presented approach is extensively validated in diverse simulations and real experiments, and compared against other similar works in the literature. ASFM attains remarkable results and shows that it is a valuable framework for social robotics applications, such as guiding people or human-robot interaction. Anais Garrell, Luis Garza-Elizondo, Michael Villamizar, Fernando Herrero, Alberto Sanfeliu |
IROS | 5 |
| 2017 | On-line adaptive side-by-side human robot companion in dynamic urban environmentsabstractThis paper presents an adaptive side-by-side human-robot companion approach for navigation in urban dynamic environments, based on the anticipative kinodynamic planning. The adaptive means that the robot is capable of adjusting its motion to the behavior of the person being accompanied. Our main objective is to optimize in real time the path performed by the pair human-robot, by modifying dynamically the angle and distance between both throughout different locations of the path. We have defined a new cost function for finding the best planned path that takes into account the cost of the geometrical configuration between the human and the robot. Moreover, we have modified the Extended Social Force Model (SFM) to include the required forces to maintain the angle and distance between the robot and human while the human-robot pair is moving towards the shared goal. The method has been validated throughout a large set of simulations and real-live experiments. Ely Repiso-Polo, Gonzalo Ferrer 0001, Alberto Sanfeliu |
IROS | 3 |
| 2017 | Random clustering ferns for multimodal object recognition
Michael Villamizar, Anais Garrell, Alberto Sanfeliu, Francesc Moreno-Noguer |
Neural Comput. Appl. | 3 |
| 2016 | Learning the hidden human knowledge of UAV pilots when navigating in a cluttered environment for improving path planningabstractWe propose in this work a new model of how the hidden human knowledge (HHK) of UAV pilots can be incorporated in the UAVs path planning generation. We intuitively know that human's pilots barely manage or even attempt to drive the UAV through a path that is optimal attending to some criteria as an optimal planner would suggest. Although human pilots might get close but not reach the optimal path proposed by some planner that optimizes over time or distance, the final effect of this differentiation could be not only surprisingly better, but also desirable. In the best scenario for optimality, the path that human pilots generate would deviate from the optimal path as much as the hidden knowledge that its perceives is injected into the path. The aim of our work is to use real human pilot paths to learn the hidden knowledge using repulsion fields and to incorporate this knowledge afterwards in the environment obstacles as cause of the deviation from optimality. We present a strategy of learning this knowledge based on attractor and repulsors, the learning method and a modified RRT* that can use this knowledge for path planning. Finally we do real-life tests and we compare the resulting paths with and without this knowledge. Ignacio Alzugaray, Alberto Sanfeliu |
IROS | 2 |
| 2016 | Interactive multiple object learning with scanty human supervision
Michael Villamizar, Anais Garrell, Alberto Sanfeliu, Francesc Moreno-Noguer |
Comput. Vis. Image Underst. | 3 |
| 2016 | Filtering graphs to check isomorphism and extracting mapping by using the Conductance Electrical Model
Manuel Igelmo, Alberto Sanfeliu |
Pattern Recognit. | 2 |
| 2015 | Efficient monocular pose estimation for complex 3D modelsabstractWe propose a robust and efficient method to estimate the pose of a camera with respect to complex 3D textured models of the environment that can potentially contain more than 100; 000 points. To tackle this problem we follow a top down approach where we combine high-level deep network classifiers with low level geometric approaches to come up with a solution that is fast, robust and accurate. Given an input image, we initially use a pre-trained deep network to compute a rough estimation of the camera pose. This initial estimate constrains the number of 3D model points that can be seen from the camera viewpoint. We then establish 3D-to-2D correspondences between these potentially visible points of the model and the 2D detected image features. Accurate pose estimation is finally obtained from the 2D-to-3D correspondences using a novel PnP algorithm that rejects outliers without the need to use a RANSAC strategy, and which is between 10 and 100 times faster than other methods that use it. Two real experiments dealing with very large and complex 3D models demonstrate the effectiveness of the approach. Antonio Rubio 0001, Michael Villamizar, Luis Ferraz, Adrián Peñate Sánchez, Arnau Ramisa, Edgar Simo-Serra, Alberto Sanfeliu, Francesc Moreno-Noguer |
ICRA | 7 |
| 2015 | Modeling robot's world with minimal effortabstractWe propose an efficient Human Robot Interaction approach to efficiently model the appearance of all relevant objects in robot's environment. Given an input video stream recorded while the robot is navigating, the user just needs to annotate a very small number of frames to build specific classifiers for each of the objects of interest. At the core of the method, there are several random ferns classifiers that share the same features and are updated online. The resulting methodology is fast (runs at 8 fps), versatile (it can be applied to unconstrained scenarios), scalable (real experiments show we can model up to 30 different object classes), and minimizes the amount of human intervention by leveraging the uncertainty measures associated to each classifier. We thoroughly validate the approach on synthetic data and on real sequences acquired with a mobile platform in outdoor and challenging scenarios containing a multitude of different objects. We show that the human can, with minimal effort, provide the robot with a detailed model of the objects in the scene. Michael Villamizar, Anais Garrell, Alberto Sanfeliu, Francesc Moreno-Noguer |
ICRA | 3 |
| 2015 | Multi-objective cost-to-go functions on robot navigation in dynamic environmentsabstractIn our previous work [1] we introduced the Anticipative Kinodynamic Planning (AKP): a robot navigation algorithm in dynamic urban environments that seeks to minimize its disruption to nearby pedestrians. In the present paper, we maintain all the advantages of the AKP, and we overcome the previous limitations by presenting novel contributions to our approach. Firstly, we present a multi-objective cost function to consider different and independent criteria and a well-posed procedure to build a joint cost function in order to select the best path. Then, we improve the construction of the planner tree by introducing a cost-to-go function that will be shown to outperform a classical Euclidean distance approach. In order to achieve real time calculations, we have used a steering heuristic that dramatically speeds up the process. Plenty of simulations and real experiments have been carried out to demonstrate the success of the AKP. Gonzalo Ferrer 0001, Alberto Sanfeliu |
IROS | 2 |
| 2015 | Combining where and what in change detection for unsupervised foreground learning in surveillance
Ivan Huerta Casado, Marco Pedersoli, Jordi Gonzàlez 0001, Alberto Sanfeliu |
Pattern Recognit. | 4 |
| 2014 | Segmentation-Aware Deformable Part ModelsabstractIn this work we propose a technique to combine bottom-up segmentation, coming in the form of SLIC superpixels, with sliding window detectors, such as Deformable Part Models (DPMs). The merit of our approach lies in "cleaning up" the low-level HOG features by exploiting the spatial support of SLIC superpixels, this can be understood as using segmentation to split the feature variation into object-specific and background changes. Rather than committing to a single segmentation we use a large pool of SLIC superpixels and combine them in a scale-, position- and object-dependent manner to build soft segmentation masks. The segmentation masks can be computed fast enough to repeat this process over every candidate window, during training and detection, for both the root and part filters of DPMs. We use these masks to construct enhanced, background-invariant features to train DPMs. We test our approach on the PASCAL VOC 2007, outperforming the standard DPM in 17 out of 20 classes, yielding an average increase of 1.7% AP. Additionally, we demonstrate the robustness of this approach, extending it to dense SIFT descriptors for large displacement optical flow. Eduard Trulls, Stavros Tsogkas, Iasonas Kokkinos, Alberto Sanfeliu, Francesc Moreno-Noguer |
CVPR | 4 |
| 2014 | On-board real-time pose estimation for UAVs using deformable visual contour registrationabstractWe present a real time algorithm for estimating the pose of non-planar objects on which we have placed a visual marker. It is designed to overcome the limitations of small aerial robots, such as slow CPUs, low image resolution and geometric distortions produced by wide angle lenses or viewpoint changes. The method initially registers the shape of a known marker to the contours extracted in an image. For this purpose, and in contrast to state-of-the art, we do not seek to match textured patches or points of interest. Instead, we optimize a geometric alignment cost computed directly from raw polygonal representations of the observed regions using very simple and efficient clipping algorithms. Further speed is achieved by performing the optimization in the polygon representation space, avoiding the need of 2D image processing operations. Deformation modes are easily included in the optimization scheme, allowing an accurate registration of different markers attached to curved surfaces using a single deformable prototype. Once this initial registration is solved, the object pose is retrieved using a standard PnP approach. As a result, the method achieves accurate object pose estimation in real-time, which is very important for interactive UAV tasks, for example for short distance surveillance or bar assembly. We present experiments where our method yields, at about 30Hz, an average error of less than 5mm in estimating the position of a 19×19mm marker placed at 0.7m of the camera. Adrian Amor-Martinez, Alberto Ruiz, Francesc Moreno-Noguer, Alberto Sanfeliu |
ICRA | 4 |
| 2014 | Behavior estimation for a complete framework for human motion prediction in crowded environmentsabstractIn the present work, we propose and validate a complete probabilistic framework for human motion prediction in urban or social environments. Additionally, we formulate a powerful and useful tool: the human motion behavior estimator. Three different basic behaviors have been detected: Aware, Balanced and Unaware. Our approach is based on the Social Force Model (SFM) and the intentionality prediction BHMIP. The main contribution of the present work is to make use of the behavior estimator for formulating a reliable prediction framework of human trajectories under the influence of dynamic crowds, robots, and in general any moving obstacle. Accordingly, we have demonstrated the great performance of our long-term prediction algorithm, in real scenarios, comparing to other prediction methods. Gonzalo Ferrer 0001, Alberto Sanfeliu |
ICRA | 2 |
| 2014 | Fast online learning and detection of natural landmarks for autonomous aerial robotsabstractWe present a method for efficiently detecting natural landmarks that can handle scenes with highly repetitive patterns and targets progressively changing its appearance. At the core of our approach lies a Random Ferns classifier, that models the posterior probabilities of different views of the target using multiple and independent Ferns, each containing features at particular positions of the target. A Shannon entropy measure is used to pick the most informative locations of these features. This minimizes the number of Ferns while maximizing its discriminative power, allowing thus, for robust detections at low computational costs. In addition, after offline initialization, the new incoming detections are used to update the posterior probabilities on the fly, and adapt to changing appearances that can occur due to the presence of shadows or occluding objects. All these virtues, make the proposed detector appropriate for UAV navigation. Besides the synthetic experiments that will demonstrate the theoretical benefits of our formulation, we will show applications for detecting landing areas in regions with highly repetitive patterns, and specific objects under the presence of cast shadows or sudden camera motions. Michael Villamizar, Alberto Sanfeliu, Francesc Moreno-Noguer |
ICRA | 2 |
| 2014 | Proactive kinodynamic planning using the Extended Social Force Model and human motion prediction in urban environmentsabstractThis paper presents a novel approach for robot navigation in crowded urban environments where people and objects are moving simultaneously while a robot is navigating. Avoiding moving obstacles at their corresponding precise moment motivates the use of a robotic planner satisfying both dynamic and nonholonomic constraints, also referred as kynodynamic constraints.We present a proactive navigation approach with respect its environment, in the sense that the robot calculates the reaction produced by its actions and provides the minimum impact on nearby pedestrians. As a consequence, the proposed planner integrates seamlessly planning and prediction and calculates a complete motion prediction of the scene for each robot propagation. Making use of the Extended Social Force Model (ESFM) allows an enormous simplification for both the prediction model and the planning system under differential constraints. Simulations and real experiments have been carried out to demonstrate the success of the proactive kinodynamic planner. Gonzalo Ferrer 0001, Alberto Sanfeliu |
IROS | 2 |
| 2014 | Bayesian Human Motion Intentionality Prediction in urban environments
Gonzalo Ferrer 0001, Alberto Sanfeliu |
Pattern Recognit. Lett. | 2 |
| 2013 | Dense Segmentation-Aware DescriptorsabstractIn this work we exploit segmentation to construct appearance descriptors that can robustly deal with occlusion and background changes. For this, we downplay measurements coming from areas that are unlikely to belong to the same region as the descriptor's center, as suggested by soft segmentation masks. Our treatment is applicable to any image point, i.e. dense, and its computational overhead is in the order of a few seconds. We integrate this idea with Dense SIFT, and also with Dense Scale and Rotation Invariant Descriptors (SID), delivering descriptors that are densely computable, invariant to scaling and rotation, and robust to background changes. We apply our approach to standard benchmarks on large displacement motion estimation using SIFT-flow and wide-baseline stereo, systematically demonstrating that the introduction of segmentation yields clear improvements. Eduard Trulls, Iasonas Kokkinos, Alberto Sanfeliu, Francesc Moreno-Noguer |
CVPR | 3 |
| 2013 | Solving Perception Uncertainty Problems in Robotics
Alberto Sanfeliu |
ICPRAM | 1 |
| 2013 | Robot companion: A social-force based approach with human awareness-navigation in crowded environmentsabstractRobots accompanying humans is one of the core capacities every service robot deployed in urban settings should have. We present a novel robot companion approach based on the so-called Social Force Model (SFM). A new model of robot-person interaction is obtained using the SFM which is suited for our robots Tibi and Dabo. Additionally, we propose an interactive scheme for robot's human-awareness navigation using the SFM and prediction information. Moreover, we present a new metric to evaluate the robot companion performance based on vital spaces and comfortableness criteria. Also, a multimodal human feedback is proposed to enhance the behavior of the system. The validation of the model is accomplished throughout an extensive set of simulations and real-life experiments. Gonzalo Ferrer 0001, Anais Garrell, Alberto Sanfeliu |
IROS | 3 |
| 2013 | Proactive behavior of an autonomous mobile robot for human-assisted learningabstractDuring the last decade, there has been a growing interest in making autonomous social robots able to interact with people. However, there are still many open issues regarding the social capabilities that robots should have in order to perform these interactions more naturally. In this paper we present the results of several experiments conducted at the Barcelona Robot Lab in the campus of the “Universitat Politècnica de Catalunya” in which we have analyzed different important aspects of the interaction between a mobile robot and nontrained human volunteers. First, we have proposed different robot behaviors to approach a person and create an engagement with him/her. In order to perform this task we have provided the robot with several perception and action capabilities, such as that of detecting people, planning an approach and verbally communicating its intention to initiate a conversation. Once the initial engagement has been created, we have developed further communication skills in order to let people assist the robot and improve its face recognition system. After this assisted and online learning stage, the robot becomes able to detect people under severe changing conditions, which, in turn enhances the number and the manner that subsequent human-robot interactions are performed. Anais Garrell, Michael Villamizar, Francesc Moreno-Noguer, Alberto Sanfeliu |
RO-MAN | 4 |
| 2012 | Spatiotemporal Descriptor for Wide-Baseline Stereo Reconstruction of Non-rigid and Ambiguous Scenes
Eduard Trulls, Alberto Sanfeliu, Francesc Moreno-Noguer |
ECCV (3) | 2 |
| 2012 | Work in progress: A constructivist didactic methodology for a humanoid robotics workshopabstractThis paper presents a constructivist methodology oriented to training Robotics in its wide sense, and its implementation in higher education. A humanoid robotics workshop is included as a part of the curriculum of `Industrial Robotics' in the Industrial Engineering. This workshop covers one third of the laboratory practices of this subject. A constructivist view for learning is adopted, whereby robotic technologies are not seen as mere tools, but rather as potential vehicles of new ways of thinking about teaching, learning and education at large. Students, in a constructivist learning environment, are invited to work on experiments and authentic problem-solving; with selective use of available resources according to their own interests, research and learning strategies. The authors, as trainers of this workshop, chose ROBONOVA humanoid robots, among other devices, which attempt to partner technology with ideas of constructivism. The materials used in the workshop offer building parts, sensors thus connecting a robot with the external environment and programming software with a simple graphical interface intended for the creation of robot behavioral movements. The idea is “learning by design”, which is central in the constructivist pedagogy introduced firstly by Resnick. In this workshop we corroborate this idea through the project-based learning approach in robotics. Learning tasks of the workshop are organized as projects (small and large) that encourage students to develop their own designs. Projects are either instructor led or completely arisen from students. Alexandre Miranda Añon, Yolanda Bolea, Antoni Grau-Saldes, Alberto Sanfeliu |
FIE | 4 |
| 2012 | Probabilistic invariant image representation and associated distance measure
Jorge Scandaliaris, Alberto Sanfeliu |
ICPR | 2 |
| 2012 | Online human-assisted learning using Random Ferns
Michael Villamizar, Anais Garrell, Alberto Sanfeliu, Francesc Moreno-Noguer |
ICPR | 3 |
| 2012 | On the Graph Edit Distance Cost: Properties and ApplicationsabstractWe model the edit distance as a function in a labeling space. A labeling space is an Euclidean space where coordinates are the edit costs. Through this model, we define a class of cost. A class of cost is a region in the labeling space that all the edit costs have the same optimal labeling. Moreover, we characterize the distance value through the labeling space. This new point of view of the edit distance gives us the opportunity of defining some interesting properties that are useful for a better understanding of the edit distance. Finally, we show the usefulness of these properties through some applications. Albert Solé-Ribalta, Francesc Serratosa, Alberto Sanfeliu |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2012 | Bootstrapping Boosted Random Ferns for discriminative and efficient object classification
Michael Villamizar, Juan Andrade-Cetto, Alberto Sanfeliu, Francesc Moreno-Noguer |
Pattern Recognit. | 3 |
| 2011 | Efficient 3D Object Detection using Multiple Pose-Specific ClassifiersabstractWe propose an efficient method for object localization and 3D pose estimation. A two-step approach is used. In the first step, a pose estimator is evaluated in the input images in order to estimate potential object locations and poses. These candidates are then validated, in the second step, by the corresponding pose-specific classifier. The result is a detection approach that avoids the inherent and expensive cost of testing the complete set of specific classifiers over the entire image. A further speedup is achieved by feature sharing. Features are computed only once and are then used for evaluating the pose estimator and all specific classifiers. The proposed method has been validated on two public datasets for the problem of detecting of cars under several views. The results show that the proposed approach yields high detection rates while keeping efficiency. © 2011. The copyright of this document resides with its authors. Michael Villamizar, Helmut Grabner, Francesc Moreno-Noguer, Juan Andrade-Cetto, Luc Van Gool, Alberto Sanfeliu |
BMVC | 6 |
| 2011 | Comparative analysis of human motion trajectory prediction using minimum variance curvatureabstractThe prediction of human motion intentionality is a key issue towards intelligent human robot interaction and robot navigation. In this work we present a comparative study of several prediction functions that are based on the minimum curvature variance from the current position to all the potential destination points, that means, the points that are relevant for people motion intentionality. The proposed predictor computes, at each interval of time, the trajectory from the present to the destination positions, and makes a prediction of the human motion at each interval of time using only the criterion of minimum curvature variation. The method has been validated in the Edinburgh Informatics Forum Pedestrian database. Gonzalo Ferrer 0001, Alberto Sanfeliu |
HRI | 2 |
| 2010 | Efficient rotation invariant object detection using boosted Random FernsabstractWe present a new approach for building an efficient and robust classifier for the two class problem, that localizes objects that may appear in the image under different orientations. In contrast to other works that address this problem using multiple classifiers, each one specialized for a specific orientation, we propose a simple two-step approach with an estimation stage and a classification stage. The estimator yields an initial set of potential object poses that are then validated by the classifier. This methodology allows reducing the time complexity of the algorithm while classification results remain high. The classifier we use in both stages is based on a boosted combination of Random Ferns over local histograms of oriented gradients (HOGs), which we compute during a preprocessing step. Both the use of supervised learning and working on the gradient space makes our approach robust while being efficient at run-time. We show these properties by thorough testing on standard databases and on a new database made of motorbikes under planar rotations, and with challenging conditions such as cluttered backgrounds, changing illumination conditions and partial occlusions. Michael Villamizar, Francesc Moreno-Noguer, Juan Andrade-Cetto, Alberto Sanfeliu |
CVPR | 4 |
| 2010 | Guiding and regrouping people missions in urban areas using cooperative Multi-Robot Task Allocation
Anais Garrell, Josep M. Mirats Tur, Oscar Sandoval Torres, Alberto Sanfeliu |
ETFA | 4 |
| 2010 | Computing the Barycenter Graph by Means of the Graph Edit DistanceabstractThe barycenter graph has been shown as an alternative to obtain the representative of a given set of graphs. In this paper we propose an extension of the original algorithm which makes use of the graph edit distance in conjunction with the weighted mean of a pair of graphs. Our main contribution is that we can apply the method to attributed graphs with any kind of labels in both the nodes and the edges, equipped with a distance function less constrained than in previous approaches. Experiments done on four different datasets support the validity of the method giving good approximations of the barycenter graph. Itziar Bardají, Miquel Ferrer, Alberto Sanfeliu |
ICPR | 3 |
| 2010 | A Conductance Electrical Model for Representing and Matching Weighted Undirected GraphsabstractIn this paper we propose a conductance electrical model to represent weighted undirected graphs that allows us to efficiently compute approximate graph isomorphism in large graphs. The model is built by transforming a graph into an electrical circuit. Edges in the graph become conductances in the electrical circuit. This model follows the laws of the electrical circuit theory and we can potentially use all the existing theory and tools of this field to derive other approximate techniques for graph matching. In the present work, we use the proposed circuital model to derive approximated graph isomorphism solutions. Manuel Igelmo, Alberto Sanfeliu, Miquel Ferrer |
ICPR | 2 |
| 2010 | Discriminant and Invariant Color Model for Tracking under Abrupt Illumination ChangesabstractThe output from a color imaging sensor, or apparent color, can change considerably due to illumination conditions and scene geometry changes. In this work we take into account the dependence of apparent color with illumination an attempt to find appropriate color models for the typical conditions found in outdoor settings. We evaluate three color based trackers, one based on hue, another based on an intrinsic image representation and the last one based on a proposed combination of a chromaticity model with a physically reasoned adaptation of the target model. The evaluation is done on outdoor sequences with challenging illumination conditions, and shows that the proposed method improves the average track completeness by over 22% over the hue-based tracker and the closeness of track by over 7% over the tracker based on the intrinsic image representation. Jorge Scandaliaris, Alberto Sanfeliu |
ICPR | 2 |
| 2010 | Comparative Analysis for Detecting Objects Under Cast Shadows in Video ImagesabstractCast shadows add additional difficulties on detecting objects because they locally modify image intensity and color. Shadows may appear or disappear in an image when the object, the camera, or both are free to move through a scene. This work evaluates the performance of an object detection method based on boosted HOG paired with three different image representations in outdoor video sequences. We follow and extend on the taxonomy from van de Sande with considerations on the constraints assumed by each descriptor on the spatial variation of the illumination. We show that the intrinsic image representation consistently gives the best results. This proves the usefulness of this representation for object detection in varying illumination conditions, and supports the idea that in practice local assumptions in the descriptors can be violated. Jorge Scandaliaris, Michael Villamizar, Alberto Sanfeliu |
ICPR | 3 |
| 2010 | Shared Random Ferns for Efficient Detection of Multiple CategoriesabstractWe propose a new algorithm for detecting multiple object categories that exploits the fact that different categories may share common features but with different geometric distributions. This yields an efficient detector which, in contrast to existing approaches, considerably reduces the computation cost at runtime, where the feature computation step is traditionally the most expensive. More specifically, at the learning stage we compute common features by applying the same Random Ferns over the Histograms of Oriented Gradients on the training images. We then apply a boosting step to build discriminative weak classifiers, and learn the specific geometric distribution of the Random Ferns for each class. At runtime, only a few Random Ferns have to be densely computed over each input image, and their geometric distribution allows performing the detection. The proposed method has been validated in public datasets achieving competitive detection results, which are comparable with state-of-the-art methods that use specific features per class. Michael Villamizar, Francesc Moreno-Noguer, Juan Andrade-Cetto, Alberto Sanfeliu |
ICPR | 4 |
| 2010 | Local optimization of cooperative robot movements for guiding and regrouping people in a guiding missionabstractThis article presents a novel approach for optimizing locally the work of cooperative robots and obtaining the minimum displacement of humans in a guiding people mission. Unlike other methods, we consider situations where individuals can move freely and can escape from the formation, moreover they must be regrouped by multiple mobile robots working cooperatively. The problem is addressed by introducing a “Discrete Time Motion” model (DTM) and a new cost function that minimizes the work required by robots for leading and regrouping people. The guiding mission is carried out in urban areas containing multiple obstacles and building constraints. Furthermore, an analysis of forces actuating among robots and humans is presented throughout simulations of different situations of robot and human configurations and behaviors. Anais Garrell, Alberto Sanfeliu |
IROS | 2 |
| 2010 | Model validation: Robot behavior in people guidance mission using DTM model and estimation of human motion behaviorabstractThis paper describes the validation process of a simulation model that have been used to explore the new possibilities of interaction when humans are guided by teams of robots that work cooperatively in urban areas. The set of experiments, which have been recorded as video sequences, show a group of people being guided by a team of three people (who play the role of the guide robots). The model used in the simulation process is called Discrete Time Motion model (DTM) described in [7], where the environment is modeled using a set of potential fields, and people's motion is represented through tension functions. The video sequences were recorded in an urban space of 10:000 m2denominated Barcelona Robot Lab, where people move in the urban space following diverse trajectories. The motion (pose and velocity) of people and robots extracted from the video sequences were compared against the predictions of the DTM model. Finally, we checked the proper functioning of the model by studying the position error differences of the recorded and simulated sequences. Anais Garrell, Alberto Sanfeliu |
IROS | 2 |
| 2010 | Autonomous navigation for urban service mobile robotsabstractWe present to the robotic community a fully autonomous navigation solution for mobile robots operating in urban pedestrian areas. We introduce our robots and the experimental zone, overview the architecture of the navigation framework, and present the results after 3.5km of autonomous navigation. We expose the main lessons learnt by the scientific team and identify the issues to improve future works. Andreu Corominas Murtra, Eduard Trulls, Oscar Sandoval Torres, Joan Pérez-Ibarz, Dizan Vasquez, Josep M. Mirats Tur, Miquel Ferrer, Alberto Sanfeliu |
IROS | 8 |
| 2010 | Action Selection for Single-Camera SLAMabstractA method for evaluating, at video rate, the quality of actions for a single camera while mapping unknown indoor environments is presented. The strategy maximizes mutual information between measurements and states to help the camera avoid making ill-conditioned measurements that are appropriate to lack of depth in monocular vision systems. Our system prompts a user with the appropriate motion commands during 6-DOF visual simultaneous localization and mapping with a handheld camera. Additionally, the system has been ported to a mobile robotic platform, thus closing the control-estimation loop. To show the viability of the approach, simulations and experiments are presented for the unconstrained motion of a handheld camera and for the motion of a mobile robot with nonholonomic constraints. When combined with a path planner, the technique safely drives to a marked goal while, at the same time, producing an optimal estimated map. Teresa Vidal-Calleja, Alberto Sanfeliu, Juan Andrade-Cetto |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2009 | Combining color-based invariant gradient detector with HoG descriptors for robust image detection in scenes under cast shadowsabstractIn this work we present a robust detection method in outdoor scenes under cast shadows using color based invariant gradients in combination with HoG local features. The method achieves good detection rates in urban scene classification and person detection outperforming traditional methods based on intensity gradient detectors which are sensible to illumination variations but not to cast shadows. The method uses color based invariant gradients that emphasize material changes and extract relevant and invariant features for detection while neglecting shadow contours. This method allows to train and detect objects and scenes independently of scene illumination, cast and self shadows. Moreover, it allows to do training in one shot, that is, when the robot visits the scene for the first time. Michael Villamizar, Jorge Scandaliaris, Alberto Sanfeliu, Juan Andrade-Cetto |
ICRA | 3 |
| 2009 | Discrete time motion model for guiding people in urban areas using multiple robotsabstractWe present a new model for people guidance in urban settings using several mobile robots, that overcomes the limitations of existing approaches, which are either tailored to tightly bounded environments, or based on unrealistic human behaviors. Although the robots motion is controlled by means of a standard particle filter formulation, the novelty of our approach resides in how the environment and human and robot motions are modeled. In particular we define a ¿Discrete-Time- Motion¿ model, which from one side represents the environment by means of a potential field, that makes it appropriate to deal with open areas, and on the other hand the motion models for people and robots respond to realistic situations, and for instance human behaviors such as ¿leaving the group¿ are considered. Anais Garrell, Alberto Sanfeliu, Francesc Moreno-Noguer |
IROS | 2 |
| 2009 | Integrating asynchronous observations for mobile robot position tracking in cooperative environmentsabstractThis paper presents an asynchronous particle filter algorithm for mobile robot position tracking, taking into account time considerations when integrating observations being delayed or advanced from the prior estiamate time point. The interest of that filter lies in cooperative environments and in fast vehicles. The paper studies the first case, where a sensor network shares perception data with running robots that receive accurate obeservations with large delays due to acquisition, processing and wireless communications. Promising simulated results comparing a basic particle filter and the proposed one are shown. The paper also investigates a situation where a robot is tracking its position, fusing only odometry and observations from a camera network partially covering the robot path. Andreu Corominas Murtra, Josep M. Mirats Tur, Alberto Sanfeliu |
IROS | 3 |
| 2008 | Efficient active global localization for mobile robots operating in large and cooperative environmentsabstractThis paper presents a novel and efficient framework to the active map-based global localization problem for mobile robots operating in large and cooperative environments. The paper proposes a rational criteria to select the action that minimizes the expected number of remaining position hypotheses, for the single robot case and for the cooperative case, where the lost robot takes advantage of observations coming from a sensor network deployed on the environment or from other localized robots. Efficiency in time complexity is achieved thanks to reasoning in terms of the number of hypotheses instead of in terms of the belief function. Simulation results in a real outdoor environment of 10.000m2are presented validating the presented approach and showing different behaviours for the single robot case and for the cooperative one. Andreu Corominas Murtra, Josep M. Mirats Tur, Alberto Sanfeliu |
ICRA | 3 |
| 2008 | Dependent Multiple Cue Integration for Robust TrackingabstractWe propose a new technique for fusing multiple cues to robustly segment an object from its background in video sequences that suffer from abrupt changes of both illumination and position of the target. Robustness is achieved by the integration of appearance and geometric object features and by their estimation using Bayesian filters, such as Kalman or particle filters. In particular, each filter estimates the state of a specific object feature, conditionally dependent on another feature estimated by a distinct filter. This dependence provides improved target representations, permitting to segment it out from the background even in non-stationary sequences. Considering that the procedure of the Bayesian filters may be described by a "hypotheses generation--hypotheses correction" strategy, the major novelty of our methodology compared to previous approaches is that the mutual dependence between filters is considered during the feature observation, i.e, into the "hypotheses correction" stage,instead of considering it when generating the hypotheses. This proves to be much more effective in terms of accuracy and reliability. The proposed method is analytically justified and applied to develop a robust tracking system that adapts online and simultaneously the color space where the image points are represented, the color distributions, the contour of the object and its bounding box. Results with synthetic data and real video sequences demonstrate the robustness and versatility of our method. Francesc Moreno-Noguer, Alberto Sanfeliu, Dimitris Samaras |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Pattern recognition in interdisciplinary perception and intelligence
Antonio Fernández-Caballero 0001, Alberto Sanfeliu, Yoshiaki Shirai |
Pattern Recognit. Lett. | 2 |
| 2007 | Robust Color Contour Object Detection Invariant to Shadows
Jorge Scandaliaris, Michael Villamizar, Juan Andrade-Cetto, Alberto Sanfeliu |
CIARP | 4 |
| 2007 | A New Algorithm to Compute the Distance Between Multi-dimensional Histograms
Francesc Serratosa, Gerard Sanroma, Alberto Sanfeliu |
CIARP | 3 |
| 2007 | On the Observability of Bearing-only SLAMabstractIn this paper we present an observability analysis for a mobile robot performing SLAM with a single monocular camera. The aim is to get a better understanding of the well known intuitive behavior of these systems, such as the need for triangulation to features from different positions in order to get accurate relative pose estimates. The characterisation of the unobservable directions is made using the nullspace basis of the stripped observability matrix. This allows us to identify which vehicle motions are required to maximise the number of observable states in the system, which in turn affects accuracy in the estimation process. The analysis is performed by modelling the system in the continuous time domain as piecewise constant. Simulation results using an extended information filter are shown to verify the results of the observability analysis. Teresa Vidal-Calleja, Mitch Bryson, Salah Sukkarieh, Alberto Sanfeliu, Juan Andrade-Cetto |
ICRA | 4 |
| 2007 | Vision-based loop closing for delayed state robot mappingabstractThis paper shows results on outdoor vision-based loop closing for simultaneous localization and mapping. Our experiments show that for loops of over 50 m, the pose estimates maintained with a delayed-state extended information filter are consistent enough to guarantee assertion of vision- based pose constraints for loop closure, provided no necessary information links are added to the estimator. The technique computes relative pose constraints via a robust least squares minimization of 3D point correspondences, which are in turn obtained from the matching of SIFT features over candidate image pairs. We propose a loop closure test that checks both for closeness of means and for highly informative updates at the same time. Viorela Ila, Juan Andrade-Cetto, Rafael Valencia, Alberto Sanfeliu |
IROS | 4 |
| 2007 | Integration of deformable contours and a multiple hypotheses Fisher color model for robust tracking in varying illuminant environments
Francesc Moreno-Noguer, Alberto Sanfeliu, Dimitris Samaras |
Image Vis. Comput. | 2 |
| 2006 | Orientation Invariant Features for Multiclass Object Recognition
Michael Villamizar, Alberto Sanfeliu, Juan Andrade-Cetto |
CIARP | 2 |
| 2006 | Integration of Dependent Bayesian Filters for Robust TrackingabstractRobotics applications based on computer vision algorithms are highly constrained to indoor environments where conditions may be controlled. The development of robust visual algorithms is necessary for improving the capabilities of many autonomous systems in outdoor and dynamic environments. In particular, this paper proposes a tracking algorithm robust to several artifacts which may be found in real world applications, such as lighting changes, cluttered backgrounds and unexpected target movements. In order to deal with these difficulties the proposed tracking methodology integrates several Bayesian filters. Each filter estimates the state of a particular object feature which is conditionally dependent on another feature estimated by a distinct filter. This dependence provides improved representations of the target, allowing to segment it out from the background of the image. We describe the updating procedure of the Bayesian filters by a 'hypotheses generation and correction' scheme. The main difference with respect to previous approaches is that the dependence between filters is considered during the feature observation, i.e., into the 'hypotheses correction' stage, instead of considering it when generating the hypotheses. This proves to be much more effective in terms of accuracy and reliability Francesc Moreno-Noguer, Alberto Sanfeliu, Dimitris Samaras |
ICRA | 2 |
| 2006 | Signatures versus histograms: Definitions, distances and algorithms
Francesc Serratosa, Alberto Sanfeliu |
Pattern Recognit. | 2 |
| 2005 | A Fast Distance Between Histograms
Francesc Serratosa, Alberto Sanfeliu |
CIARP | 2 |
| 2005 | Integration of Conditionally Dependent Object Features for Robust Figure/Background SegmentationabstractWe propose a new technique for focusing multiple cues to robustly segment an object from its background in video sequences that suffer from abrupt changes of both illumination and position of the target. Robustness is achieved by tile integration of appearance and geometric object features and by their description using particle filters. Previous approaches assume independence of the object cues or apply the particle filter formulation to only one of the features, and assume a smooth change in the rest, which can prove is very limiting, especially when the state of some features needs to be updated using other cues or when their dynamics follow non-linear and unpredictable paths. Our technique offers a general framework to model the probabilistic relationship between features. The proposed method is analytically justified and applied to develop a robust tracking system that adapts online and simultaneously the color space where the image points are represented, the color distributions, and the contour of the object. Results with synthetic data and real video sequences demonstrate the robustness and versatility of our method Francesc Moreno-Noguer, Alberto Sanfeliu, Dimitris Samaras |
ICCV | 2 |
| 2005 | Robot Perception for Navigation in Indoor Buildings
Alberto Sanfeliu |
ICINCO | 1 |
| 2005 | Unscented Transformation of Vehicle States in SLAMabstractIn this article we propose an algorithm to reduce the effects caused by linearization in the typical EKF approach to SLAM. The technique consists in computing the vehicle prior using an Unscented Transformation. The UT allows a better nonlinear mean and variance estimation than the EKF. There is no need however in using the UT for the entire vehicle-map state, given the linearity in the map part of the model. By applying the UT only to the vehicle states we get more accurate covariance estimates. The a posteriori estimation is made using a fully observable EKF step, thus preserving the same computational complexity as the EKF with sequential innovation. Experiments over a standard SLAM data set show the behavior of the algorithm. Juan Andrade-Cetto, Teresa Vidal-Calleja, Alberto Sanfeliu |
ICRA | 3 |
| 2005 | Object and image indexing based on region connection calculus and oriented matroid theory
Ernesto Staffetti, Antoni Grau-Saldes, Francesc Serratosa, Alberto Sanfeliu |
Discret. Appl. Math. | 4 |
| 2005 | An approach of visual motion analysis
Alberto Sanfeliu, Juan José Villanueva |
Pattern Recognit. Lett. | 1 |
| 2005 | The Effects of Partial Observability When Building Fully Correlated MapsabstractThis paper presents an analysis of the fully correlated approach to the simultaneous localization and map building (SLAM) problem from a control systems theory point of view, both for linear and nonlinear vehicle models. We show how partial observability hinders full reconstructibility of the state space, making the final map estimate dependent on the initial observations. Nevertheless, marginal filter stability guarantees convergence of the state error covariance to a positive semidefinite covariance matrix. By characterizing the form of the total Fisher information, we are able to determine the unobservable state space directions. Moreover, we give a closed-form expression that links the amount of reconstruction error to the number of landmarks used. The analysis allows the formulation of measurement models that make SLAM observable. Juan Andrade-Cetto, Alberto Sanfeliu |
IEEE Trans. Robotics | 2 |
| 2004 | Structural Pattern Recognition for Industrial Machine Sounds Based on Frequency Spectrum Analysis
Yolanda Bolea, Antoni Grau-Saldes, Arthur Pelissier, Alberto Sanfeliu |
CIARP | 4 |
| 2004 | Adaptive Color Model for Figure-Ground Segmentation in Dynamic Environments
Francesc Moreno-Noguer, Alberto Sanfeliu |
CIARP | 2 |
| 2004 | A Color Constancy Algorithm for the Robust Description of Images Collected from a Mobile Robot
Jaume Vergés-Llahí, Alberto Sanfeliu |
CIARP | 2 |
| 2004 | The Effects of Partial Observability in SLAMabstractIn this article, we show that partial observability hinders full reconstructibility of the state space in SLAM, making the final map estimate dependent on the initial observations, and not guaranteeing convergence to a positive semi-definite covariance matrix. By characterizing the form of the total Fisher information we are able to determine the unobservable state space directions. To overcome this problem, we formulate new fully observable measurement models that make SLAM stable. Juan Andrade-Cetto, Alberto Sanfeliu |
ICRA | 2 |
| 2004 | Conditions for suboptimal filter stability in SLAMabstractIn this article, we show marginal stability in SLAM, guaranteeing convergence to a non-zero mean state error estimate bounded by a constant value. Moreover, marginal stability guarantees also convergence of the Riccati equation of the one-step ahead state error covariance to at least one psd steady state solution. In the search for real-time implementations of SLAM, covariance inflation methods produce a suboptimal filter that eventually may lead to the computation of an unbounded state error covariance. We provide tight constraints in the amount of decorrelation possible, to guarantee convergence of the state error covariance, and at the same time, a linear-time implementation of SLAM. Teresa Vidal-Calleja, Juan Andrade-Cetto, Alberto Sanfeliu |
IROS | 3 |
| 2004 | Second-Order Random Graphs For Modeling Sets Of Attributed Graphs And Their Application To Object Learning And RecognitionabstractThe aim of this article is to present a random graph representation, that is based on second-order relations between graph elements, for modeling sets of attributed graphs (AGs). We refer to these models as Second-Order Random Graphs (SORGs). The basic feature of SORGs is that they include both marginal probability functions of graph elements and second-order joint probability functions. This allows a more precise description of both the structural and semantic information contents in a set of AGs and, consequently, an expected improvement in graph matching and object recognition. The article presents a probabilistic formulation of SORGs that includes as particular cases the two previously proposed approaches based on random graphs, namely the First-Order Random Graphs (FORGs) and the Function-Described Graphs (FDGs). We then propose a distance measure derived from the probability of instantiating a SORG into an AG and an incremental procedure to synthesize SORGs from sequences of AGs. Finally, SORGs are shown to improve the performance of FORGs, FDGs and direct AG-to-AG matching in three experimental recognition tasks: one in which AGs are randomly generated and the other two in which AGs represent multiple views of 3D objects (either synthetic or real) that have been extracted from color images. In the last case, object learning is achieved through the synthesis of SORG models. Alberto Sanfeliu, Francesc Serratosa, René Alquézar |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2003 | Non-speech Sound Feature Extraction Based on Model Identification for Robot Navigation
Yolanda Bolea, Antoni Grau-Saldes, Alberto Sanfeliu |
CIARP | 3 |
| 2003 | Robot Vision for Autonomous Object Learning and Tracking
Alberto Sanfeliu |
CIARP | 1 |
| 2003 | A Colour Constancy Algorithm Based on the Histogram of Feasible Colour Mappings
Jaume Vergés-Llahí, Alberto Sanfeliu |
CIARP | 2 |
| 2003 | Learning and recognising 3D models represented by multiple views by means of methods based on random graphsabstractThe aim of this article is to describe and compare the methods based on random graphs (RGs) which are applied to learn and recognize 3D objects represented by multiple views. These methods are based on modelling the objects by means of probabilistic structures that keep 1/sup st/ and 2/sup nd/-order probabilities. That is, multiple views of a 3D object are represented by few RGs. The most important probabilistic structures presented in the literature are first-order random graphs (FORGs), function-described graphs (FDGs) and second-order random graphs(SORGs). In the learning process, each one of the 3D-object views are represented by an attribute graph (AC), and a group of AGs are synthesized in a RG. In the recognizing process, the view of the object is represented by an AG and then it is compared with the RG that model each one of the 3D-object prototypes. In this paper, it is explained the modelling of the 3D-objects and the methods of learning and recognition based on FORGs, FDGs and SORGs. We show some results of the methods for real 3D objects. Alberto Sanfeliu, Francesc Serratosa |
ICIP (2) | 1 |
| 2003 | Temporal landmark validation in CabstractCurrent techniques to concurrent map building and localization (CML) have been devised for static environments, and lack robustness in more realistic situations. In this communication we provide new ideas that extend the typical stochastic estimation approach to CML, to take into account the dynamics of the environment. The basic idea consists on using the history of data association mismatches for the computation of the likelihood of future data association. The incorporation of a novel temporal landmark quality test, together with the spatial compatibility tests already available, help alleviate the difficulty of data association. We propose a pair of temporal landmark quality functions to aid in those situations in which landmark observations might not be consistent in time; and show how by incorporating these functions, the overall estimation-theoretic approach to CML is improved. Special attention is paid in that the removal of landmarks from the map does not violate the basic convergence properties of the localization and map building algorithms already described in the literature. Namely, asymptotic convergence and full correlation. Juan Andrade-Cetto, Alberto Sanfeliu |
ICRA | 2 |
| 2003 | Function-described graphs for modelling objects represented by sets of attributed graphs
Francesc Serratosa, René Alquézar, Alberto Sanfeliu |
Pattern Recognit. | 3 |
| 2002 | Concurrent Map Building and Localization on Indoor Dynamic EnvironmentsabstractA system that builds and maintains a dynamic map for a mobile robot is presented. A learning rule associated to each observed landmark is used to compute its robustness. The position of the robot during map construction is estimated by combining sensor readings, motion commands, and the current map state by means of an Extended Kalman Filter. The combination of landmark strength validation and Kalman filtering for map updating and robot position estimation allows for robust learning of moderately dynamic indoor environments. Juan Andrade-Cetto, Alberto Sanfeliu |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2002 | Synthesis of Function-Described Graphs and Clustering of Attributed GraphsabstractFunction-Described Graphs (FDGs) have been introduced by the authors as a representation of an ensemble of Attributed Graphs (AGs) for structural pattern recognition alternative to first-order random graphs. Both optimal and approximate algorithms for error-tolerant graph matching, which use a distance measure between AGs and FDGs, have been reported elsewhere. In this paper, both the supervised and the unsupervised synthesis of FDGs from a set of graphs is addressed. First, two procedures are described to synthesize an FDG from a set of commonly labeled AGs or FDGs, respectively. Then, the unsupervised synthesis of FDGs is studied in he context of clustering a set of AGs and obtaining an FDG model for each cluster. Two algorithms based on incremental and hierarchical clustering, respectively, are proposed, which are parameterized by a graph matching method. Some experimental results both on synthetic data and a real 3D-object recognition application show that the proposed algorithms are effective for clustering a set of AGs and synthesizing the FDGs that describe the classes. Moreover, the synthesized FDGs are shown to be useful for pattern recognition thanks to the distance measure and matching algorithm previously reported. Francesc Serratosa, René Alquézar, Alberto Sanfeliu |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2002 | Graph-based representations and techniques for image processing and image analysis
Alberto Sanfeliu, René Alquézar, J. Andrade, Joan Climent, Francesc Serratosa, Jaume Vergés-Llahí |
Pattern Recognit. | 1 |
| 2001 | Localization of human faces fusing color segmentation and depth from stereoabstractDescribes a method to localize faces in color images based on the fusion of the information gathered from a stereo vision system and the analysis of color images. Our method generates a depth map of the scene and tries to fit a head model taking into account the shape of the model and skin color information. The method is tailored for its use in factory automation applications where the detection and localization of humans is necessary for the completion or interruption of a particular task, such as robot manipulator safety or the interaction of service robots with humans. Francesc Moreno-Noguer, Juan Andrade-Cetto, Alberto Sanfeliu |
ETFA (2) | 3 |
| 2000 | Integration of Perceptual Grouping and DepthabstractDifferent data acquisition methods are tailored at extracting particular characteristics from a scene and by combining their results a more robust scene description can be created. A method to fuse perceptual groupings extracted from color-based segmentation and depth information from stereo using supervised classification is presented. The merging of data from these two acquisition modules allows for a spatially coherent blend of smooth regions and detail in an image. Depth cues are used to limit the area of interest in the scene and to improve perceptual grouping solving subsegmentation and oversegmentation of the original images. The complexity of the algorithm does not exceed that of the individual acquisition modules. The resulting scene description can then be fed to an object recognition modules for scene interpretation. Juan Andrade-Cetto, Alberto Sanfeliu |
ICPR | 2 |
| 2000 | Clustering of Attributed Graphs and Unsupervised Synthesis of Function-Described GraphsabstractFunction-described graphs (FDGs) have been introduced by the authors as a representation of an ensemble of attributed graphs (AGs) for structural pattern recognition as an alternative to first-order random graphs. The unsupervised synthesis of FDGs is studied in the context of clustering a set of AGs and obtaining an FDG model for each cluster. Two algorithms based on incremental and hierarchical clustering, respectively, are proposed, which are parameterized by a graph matching method. Results on 3D object recognition show that these algorithms are effective for clustering a set of AGs and synthesising the FDGs that describe the classes. Alberto Sanfeliu, René Alquézar, Francesc Serratosa |
ICPR | 1 |
| 2000 | Efficient Algorithms for Matching Attributed Graphs and Function-Described GraphsabstractFunction-described graphs (FDG) have been introduced by the authors as a representation of an ensemble of attributed graphs (AG) for structural pattern recognition alternative to first-order random graphs. In previous works, algorithms for the synthesis of FDG and a branch-and-bound algorithm for error-tolerant graph matching, which computes a distance measure between AG and FDG, have been reported. Since the worst-case complexity of that matching algorithm is exponential in the number of nodes, an approximate algorithm to compute a suboptimal measure is proposed in this paper. Results in 3D-object recognition show that, although the computational time is reduced, there is only a slight decrease of effectiveness while classifying an AG against a set of FDG. Francesc Serratosa, René Alquézar, Alberto Sanfeliu |
ICPR | 3 |
| 2000 | Color Image Segmentation Solving Hard-Constraints on Graph Partitioning Greedy AlgorithmsabstractA graph partitioning greedy algorithm is presented. This algorithm avoids the hard-constraints of others similar approaches such as the impossibility for some regions to grow after certain step of the algorithm and the uniqueness of the solution. Nevertheless, it allows attaining global results by local approximations using a generalised concept of not over-segmentation, which includes an energy function, and eliminating the not sub-segmentation criterion using a probabilistic criterion similar to that of annealing. The high-variability region problems such as borders are also eliminated identifying them and distributing their pixels among the other neighbour regions. Thus, it is possible to keep the time complexity of usual graph partitioning greedy algorithm and avoiding its high variability region problems, obtaining better results. Jaume Vergés-Llahí, Alberto Sanfeliu, Joan Climent |
ICPR | 2 |
| 1998 | Low cost architecture for structure measure distance computationabstractHuge and expensive computation resources are usually required to perform graph labelling at high speed. This fact restricts an extensive use of this methodology in industrial applications such as visual inspection. A new systolic architecture is presented which computes structural distances between cliques of different graphs based on a modified incremental Levenshtein distance algorithm. The distances obtained are used as a support function for graph labelling using probabilistic relaxation techniques. The proposed architecture computes the distances between k input cliques of an input graph and one reference clique of a reference graph. It does not limit the number of cliques nor cliques complexity of the input graph, so any input graph can be labelled. A low cost solution has been implemented based on FPGAs. Joan Aranda, Joan Climent, Antoni Grau-Saldes, Alberto Sanfeliu |
ICPR | 4 |
| 1997 | Recognition and learning of a class of context-sensitive languages described by augmented regular expressions
René Alquézar, Alberto Sanfeliu |
Pattern Recognit. | 2 |
| 1996 | Learning of context-sensitive languages described by augmented regular expressionsabstractRecently augmented regular expressions (AREs) have been proposed as a formalism to describe and recognize a non-trivial class of context-sensitive languages (CSLs). AREs augment the expressive power of regular expressions (REs) by including a set of constraints, that involve the number of instances in a string of the operands of the star operations of an RE. Although it has been demonstrated that not all the CSLs can be described by AREs, the class of representable objects includes planar shapes with symmetries, which is important for pattern recognition tasks. In this paper a general method to infer AREs from string examples is presented. The method consists of a regular grammatical inference step, aimed at obtaining a regular superset of the target language, followed by a constraint induction process, which reduces the extension of the inferred language attempting to discover the maximal number of context relations. Hence, this approach avoids the difficulty of learning context-sensitive grammars. René Alquézar, Alberto Sanfeliu |
ICPR | 2 |
| 1996 | Learning bidimensional context-dependent models using a context-sensitive languageabstractAutomatic generation of models from a set of positive and negative samples and a-priori knowledge (if available) is a crucial issue for pattern recognition applications. Grammatical inference can play an important role in this issue since it can be used to generate the set of model classes, where each class consists on the rules to generate the models. In this paper we present the process of learning context dependent bidimensional objects from outdoors images as context sensitive languages. We show how the process is conceived to overcome the problem of generalizing rules based on a set of samples which have small differences due to noisy pixels. The learned models can be used to identify objects in outdoors images irrespectively of their size and partial occlusions. Some results of the inference procedure are shown in the paper. Miguel Sainz, Alberto Sanfeliu |
ICPR | 2 |
| 1995 | An Algebraic Framework to Represent Finite State Machines in Single-Layer Recurrent Neural NetworksabstractIn this paper we present an algebraic framework to represent finite state machines (FSMs) in single-layer recurrent neural networks (SLRNNs), which unifies and generalizes some of the previous proposals. This framework is based on the formulation of both the state transition function and the output function of an FSM as a linear system of equations, and it permits an analytical explanation of the representational capabilities of first-order and higher-order SLRNNs. The framework can be used to insert symbolic knowledge in RNNs prior to learning from examples and to keep this knowledge while training the network. This approach is valid for a wide range of activation functions, whenever some stability conditions are met. The framework has already been used in practice in a hybrid method for grammatical inference reported elsewhere (Sanfeliu and Alquézar 1994). René Alquézar, Alberto Sanfeliu |
Neural Comput. | 2 |
| 1992 | Sensibility, relative error and error probability of projective invariants of planar surfaces of 3D objectsabstractPresents the study of the sensibility, relative error and error probability of projective invariants of planar objects. The study is applied to configurations of groups of five points which define a planar object. The results are general for any type of projective invariant. The paper shows that given the precision (or tolerance) and the maximum allowed error probability, the sensibility and relative error can be used to decide the correct configurations for the matching procedure between the model and the projective view of a physical object. The last result is general for any type of geometric invariant.> Alberto Sanfeliu, Antoni Llorens, Walter Emde |
ICPR (1) | 1 |
| 1988 | An architecture based on hybrid systems for analyzing 3D industrial scenesabstractA vision system architecture called PIRS (Perception Interpretation Robotic System) is described which analyzes scenes for robotic tasks, for purposes of identifying partially occluded and overlapped 3D objects, verifying modifications in a scene for task execution, and looking for an object in processes of task recovery. The objective of this system is to integrate different techniques for a research environment and connect them to diverse sensors and equipment. The system is based on a hybrid architecture which uses logic-based and structural techniques as well as several types of knowledge-base structures. The system consists of a supervisor which controls four main modules: classifier, verifier, predictor and vision processor.> Alberto Sanfeliu, Joan Font-Rosselló, Ignacio Orteu |
ICPR | 1 |
| 1987 | An Application of a Graph Distance Measure to the Classification of muscle Tissue PatternsabstractThis paper describes an application of a distance measure between attributed relational graphs to the classification of muscle tissue patterns. The goal is to classify these patterns according to an increasing degree of abnormality. These degrees of abnormality are from 1 to 5 with 1 for completely normal muscles and 5 for completely abnormal muscles. The muscle tissue patterns are described as graphs with the nodes of the graphs representing groups of fibers, and the branches, the fibers which separate the nodes. The distance measure defined in this paper computes the number of modifications required to transform an input muscle tissue pattern to the reference one. For the application of the proposed distance measure a reference or ideal graph representing a normal muscle tissue pattern is created. Then the distance is calculated between the input pattern and the reference, and the numerical value of this distance is used for classification. Zero distance implies that the input pattern is normal and a distance greater than zero indicates the degree of abnormality. Alberto Sanfeliu, King-Sun Fu, Judith M. S. Prewitt |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1986 | Introduction
Horst Bunke, Alberto Sanfeliu |
Pattern Recognit. | 2 |
| 1983 | A distance measure between attributed relational graphs for pattern recognitionabstractA method to determine a distance measure between two nonhierarchical attributed relational graphs is presented. In order to apply this distance measure, the graphs are characterised by descriptive graph grammars (DGG). The proposed distance measure is based on the computation of the minimum number of modifications required to transform an input graph into the reference one. Specifically, the distance measure is defined as the cost of recognition of nodes plus the number of transformations which include node insertion, node deletion, branch insertion, branch deletion, node label substitution and branch label substitution. The major difference between the proposed distance measure and the other ones is the consideration of the cost of recognition of nodes in the distance computation. In order to do this, the principal features of the nodes are described by one or several cost functions which are used to compute the similarity between the input nodes and the reference ones. Finally, an application of this distance measure to the recognition of lower case handwritten English characters is presented. Alberto Sanfeliu, King-Sun Fu |
IEEE Trans. Syst. Man Cybern. | 1 |