VLDB 2026 Research / reviewers in the wild / expert
José Santos-Victor
dblp:11/1265
· DBLP profile ↗
100ranked-venue papers
8as first author
10since 2021 · last 2025
0000-0002-9036-1728ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 88 · 7 first-author · 10 since 2021Systems, architecture and hardware · 48 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 7Human-computer interaction and ubiquitous computing · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GRASPLAT: Enabling dexterous grasping through novel view synthesisabstractAchieving dexterous robotic grasping with multi-fingered hands remains a significant challenge. While existing methods rely on complete 3D scans to predict grasp poses, these approaches face limitations due to the difficulty of acquiring high-quality 3D data in real-world scenarios. In this paper, we introduce GRASPLAT, a novel grasping framework that leverages consistent 3D information while being trained solely on RGB images. Our key insight is that by synthesizing physically plausible images of a hand grasping an object, we can regress the corresponding hand joints for a successful grasp. To achieve this, we utilize 3D Gaussian Splatting to generate high-fidelity novel views of real hand-object interactions, enabling end-to-end training with RGB data. Unlike prior methods, our approach incorporates a photometric loss that refines grasp predictions by minimizing discrepancies between rendered and real images. We conduct extensive experiments on both synthetic and real-world grasping datasets, demonstrating that GRASPLAT improves grasp success rates up to 36.9% over existing image-based methods. Project page: https://mbortolon97.github.io/grasplat/ Matteo Bortolon, Nuno Ferreira Duarte, Plinio Moreno, Fabio Poiesi, José Santos-Victor, Alessio Del Bue |
IROS | 5 |
| 2025 | Measuring Uncertainty in Shape Completion to Improve Grasp QualityabstractShape completion networks have been used recently in real-world robotic experiments to complete the missing/hidden information in environments where objects are only observed in one or few instances where self-occlusions are bound to occur. Nowadays, most approaches rely on deep neural networks that handle rich 3D point cloud data that lead to more precise and realistic object geometries. However, these models still suffer from inaccuracies due to its nondeterministic/stochastic inferences which could lead to poor performance in grasping scenarios where these errors compound to unsuccessful grasps. We present an approach to calculate the uncertainty of a 3D shape completion model during inference of single view point clouds of an object on a table top. In addition, we propose an update to grasp pose algorithms quality score by introducing the uncertainty of the completed point cloud present in the grasp candidates. To test our full pipeline we perform real world grasping with a 7dof robotic arm with a 2 finger gripper on a large set of household objects and compare against previous approaches that do not measure uncertainty. Our approach ranks the grasp quality better, leading to higher grasp success rate for the rank 5 grasp candidates compared to state of the art. Project Page: https://nunoduarte.github.io/pages.3dsgrasp++ Nuno Ferreira Duarte, Seyed Saber Mohammadi, Plinio Moreno, Alessio Del Bue, José Santos-Victor |
IROS | 5 |
| 2024 | The Gaze Dialogue Model: Nonverbal Communication in HHI and HRIabstractWhen humans interact with each other, eye gaze movements have to support motor control as well as communication. On the one hand, we need to fixate the task goal to retrieve visual information required for safe and precise action-execution. On the other hand, gaze movements fulfil the purpose of communication, both for reading the intention of our interaction partners, as well as to signal our action intentions to others. We study this Gaze Dialogue between two participants working on a collaborative task involving two types of actions: 1) individual action and 2) action-in-interaction. We recorded the eye-gaze data of both participants during the interaction sessions in order to build a computational model, the Gaze Dialogue, encoding the interplay of the eye movements during the dyadic interaction. The model also captures the correlation between the different gaze fixation points and the nature of the action. This knowledge is used to infer the type of action performed by an individual. We validated the model against the recorded eye-gaze behavior of one subject, taking the eye-gaze behavior of the other subject as the input. Finally, we used the model to design a humanoid robot controller that provides interpersonal gaze coordination in human-robot interaction scenarios. During the interaction, the robot is able to: 1) adequately infer the human action from gaze cues; 2) adjust its gaze fixation according to the human eye-gaze behavior; and 3) signal nonverbal cues that correlate with the robot's own action intentions. Mirko Rakovic, Nuno Ferreira Duarte, Jorge S. Marques, Aude Billard, José Santos-Victor |
IEEE Trans. Cybern. | 5 |
| 2023 | 3DSGrasp: 3D Shape-Completion for Robotic GraspabstractReal-world robotic grasping can be done robustly if a complete 3D Point Cloud Data (PCD) of an object is available. However, in practice, PCDs are often incomplete when objects are viewed from few and sparse viewpoints before the grasping action, leading to the generation of wrong or inaccurate grasp poses. We propose a novel grasping strategy, named 3DSGrasp, that predicts the missing geometry from the partial PCD to produce reliable grasp poses. Our proposed PCD completion network is a Transformer-based encoder-decoder network with an Offset-Attention layer. Our network is inherently invariant to the object pose and point's permutation, which generates PCDs that are geometrically consistent and completed properly. Experiments on a wide range of partial PCD show that 3DSGrasp outperforms the best state-of-the-art method on PCD completion tasks and largely improves the grasping success rate in real-world scenarios. The code and dataset are available at: https://github.com/NunoDuarte/3DSGrasp. Seyed Saber Mohammadi, Nuno Ferreira Duarte, Dimitrios Dimou, Yiming Wang 0002, Matteo Taiana, Pietro Morerio, Atabak Dehban, Plinio Moreno, Alexandre Bernardino, Alessio Del Bue, José Santos-Victor |
ICRA | 11 |
| 2023 | Expressing and Inferring Action Carefulness in Human-to-Robot HandoversabstractImplicit communication plays such a crucial role during social exchanges that it must be considered for a good experience in human-robot interaction. This work addresses implicit communication associated with the detection of physical properties, transport, and manipulation of objects. We propose an ecological approach to infer object characteristics from subtle modulations of the natural kinematics occurring during human object manipulation. Similarly, we take inspiration from human strategies to shape robot movements to be communica-tive of the object properties while pursuing the action goals. In a realistic HRI scenario, participants handed over cups - filled with water or empty - to a robotic manipulator that sorted them. We implemented an online classifier to differentiate careful/not careful human movements, associated with the cups' content. We compared our proposed “expressive” controller, which modulates the movements according to the cup filling, against a neutral motion controller. Results show that human kinematics is adjusted during the task, as a function of the cup content, even in reach-to-grasp motion. Moreover, the carefulness during the handover of full cups can be reliably inferred online, well before action completion. Finally, although questionnaires did not reveal explicit preferences from partici-pants, the expressive robot condition improved task efficiency. Linda Lastrico, Nuno Ferreira Duarte, Alessandro Carfì, Francesco Rea, Alessandra Sciutti, Fulvio Mastrogiovanni, José Santos-Victor |
IROS | 7 |
| 2022 | Robot Learning Physical Object Properties from Human Visual Cues: A novel approach to infer the fullness level in containersabstractFor collaborative tasks, involving handovers, humans are able to exploit visual, non-verbal cues, to infer physical object properties, like mass, to modulate their actions. In this paper, we investigate how the different levels of liquid inside a cup can be inferred from the observation of the movement of the person handling the cup. We model this mechanism from human experiments and incorporate it in an online human-to-robot handover. Finally, we provide a new dataset with human eye+head+hand motion data for human-to-human handovers and human pick-and-place of a cup with three levels of liquid: empty, half-full, and full of water. Our results show that it is possible to model (non-verbal) signals exchanged by humans during interaction and classify the level of water inside the cup being handed over. Nuno Ferreira Duarte, Mirko Rakovic, José Santos-Victor |
ICRA | 3 |
| 2021 | Learning Conditional Postural Synergies for Dexterous Hands: A Generative Approach Based on Variational Auto-Encoders and Conditioned on Object Size and CategoryabstractPostural synergies are used in robotics to facilitate the control of dexterous artificial hands. This is achieved by learning a latent space (synergy space) from grasp postures and directly controlling the hand in this space. In this work, we propose the use of a non-linear conditional model for learning the latent space, that can incorporate the object shape and size as additional variables. While on most of the previous works the evaluation criterion is the reconstruction error, we propose to use the smoothness of the latent space. Our model ranks better than other non-linear models in smoothness, which is a better criterion to evaluate in-hand manipulation tasks. We validate our arguments by executing regrasp trajectories in which our model outperforms all previous approaches. Dimitrios Dimou, José Santos-Victor, Plinio Moreno |
ICRA | 2 |
| 2021 | Learning Motor Resonance in Human-Human and Human-Robot Interaction with Coupled Dynamical SystemsabstractHuman interaction involves very sophisticated non-verbal communication skills like understanding the goals and actions of others and coordinating our own actions accordingly. Neuroscience refers to this mechanism as motor resonance, in the sense that the perception of another persons actions and sensory experiences activates the observer’s brain as if (s)he would be performing the same actions and having the same experiences.We analyze and model the non-verbal cues exchanged between two humans in handover actions. The contributions of this paper are the following: (i) computational models, using recorded motion data, describing the motor behaviour of each actor in action-in-interaction situations; (ii) a computational model that captures the behaviour of the "giver" and "receiver" during an object handover action, by coupling the wrist kinematic motion of the actors; and (iii) the transfer of these models to the iCub robot for both action execution and recognition.Our results show that: (i) the robot is able to interpret the human wrist motion and infer whether or not the observed action is an "handover"; and (ii) use the motor resonance model to coordinate its actions with the human partner, during handover actions. Nuno Ferreira Duarte, Mirko Rakovic, José Santos-Victor |
ICRA | 3 |
| 2021 | SENSORIMOTOR GRAPH: Action-Conditioned Graph Neural Network for Learning Robotic Soft Hand DynamicsabstractSoft robotics is a thriving branch of robotics which takes inspiration from nature and uses affordable flexible materials to design adaptable non-rigid robots. However, their flexible behavior makes these robots hard to model, which is essential for a precise actuation and for optimal control. For system modelling, learning-based approaches have demonstrated good results, yet they fail to consider the physical structure underlying the system as an inductive prior. In this work, we take inspiration from sensorimotor learning, and apply a Graph Neural Network to the problem of modelling a non-rigid kinematic chain (i.e. a robotic soft hand) taking advantage of two key properties: 1) the system is compositional, that is, it is composed of simple interacting parts connected by edges, 2) it is order invariant, i.e. only the structure of the system is relevant for predicting future trajectories. We denote our model as the "Sensorimotor Graph" since it learns the system connectivity from observation and uses it for dynamics prediction. We validate our model in different scenarios and show that it outperforms the non-structured baselines in dynamics prediction while being more robust to configurational variations, tracking errors or node failures. João Damião Almeida, Paul Schydlo, Atabak Dehban, José Santos-Victor |
IROS | 4 |
| 2021 | Action anticipation for collaborative environments: The impact of contextual information and uncertainty-based prediction
Clebeson Canuto dos Santos, Plinio Moreno, Jorge Leonid Aching Samatelo, Raquel Frizera Vassallo, José Santos-Victor |
Neurocomputing | 5 |
| 2020 | Action-conditioned Benchmarking of Robotic Video Prediction Models: a Comparative StudyabstractA defining characteristic of intelligent systems is the ability to make action decisions based on the anticipated outcomes. Video prediction systems have been demonstrated as a solution for predicting how the future will unfold visually, and thus, many models have been proposed that are capable of predicting future frames based on a history of observed frames (and sometimes robot actions). However, a comprehensive method for determining the fitness of different video prediction models at guiding the selection of actions is yet to be developed.Current metrics assess video prediction models based on human perception of frame quality. In contrast, we argue that if these systems are to be used to guide action, necessarily, the actions the robot performs should be encoded in the predicted frames. In this paper, we are proposing a new metric to compare different video prediction models based on this argument. More specifically, we propose an action inference system and quantitatively rank different models based on how well we can infer the robot actions from the predicted frames. Our extensive experiments show that models with high perceptual scores can perform poorly in the proposed action inference tests and thus, may not be suitable options to be used in robot planning systems. Manuel Serra Nunes, Atabak Dehban, Plinio Moreno, José Santos-Victor |
ICRA | 4 |
| 2019 | The Impact of Domain Randomization on Object Detection: A Case Study on Parametric Shapes and Synthetic Textures*abstractRecent advances in deep learning-based object detection techniques have revolutionized their applicability in several fields. However, since these methods rely on unwieldy and large amounts of data, a common practice is to download models pre-trained on standard datasets and fine-tune them for specific application domains with a small set of domain-relevant images. In this work, we show that using synthetic datasets that are not necessarily photo-realistic can be a better alternative to simply fine-tune pre-trained networks. Specifically, our results show an impressive 25%improvement in the mAP metric over a fine-tuning baseline when only about 200 labelled images are available to train. Finally, an ablation study of our results is presented to delineate the individual contribution of different components in the randomization pipeline. Atabak Dehban, João Borrego, Rui Pimentel de Figueiredo, Plinio Moreno, Alexandre Bernardino, José Santos-Victor |
IROS | 6 |
| 2019 | Coupling of Arm Movements during Human-Robot Interaction: the handover caseabstractCollaboration involves understanding the action of others, as well as acting in a way that can be understood by others. One of those tasks is the handover. In this paper, we study the behaviour of humans during the handover and design the mechanisms allowing a robot to learn from that behaviour.We analyse and model the arm movements of humans while handing over objects to one another. The contributions of this paper are the following: (i) a computational model that captures the behaviour of the “giver” and “receiver” of the object, by coupling the arm motion; (ii) discuss this approach amidst a previous coupling strategy; and (iii) embedded the model in the iCub robot for human to robot handovers.Our results show that: (i) the robot can coordinate with the human to timely and safely receive the object; (ii) the robot behaves in a “human-like” manner while receiving the object; and (iii) our approach has significant advantages to the previous approach. Nuno Ferreira Duarte, Mirko Rakovic, José Santos-Victor |
RO-MAN | 3 |
| 2018 | Anticipation in Human-Robot Cooperation: A Recurrent Neural Network Approach for Multiple Action Sequences PredictionabstractClose human-robot cooperation is a key enabler for new developments in advanced manufacturing and assistive applications. Close cooperation require robots that can predict human actions and intent, understanding human non-verbal cues. Recent approaches based on neural networks have led to encouraging results in the human action prediction problem both in continuous and discrete spaces. Our approach extends the research in this direction. Our contributions are three-fold. First, we validate the use of gaze and body pose cues as a means of predicting human action through a feature selection method. Next, we address two shortcomings of existing literature: predicting multiple and variable-length action sequences. This is achieved by applying an encoder-decoder recurrent neural network topology in the discrete action prediction problem. In addition, we theoretically demonstrate the importance of predicting multiple action sequences as a means of estimating the stochastic reward in a human robot cooperation scenario. Finally, we show the ability to effectively train the prediction model on an action prediction dataset, involving human motion data, and explore the influence of the model's parameters on its performance. Paul Schydlo, Mirko Rakovic, Lorenzo Jamone, José Santos-Victor |
ICRA | 4 |
| 2017 | Low-cost 3-axis soft tactile sensors for the human-friendly robot VizzyabstractIn this paper we present a low-cost and easy to fabricate 3-axis tactile sensor based on magnetic technology. The sensor consists in a small magnet immersed in a silicone body with an Hall-effect sensor placed below to detect changes in the magnetic field caused by displacements of the magnet, generated by an external force applied to the silicone body. The use of a 3-axis Hall-effect sensor allows to detect the three components of the force vector, and the proposed design assures high sensitivity, low hysteresis and good repeatability of the measurement: notably, the minimum sensed force is about 0.007N. All components are cheap and easy to retrieve and to assemble; the fabrication process is described in detail and it can be easily replicated by other researchers. Sensors with different geometries have been fabricated, calibrated and successfully integrated in the hand of the human-friendly robot Vizzy. In addition to the sensor characterization and validation, real world experiments of object manipulation are reported, showing proper detection of both normal and shear forces. Tiago Paulino, Pedro Ribeiro 0006, Susana Cardoso, Alexander Schmitz, José Santos-Victor, Alexandre Bernardino, Lorenzo Jamone |
ICRA | 6 |
| 2016 | Person Re-identification in Frontal Gait Sequences via Histogram of Optic Flow Energy Image
Athira Nambiar, Jacinto C. Nascimento, Alexandre Bernardino, José Santos-Victor |
ACIVS | 4 |
| 2016 | Denoising auto-encoders for learning of objects and tools affordances in continuous spaceabstractThe concept of affordances facilitates the encoding of relations between actions and effects in an environment centered around the agent. Such an interpretation has important impacts on several cognitive capabilities and manifestations of intelligence, such as prediction and planning. In this paper, a new framework based on denoising Auto-encoders (dA) is proposed which allows an agent to explore its environment and actively learn the affordances of objects and tools by observing the consequences of acting on them. The dA serves as a unified framework to fuse multi-modal data and retrieve an entire missing modality or a feature within a modality given information about other modalities. This work has two major contributions. First, since training the dA is done in continuous space, there will be no need to discretize the dataset and higher accuracies in inference can be achieved with respect to approaches in which data discretization is required (e.g. Bayesian networks). Second, by fixing the structure of the dA, knowledge can be added incrementally making the architecture particularly useful in online learning scenarios. Evaluation scores of real and simulated robotic experiments show improvements over previous approaches while the new model can be applied in a wider range of domains. Atabak Dehban, Lorenzo Jamone, Adam R. Kampff, José Santos-Victor |
ICRA | 4 |
| 2016 | On Stereo Confidence Measures for Global Methods: Evaluation, New Model and Integration into Occupancy GridsabstractStereo confidence measures are important functions for global reconstruction methods and some applications of stereo. In this article we evaluate and compare several models of confidence which are defined at the whole disparity range. We propose a new stereo confidence measure to which we call the Histogram Sensor Model (HSM), and show how it is one of the best performing functions overall. We also introduce, for parametric models, a systematic method for estimating their parameters which is shown to lead to better performance when compared to parameters as computed in previous literature. All models were evaluated when applied to two different cost functions at different window sizes and model parameters. Contrary to previous stereo confidence measure benchmark literature, we evaluate the models with criteria important not only to winner-take-all stereo, but also to global applications. To this end, we evaluate the models on a real-world application using a recent formulation of 3D reconstruction through occupancy grids which integrates stereo confidence at all disparities. We obtain and discuss our results on both indoors' and outdoors' publicly available datasets. Martim Brandão, Ricardo Ferreira 0002, Kenji Hashimoto, Atsuo Takanishi, José Santos-Victor |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2016 | Footstep Planning for Slippery and Slanted Terrain Using Human-Inspired ModelsabstractEnergy efficiency and robustness of locomotion to different terrain conditions are important problems for humanoid robots deployed in the real world. In this paper, we propose a footstep-planning algorithm for humanoids that is applicable to flat, slanted, and slippery terrain, which uses simple principles and representations gathered from human gait literature. The planner optimizes a center-of-mass (COM) mechanical work model subject to motion feasibility and ground friction constraints using a hybrid A* search and optimization approach. Footstep placements and orientations are discrete states searched with an A* algorithm, while other relevant parameters are computed through continuous optimization on state transitions. These parameters are also inspired by human gait literature and include footstep timing (double-support and swing time) and parameterized COM motion using knee flexion angle keypoints. The planner relies on work, the required coefficient of friction (RCOF), and feasibility models that we estimate in a physics simulation. We show through simulation experiments that the proposed planner leads to both low electrical energy consumption and human-like motion on a variety of scenarios. Using the planner, the robot automatically opts between avoiding or (slowly) traversing slippery patches depending on their size and friction, and it chooses energy-optimal stairs and climbing angles in slopes. The obtained motion is also consistent with observations found in human gait literature, such as human-like changes in RCOF, step length and double-support time on slippery terrain, and human-like curved walking on steep slopes. Finally, we compare COM work minimization with other choices of the objective function. Martim Brandão, Kenji Hashimoto, José Santos-Victor, Atsuo Takanishi |
IEEE Trans. Robotics | 3 |
| 2015 | A Vision-Based System for Movement Analysis in Medical Applications: The Example of Parkinson Disease
Sofija Spasojevic, José Santos-Victor, Tihomir Ilic, Sladan Milanovic, Veljko Potkonjak, Aleksandar Rodic 0001 |
ICVS | 2 |
| 2015 | People and Mobile Robot Classification Through Spatio-Temporal Analysis of Optical FlowabstractThe goal of this work is to distinguish between humans and robots in a mixed human-robot environment. We analyze the spatio-temporal patterns of optical flow-based features along several frames. We consider the Histogram of Optical Flow (HOF) and the Motion Boundary Histogram (MBH) features, which have shown good results on people detection. The spatio-temporal patterns are composed of groups of feature components that have similar values on previous frames. The groups of features are fed into the FuzzyBoost algorithm, which at each round selects the spatio-temporal pattern (i.e. feature set) having the lowest classification error. The search for patterns is guided by grouping feature dimensions, considering three algorithms: (a) similarity of weights from dimensionality reduction matrices, (b) Boost Feature Subset Selection (BFSS) and (c) Sequential Floating Feature Selection (SFSS), which avoid the brute force approach. The similarity weights are computed by the Multiple Metric Learning for large Margin Nearest Neighbor (MMLMNN), a linear dimensionality algorithm that provides a type of Mahalanobis metric Weinberger and Saul, J. MaCh. Learn. Res.10 (2009) 207–244. The experiments show that FuzzyBoost brings good generalization properties, better than the GentleBoost, the Support Vector Machines (SVM) with linear kernels and SVM with Radial Basis Function (RBF) kernels. The classifier was implemented and tested in a real-time, multi-camera dynamic setting. Plinio Moreno, Dario Figueira, Alexandre Bernardino, José Santos-Victor |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2014 | On the formulation, performance and design choices of Cost-Curve Occupancy Grids for stereo-vision based 3D reconstructionabstractWe present a grid-based 3D reconstruction method which integrates all costs given by stereo vision into what we call a Cost-Curve Occupancy Grid (CCOG). Occupancy probabilities of grid cells are estimated in a Bayesian formulation, from the likelihood of stereo cost measurements taken at all distance hypotheses. This is accomplished with only a small set of probabilistic assumptions which we discuss in the paper. We quantitatively characterize the method's performance under different conditions of both image noise and number of used stereo pairs, compared also to traditional algorithms. We complement the study by giving insights on design choices of CCOGs such as likelihood model, window size of the cost function and use of a hole filling method. Experiments were made on a real-world outdoors dataset with ground-truth data. Martim Brandão, Ricardo Ferreira 0002, Kenji Hashimoto, José Santos-Victor, Atsuo Takanishi |
IROS | 4 |
| 2014 | 3D to 2D bijection for spherical objects under equidistant fisheye projection
Aamir Ahmad, João M. F. Xavier, José Santos-Victor, Pedro U. Lima |
Comput. Vis. Image Underst. | 3 |
| 2013 | Multi-object detection and pose estimation in 3D point clouds: A fast grid-based Bayesian FilterabstractWe address the problem of object detection and pose estimation using 3D dense data in a multiple object library scenario. State-of-the-art object detection and pose estimation methods are able cope with background clutter and occlusion with acceptable noise levels in the single object scenario. However, with multiple object libraries, even moderate amount of noise lead to frequent object identity switches and serious pose estimation errors. To attenuate these effects, we propose a joint object-id and pose filtering approach using grid-based Recursive Bayesian Filters (RBF). The grid method considers as state variables the object label and its pose, and models the dynamics of the filter with two “inertia” parameters: one for the object label and the other for the object pose. Sensor noise characteristics are taken into account with an observation noise parameter. To allow real-time functionality we propose a selective update approach that dynamically reduces the set of hypotheses evaluated at run time. We present results in realistic scenarios and compare our approach with state-of-the-art approaches in a three object problem, with significant performance improvements. Rui Pimentel de Figueiredo, Plinio Moreno, Alexandre Bernardino, José Santos-Victor |
ICRA | 4 |
| 2013 | Online learning of humanoid robot kinematics under switching tools contextsabstractIn this paper a novel approach to kinematics learning and task space control, under switching contexts, is presented. Such non-stationary contexts may appear in many robotic tasks: in particular, the changing of the context due to the use of tools with different lengths and shapes is herein studied. We model the robot forward kinematics as a multi-valued function, in which different outputs for the same input query are related to actual different hidden contexts. To do that, we employ IMLE, a recent online learning algorithm that fits an infinite mixture of linear experts to the online stream of training data. This algorithm can directly provide multi-valued regression in a online fashion, while having, for classic single-valued regression, a performance comparable to state-of-the-art online learning algorithms. The context varying forward kinematics is learned online through exploration, not relying on any kind of prior knowledge. Using the proposed approach, the robot can dynamically learn how to use different tools, without forgetting the kinematic mappings concerning previously manipulated tools. No information is given about such tool changes to the learning algorithm, nor any assumption is made about the tool kinematics. To our knowledge this is the most general and efficient approach to learning and control under discrete varying contexts. Some experimental results obtained on a high-dimensional simulated humanoid robot provide a strong support to our approach. Lorenzo Jamone, Bruno D. Damas, José Santos-Victor, Atsuo Takanishi |
ICRA | 3 |
| 2013 | Integrating the whole cost-curve of stereo into occupancy gridsabstractExtensive literature has been written on occupancy grid mapping for different sensors. When stereo vision is applied to the occupancy grid framework it is common, however, to use sensor models that were originally conceived for other sensors such as sonar. Although sonar provides a distance to the nearest obstacle for several directions, stereo has confidence measures available for each distance along each direction. The common approach is to take the highest-confidence distance as the correct one, but such an approach disregards mismatch errors inherent to stereo. In this work, stereo confidence measures of the whole sensed space are explicitly integrated into 3D grids using a new occupancy grid formulation. Confidence measures themselves are used to model uncertainty and their parameters are computed automatically in a maximum likelihood approach. The proposed methodology was evaluated in both simulation and a real-world outdoor dataset which is publicly available. Mapping performance of our approach was compared with a traditional approach and shown to achieve less errors in the reconstruction. Martim Brandão, Ricardo Ferreira 0002, Kenji Hashimoto, José Santos-Victor, Atsuo Takanishi |
IROS | 4 |
| 2013 | Open and closed-loop task space trajectory control of redundant robots using learned modelsabstractThis paper presents a comparison of open-loop and closed-loop control strategies for tracking a task space trajectory, using redundant robots. We do not assume any knowledge of the analytical forward and inverse kinematics, relying instead on learning these models online, while executing a desired task. Specifically, we employ a recent learning algorithm that allows to learn a probabilistic model from which both the forward and inverse solutions can be obtained, as well as the Jacobian of the kinematics map. Such learned model can then be used to implement both types of control. Moreover, the multi-valued solutions provided by the learned model can be applied to redundant systems in which an infinite number of inverse solutions may exist. We present experiments with a simulated version of the iCub, a highly redundant humanoid robot, in which this learned model is employed to execute both open-loop and closed-loop trajectory control. We show the advantages and drawbacks of both control strategies, and we propose a way to combine them to deal with sensor noise and failures, showing the benefits of using a learning algorithm that can simultaneously provide forward and inverse predictions. Bruno D. Damas, Lorenzo Jamone, José Santos-Victor |
IROS | 3 |
| 2013 | Learning robot gait stability using neural networks as sensory feedback function for Central Pattern GeneratorsabstractIn this paper we present a framework to learn a model-free feedback controller for locomotion and balance control of a compliant quadruped robot walking on rough terrain. Having designed an open-loop gait encoded in a Central Pattern Generator (CPG), we use a neural network to represent sensory feedback inside the CPG dynamics. This neural network accepts sensory inputs from a gyroscope or a camera, and its weights are learned using Particle Swarm Optimization (unsupervised learning). We show with a simulated compliant quadruped robot that our controller can perform significantly better than the open-loop one on slopes and randomized height maps. Sébastien Gay, José Santos-Victor, Auke Jan Ijspeert |
IROS | 2 |
| 2013 | Online Learning of Single- and Multivalued Functions with an Infinite Mixture of Linear ExpertsabstractWe present a supervised learning algorithm for estimation of generic input-output relations in a real-time, online fashion. The proposed method is based on a generalized expectation-maximization approach to fit an infinite mixture of linear experts (IMLE) to an online stream of data samples. This probabilistic model, while not fully Bayesian, can efficiently choose the number of experts that are allocated to the mixture, this way effectively controlling the complexity of the resulting model. The result is an incremental, online, and localized learning algorithm that performs nonlinear, multivariate regression on multivariate outputs by approximating the target function by a linear relation within each expert input domain and that can allocate new experts as needed. A distinctive feature of the proposed method is the ability to learn multivalued functions: one-to-many mappings that naturally arise in some robotic and computer vision learning domains, using an approach based on a Bayesian generative model for the predictions provided by each of the mixture experts. As a consequence, it is able to directly provide forward and inverse relations from the same learned mixture model. We conduct an extensive set of experiments to evaluate the proposed algorithm performance, and the results show that it can outperform state-of-the-art online function approximation algorithms in single-valued regression, while demonstrating good estimation capabilities in a multivalued function approximation context. Bruno D. Damas, José Santos-Victor |
Neural Comput. | 2 |
| 2012 | Predictive gaze stabilization during periodic locomotion based on Adaptive Frequency OscillatorsabstractIn this paper we present an approach to the problem of stabilizating the gaze of legged robots using Adaptive Frequency Oscillators to learn the frequency, phase and amplitude of the optical flow and generate compensatory commands during robot locomotion. Assuming periodic and nearly sine shaped motion of the head of the robot, the system successfully stabilizes the gaze of the robot, whether the robot itself is moving, or an external object is moving relative to the robot. We present experiments in simulation and, for object tracking, with a real robotics setup, the Hoap 3, showing that the system can be successfully applied to gaze stabilization during locomotion, even when the feedback loop is very slow and noisy. Sébastien Gay, Auke Jan Ijspeert, José Santos-Victor |
ICRA | 3 |
| 2012 | Learning relational affordance models for robots in multi-object manipulation tasksabstractAffordances define the action possibilities on an object in the environment and in robotics they play a role in basic cognitive capabilities. Previous works have focused on affordance models for just one object even though in many scenarios they are defined by configurations of multiple objects that interact with each other. We employ recent advances in statistical relational learning to learn affordance models in such cases. Our models generalize over objects and can deal effectively with uncertainty. Two-object interaction models are learned from robotic interaction with the objects in the world and employed in situations with arbitrary numbers of objects. We illustrate these ideas with experimental results of an action recognition task where a robot manipulates objects on a shelf. Bogdan Moldovan, Plinio Moreno, Martijn van Otterlo, José Santos-Victor, Luc De Raedt |
ICRA | 4 |
| 2012 | An online algorithm for simultaneously learning forward and inverse kinematicsabstractThis paper proposes a supervised algorithm for online learning of input-output relations that is particularly suitable to simultaneously learn the forward and inverse kinematics of general manipulators - the multi-valued nature of the inverse kinematics of serial chains and forward kinematics of parallel manipulators makes it infeasible to apply state-of-the-art learning techniques to these problems, as they typically assume a single-valued function to be learned. The proposed algorithm is based on a generalized expectation maximization approach to fit an infinite mixture of linear experts to an online stream of data samples, together with an outlier probabilistic model that dynamically grows the number of linear experts allocated to the mixture, this way controlling the complexity of the resulting model. The result is an incremental, online and localized learning algorithm that performs nonlinear, multivariate regression on multivariate outputs by approximating the target function by a linear relation within each expert input domain, which can directly provide forward and inverse multi-valued estimates. The experiments presented in this paper show that it can achieve, for single-valued functions, a performance directly comparable to state-of-the-art online function approximation algorithms, while additionally providing inverse predictions and the capability to learn multi-valued functions in a natural manner. To our knowledge this is a distinctive property of the algorithm presented in this paper. Bruno D. Damas, José Santos-Victor |
IROS | 2 |
| 2012 | Online calibration of a humanoid robot head from relative encoders, IMU readings and visual dataabstractHumanoid robots are complex sensorimotor systems where the existence of internal models are of utmost importance both for control purposes and for predicting the changes in the world arising from the system's own actions. This so-called expected perception relies on the existence of accurate internal models of the robot's sensorimotor chains. Nuno Moutinho, Martim Brandão, Ricardo Ferreira 0002, José António Gaspar, Alexandre Bernardino, Atsuo Takanishi, José Santos-Victor |
IROS | 7 |
| 2012 | Fast estimation of Gaussian mixture models for image segmentation
Nicola Greggio, Alexandre Bernardino, Cecilia Laschi, Paolo Dario, José Santos-Victor |
Mach. Vis. Appl. | 5 |
| 2012 | Language Bootstrapping: Learning Word Meanings From Perception-Action AssociationabstractWe address the problem of bootstrapping language acquisition for an artificial system similarly to what is observed in experiments with human infants. Our method works by associating meanings to words in manipulation tasks, as a robot interacts with objects and listens to verbal descriptions of the interactions. The model is based on an affordance network, i.e., a mapping between robot actions, robot perceptions, and the perceived effects of these actions upon objects. We extend the affordance model to incorporate spoken words, which allows us to ground the verbal symbols to the execution of actions and the perception of the environment. The model takes verbal descriptions of a task as the input and uses temporal co-occurrence to create links between speech utterances and the involved objects, actions, and effects. We show that the robot is able form useful word-to-meaning associations, even without considering grammatical structure in the learning process and in the presence of recognition errors. These word-to-meaning associations are embedded in the robot's own understanding of its actions. Thus, they can be directly used to instruct the robot to perform tasks and also allow to incorporate context in the speech recognition task. We believe that the encouraging results with our approach may afford robots with a capacity to acquire language descriptors in their operation's environment as well as to shed some light as to how this challenging process develops with human infants. Giampiero Salvi, Luis Montesano, Alexandre Bernardino, José Santos-Victor |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2011 | Real-time Ellipse Fitting, 3D Spherical Object Localization, and Tracking for the iCub Simulator
Nicola Greggio, Alexandre Bernardino, Cecilia Laschi, Paolo Dario, José Santos-Victor |
ICINCO (2) | 5 |
| 2011 | Monocular Vs Binocular 3D Real-time Ball Tracking from 2D Ellipses
Nicola Greggio, José António Gaspar, Alexandre Bernardino, José Santos-Victor |
ICINCO (2) | 4 |
| 2011 | An expected perception architecture using visual 3D reconstruction for a humanoid robotabstractThe maintenance of a stable and coherent representation of the surrounding environment is an essential capability in cognitive robotic systems. Most systems employ some form of 3D perception to create internal representations of space (maps) to support tasks such as navigation, manipulation and interaction. The creation and update of such representations may represent a significant effort in the overall computation performed by the robot. In this paper we propose an architecture based on the concept of Expected Perception that allows lightweight map updates whenever the course of action happens according to the robot's expectations. It is only when the robot's predictions and the real world outcomes differ, that corrections must be done at its full extent. We performed experiments and show results in a real robotic platform with stereo (3D) perception where map corrections are proposed by simple image level (2D) comparisons. Nuno Moutinho, Nino Cauli, Egidio Falotico, Ricardo Ferreira 0002, José António Gaspar, Alexandre Bernardino, José Santos-Victor, Paolo Dario, Cecilia Laschi |
IROS | 7 |
| 2010 | A Practical Method for Self-adapting Gaussian Expectation Maximization
Nicola Greggio, Alexandre Bernardino, José Santos-Victor |
ICINCO (1) | 3 |
| 2010 | Bioinspired Robotics and Vision with Humanoid Robots
José Santos-Victor |
ICINCO (1) | 1 |
| 2010 | Unsupervised and Online Update of Boosted Temporal Models: The UAL2BoostabstractThe application of learning-based vision techniques to real scenarios usually requires a tunning procedure, which involves the acquisition and labeling of new data and in situ experiments in order to adapt the learning algorithm to each scenario. We address an automatic update procedure of the L2boost algorithm that is able to adapt the initial models learned off-line. Our method is named UAL2Boost and present three new contributions: (i) an on-line and continuous procedure that updates recursively the current classifier, reducing the storage constraints, (ii) a probabilistic unsupervised update that eliminates the necessity of labeled data in order to adapt the classifier and (iii) a multi-class adaptation method. We show the applicability of the on-line unsupervised adaptation to human action recognition and demonstrate that the system is able to automatically update the parameters of the L2boost with linear temporal models, thus improving the output of the models learned off-line on new video sequences, in a recursive and continuous way. The automatic adaptation of UAL2Boost follows the idea of adapting the classifier incrementally: from simple to complex. Pedro Ribeiro 0006, Plinio Moreno, José Santos-Victor |
ICMLA | 3 |
| 2010 | Unsupervised Greedy Learning of Finite Mixture ModelsabstractThis work deals with a new technique for the estimation of the parameters and number of components in a finite mixture model. The learning procedure is performed by means of a expectation maximization (EM) methodology. The key feature of our approach is related to a top-down hierarchical search for the number of components, together with the integration of the model selection criterion within a modified EM procedure, used for the learning the mixture parameters. We start with a single component covering the whole data set. Then new components are added and optimized to best cover the data. The process is recursive and builds a binary tree like structure that effectively explores the search space. We show that our approach is faster that state-of-the- art alternatives, is insensitive to initialization, and has better data fits in average. We elucidate this through a series of experiments, both with synthetic and real data. Nicola Greggio, Alexandre Bernardino, Cecilia Laschi, Paolo Dario, José Santos-Victor |
ICTAI (2) | 5 |
| 2010 | An Algorithm for the Least Square-Fitting of EllipsesabstractIn this paper we propose a new algorithm for the least square fitting of ellipses from scattered data. Originally based on the one proposed by Fitzgibbon et Al in 1999, our procedure is able to overcome the numerical instability of that algorithm. We test our approach versus the latter and another approach with different ellipses. Then, we present and discuss our results. Nicola Greggio, Alexandre Bernardino, Cecilia Laschi, Paolo Dario, José Santos-Victor |
ICTAI (2) | 5 |
| 2010 | Learning words and speech units through natural interactionsabstractThis work provides an ecological approach to learning words and speech units through natural interactions, without the need for preprogrammed linguistic knowledge in form of phonemes. Interactions such as imitation games and multimodal word learning create an initial set of words and speech units. These sets are then used to train statistical models in an unsupervised way. Index Terms: multimodal learning, ecological approach, motor learning, interactions Jonas Hörnstein, José Santos-Victor |
INTERSPEECH | 2 |
| 2010 | Integration of vision and central pattern generator based locomotion for path planning of a non-holonomic crawling humanoid robotabstractIn this paper we present our work on integrating a locomotion controller based on central pattern generator (CPG) and a motion planning algorithm using artificial potential fields for a non-holonomic crawling humanoid robot, the iCub. We also integrated a vision tracker and an inverse kinematics solver to perform reaching tasks. We study the influence of the various parameters of the potential field equations on the performance of the system and prove the efficiency of our framework by testing it on a physics-based robotics simulator and partially on the real iCub. Sébastien Gay, Sarah Dégallier-Rochat, Ugo Pattacini, Auke Jan Ijspeert, José Santos-Victor |
IROS | 5 |
| 2010 | Gaussian mixture models for affordance learning using Bayesian NetworksabstractAffordances are fundamental descriptors of relationships between actions, objects and effects. They provide the means whereby a robot can predict effects, recognize actions, select objects and plan its behavior according to desired goals. This paper approaches the problem of an embodied agent exploring the world and learning these affordances autonomously from its sensory experiences. Models exist for learning the structure and the parameters of a Bayesian Network encoding this knowledge. Although Bayesian Networks are capable of dealing with uncertainty and redundancy, previous work considered complete observability of the discrete sensory data, which may lead to hard errors in the presence of noise. In this paper we consider a probabilistic representation of the sensors by Gaussian Mixture Models (GMMs) and explicitly taking into account the probability distribution contained in each discrete affordance concept, which can lead to a more correct learning. Pedro Osório, Alexandre Bernardino, Ruben Martinez-Cantin, José Santos-Victor |
IROS | 4 |
| 2010 | Sensor-based self-calibration of the iCub's headabstractIn this paper we propose techniques for the calibration of the iCub's stereo head using vision and inertial measurements. Given that wear and tear can change the geometrical relationship between the different elements in the kinematic chain, new calibrations must be performed periodically. We propose methods that allow automatic calibration without the need for using external sensors or specially designed calibration objects. The methods can be applied at any time during the operation of the system, thus being an alternative for systems whose calibrations are imprecise or that require frequent recalibration. Results are shown both in simulations and on the iCub's stereo head. José Fragoso Santos, Alexandre Bernardino, José Santos-Victor |
IROS | 3 |
| 2010 | Self-adaptive Gaussian mixture models for real-time video segmentation and background subtractionabstractThe usage of Gaussian mixture models for video segmentation has been widely adopted. However, the main difficulty arises in choosing the best model complexity. High complex models can describe the scene accurately, but they come with a high computational requirements, too. Low complex models promote segmentation speed, with the drawback of a less exhaustive description. In this paper we propose an algorithm that first learns a description mixture for the first video frames, and then it uses these results as a starting point for the analysis of the further frames. Then, we apply it to a video sequence and show its effectiveness for real-time tracking multiple moving objects. Moreover, we integrated this procedure into a foreground/background subtraction statistical framework. We compare our procedure against the state-of-the-art alternatives, and we show both its initialization efficacy and its improved segmentation performance. Nicola Greggio, Alexandre Bernardino, Cecilia Laschi, Paolo Dario, José Santos-Victor |
ISDA | 5 |
| 2010 | The iCub humanoid robot: An open-systems platform for research in cognitive development
Giorgio Metta, Lorenzo Natale, Francesco Nori, Giulio Sandini, David Vernon, Luciano Fadiga, Claes von Hofsten, Kerstin Rosander, Manuel Lopes 0001, José Santos-Victor, Alexandre Bernardino, Luis Montesano |
Neural Networks | 10 |
| 2009 | Affordance based word-to-meaning associationabstractThis paper presents a method to associate meanings to words in manipulation tasks. We base our model on an affordance network, i.e., a mapping between robot actions, robot perceptions and the perceived effects of these actions upon objects. We extend the affordance model to incorporate words. Using verbal descriptions of a task, the model uses temporal co-occurrence to create links between speech utterances and the involved objects, actions and effects. We show that the robot is able form useful word-to-meaning associations, even without considering grammatical structure in the learning process and in the presence of recognition errors. These word-to-meaning associations are embedded in the robot's own understanding of its actions. Thus they can be directly used to instruct the robot to perform tasks and also allow to incorporate context in the speech recognition task. Verica Krunic, Giampiero Salvi, Alexandre Bernardino, Luis Montesano, José Santos-Victor |
ICRA | 5 |
| 2009 | ISROBOTNET: A testbed for sensor and robot network systemsabstractThis paper introduces a testbed for sensor and robot network systems, currently composed of 10 cameras and 5 mobile wheeled robots equipped with several sensors for self-localization, obstacle avoidance and vision cameras, and wireless communications. The testbed includes a service-oriented middleware to enable fast prototyping and implementation of algorithms previously tested in simulation, as well as to simplify integration of subsystems developed by different partners. We survey an integrated approach to human-robot interaction that has been developed supported by the testbed under an European research project. The application integrates innovative methods and algorithms for people tracking and waving detection, cooperative perception among static and mobile cameras to improve people tracking accuracy, as well as decision-theoretical approaches to sensor selection and task allocation within the sensor network. Marco Barbosa, Alexandre Bernardino, Dario Figueira, José António Gaspar, Nelson Gonçalves, Pedro U. Lima, Plinio Moreno, Abdolkarim Pahliani, José Santos-Victor, Matthijs T. J. Spaan, João Sequeira 0001 |
IROS | 9 |
| 2009 | Avoiding moving obstacles: the forbidden velocity mapabstractRobotic obstacle avoidance in cluttered and dense environments is an important issue in robotic navigation. Over the past few years a number of techniques has been proposed to deal with safe navigation among obstacles in unknown scenarios. Unfortunately many of these methods do not consider obstacle velocities, which can rise some serious questions concerning their safety [1]. This paper will deal with a novel approach to moving obstacle avoidance in holonomic robots. It proposes the Forbidden VelocityMap, a generalization of the Dynamic Window concept [2] that considers obstacle and robot shape, velocity and dynamics, resulting in a safe, reactive real-time navigation algorithm that is able to deal with navigation in unpredictable and cluttered scenarios. Bruno D. Damas, José Santos-Victor |
IROS | 2 |
| 2009 | Multimodal word learning from Infant Directed SpeechabstractWhen adults talk to infants they do that in a different way compared to how they communicate with other adults. This kind of infant directed speech (IDS) typically highlights target words using focal stress and utterance final position. Also, speech directed to infants often refers to objects, people and events in the world surrounding the infant. Because of this, the sound sequences the infant hears are very likely to co-occur with actual objects or events in the infant's visual field. In this work we present a model that is able to learn word-like structures from multimodal information sources without any pre-programmed linguistic knowlege, by taking advantage of the characteristics of IDS. The model is implemented on a humanoid robot platform and is able to extract word-like patterns and associating these to objects in the visual surrounding. Jonas Hörnstein, Lisa Gustavsson, Francisco Lacerda, José Santos-Victor |
IROS | 4 |
| 2009 | Improving the SIFT descriptor with smooth derivative filters
Plinio Moreno, Alexandre Bernardino, José Santos-Victor |
Pattern Recognit. Lett. | 3 |
| 2008 | Multimodal saliency-based bottom-up attention a framework for the humanoid robot iCubabstractThis work presents a multimodal bottom-up attention system for the humanoid robot iCub where the robot's decisions to move eyes and neck are based on visual and acoustic saliency maps. We introduce a modular and distributed software architecture which is capable of fusing visual and acoustic saliency maps into one egocentric frame of reference. This system endows the iCub with an emergent exploratory behavior reacting to combined visual and auditory saliency. The developed software modules provide a flexible foundation for the open iCub platform and for further experiments and developments, including higher levels of attention and representation of the peripersonal space. Jonas Ruesch, Manuel Lopes 0001, Alexandre Bernardino, Jonas Hörnstein, José Santos-Victor, Rolf Pfeifer |
ICRA | 5 |
| 2008 | Learning Object Affordances: From Sensory-Motor Coordination to ImitationabstractAffordances encode relationships between actions, objects, and effects. They play an important role on basic cognitive capabilities such as prediction and planning. We address the problem of learning affordances through the interaction of a robot with the environment, a key step to understand the world properties and develop social skills. We present a general model for learning object affordances using Bayesian networks integrated within a general developmental architecture for social robots. Since learning is based on a probabilistic model, the approach is able to deal with uncertainty, redundancy, and irrelevant information. We demonstrate successful learning in the real world by having an humanoid robot interacting with objects. We illustrate the benefits of the acquired knowledge in imitation games. Luis Montesano, Manuel Lopes 0001, Alexandre Bernardino, José Santos-Victor |
IEEE Trans. Robotics | 4 |
| 2007 | A unified approach to speech production and recognition based on articulatory motor representationsabstractWe present a unified approach for speech production and recognition based on articulatory motor representations. The approach is inspired by the motor theory and the discovery of mirror neurons, and use motor representations for both reproduction and recognition of speech. A model of the vocal tract is used to create sound and the created sound is then mapped back to the motor representation using a neural network. To learn the map we mimic the behavior of a child that uses a combination of babbling and interaction with its caregiver to learn how to speak. Several different phases of babbling and interaction are identified and described. These help to overcome the inversion problem. The approach has been implemented on a humanoid robot, which has successfully learned to pronounce Swedish and Portuguese vowels. We have also studied how the different phases of babbling and interaction effect the error of the map and the achieved recognition rate when presented with vowels from different subjects. Finally we compare the recognition rates obtained using motor space with recognition rates obtained by directly using the acoustic parameters. Jonas Hörnstein, José Santos-Victor |
IROS | 2 |
| 2007 | Modeling affordances using Bayesian networksabstractAffordances represent the behavior of objects in terms of the robot's motor and perceptual skills. This type of knowledge plays a crucial role in developmental robotic systems, since it is at the core of many higher level skills such as imitation. In this paper, we propose a general affordance model based on Bayesian networks linking actions, object features and action effects. The network is learnt by the robot through interaction with the surrounding objects. The resulting probabilistic model is able to deal with uncertainty, redundancy and irrelevant information. We evaluate the approach using a real humanoid robot that interacts with objects. Luis Montesano, Manuel Lopes 0001, Alexandre Bernardino, José Santos-Victor |
IROS | 4 |
| 2007 | A Developmental Roadmap for Learning by Imitation in RobotsabstractIn this paper, we present a strategy whereby a robot acquires the capability to learn by imitation following a developmental pathway consisting on three levels: 1) sensory-motor coordination; 2) world interaction; and 3) imitation. With these stages, the system is able to learn tasks by imitating human demonstrators. We describe results of the different developmental stages, involving perceptual and motor skills, implemented in our humanoid robot, Baltazar. At each stage, the system's attention is drawn toward different entities: its own body and, later on, objects and people. Our main contributions are the general architecture and the implementation of all the necessary modules until imitation capabilities are eventually acquired by the robot. Also, several other contributions are made at each level: learning of sensory-motor maps for redundant robots, a novel method for learning how to grasp objects, and a framework for learning task description from observation for program-level imitation. Finally, vision is used extensively as the sole sensing modality (sometimes in a simplified setting) avoiding the need for special data-acquisition hardware. Manuel Lopes 0001, José Santos-Victor |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2006 | Design of the Robot-cub (iCub) HeadabstractThis paper describes the design of a robot head, developed in the framework of the RobotCub project. This project goals consists on the design and construction of a humanoid robotic platform, the iCub, for studying human cognition. The final platform would be approximately 90 cm tall, with 23 kg and with a total number of 53 degrees of freedom. For its size, the iCub is the most complete humanoid robot currently being designed, in terms of kinematic complexity. The eyes can also move, as opposed to similarly sized humanoid platforms. Specifications are made based on biological anatomical and behavioral data, as well as tasks constraints. Different concepts for the neck design (flexible, parallel and serial solutions) are analyzed and compared with respect to the specifications. The eye structure and the proprioceptive sensors are presented, together with some discussion of preliminary work on the face design Ricardo Beira, Manuel Lopes 0001, Miguel Praça, José Santos-Victor, Alexandre Bernardino, Giorgio Metta, Francesco Becchi, Roque J. Saltarén |
ICRA | 4 |
| 2006 | Sound Localization for Humanoid Robots - Building Audio-Motor Maps based on the HRTFabstractBeing able to locate the origin of a sound is important for our capability to interact with the environment. Humans can locate a sound source in both the horizontal and vertical plane with only two ears, using the head related transfer function HRTF, or more specifically features like interaural time difference ITD, interaural level difference ILD, and notches in the frequency spectra. In robotics notches have been left out since they are considered complex and difficult to use. As they are the main cue for humans' ability to estimate the elevation of the sound source this have to be compensated by adding more microphones or very large and asymmetric ears. In this paper, we present a novel method to extract the notches that makes it possible to accurately estimate the location of a sound source in both the horizontal and vertical plane using only two microphones and human-like ears. We suggest the use of simple spiral-shaped ears that has similar properties to the human ears and make it easy to calculate the position of the notches. Finally we show how the robot can learn its HRTF and build audiomotor maps using supervised learning and how it automatically can update its map using vision and compensate for changes in the HRTF due to changes to the ears or the environment. Jonas Hörnstein, Manuel Lopes 0001, José Santos-Victor, Francisco Lacerda |
IROS | 3 |
| 2006 | Learning Sensory-Motor Maps for Redundant RobotsabstractHumanoid robots are routinely engaged in tasks requiring the coordination between multiple degrees of freedom and sensory inputs, often achieved through the use of sensorymotor maps (SMMs). Most of the times, humanoid robots have more degrees of freedom (DOFs) available than those necessary to solve specific tasks. Notwithstanding, the majority of approaches for learning these SMMs do not take that into account. At most, the redundant degrees of freedom (degrees of redundancy, DOR) are "frozen" with some auxiliary criteria or heuristic rule. We present a solution to the problem of learning the forward/backward model, when the map is not injective, as in redundant robots. We propose the use of a "Minimum order SMM" that takes the desired image configuration and the DORs as input variables, while the non-redundant DOFs are viewed as outputs. Since the DORs are not frozen in this process, they can be used to solve additional tasks or criteria. This method provides a global solution for positioning a robot in the workspace, without the need to move in an incremental way. We provide examples where these tasks correspond to optimization criteria that can be solved online. We show how to learn the "Minimum Order SMM" using a local statistical learning method. Extensive experimental results with a humanoid robot are discussed to validate the approach, showing how to learn the Minimum Order SMM of a redundant system and using the redundancy to accomplish auxiliary tasks. Manuel Lopes 0001, José Santos-Victor |
IROS | 2 |
| 2006 | Jacobian Learning Methods for Tasks Sequencing in Visual ServoingabstractIn this paper, the coupling between Jacobian learning and task sequencing through the redundancy approach is studied. It is well known that visual servoing is robust to modeling errors in the Jacobian matrices. This justifies why Jacobian estimation does not usually degrade the system convergence. However, we show that this is not true any more when the redundancy formalism is used. In this case the Jacobian matrix is also necessary to compute projection operators for task decomposition, which is quite sensitive to errors. We show that learning improves the servoing performance, when task sequencing is used. Conversely, sequencing improves the convergence of learning, especially for tasks involving several degrees of freedom. Eye-in-hand and eye-to-hand experiments have been performed on two robots with six degrees of freedom Nicolas Mansard, Manuel Lopes 0001, José Santos-Victor, François Chaumette |
IROS | 3 |
| 2006 | Fast IIR Isotropic 2-D Complex Gabor Filters With Boundary InitializationabstractGabor filters are widely applied in image analysis and computer vision applications. This paper describes a fast algorithm for isotropic complex Gabor filtering that outperforms existing implementations. The main computational improvement arises from the decomposition of Gabor filtering into more efficient Gaussian filtering and sinusoidal modulations. Appropriate filter initial conditions are derived to avoid boundary transients, without requiring explicit image border extension. Our proposal reduces up to 39% the number of required operations with respect to state-of-the-art approaches. A full C++ implementation of the method is publicly available. Alexandre Bernardino, José Santos-Victor |
IEEE Trans. Image Process. | 2 |
| 2006 | Detection and classification of highway lanes using vehicle motion trajectoriesabstractIntelligent vision-based traffic surveillance systems are assuming an increasingly important role in highway monitoring and road management schemes. This paper describes a low-level object tracking system that produces accurate vehicle motion trajectories that can be further analyzed to detect lane centers and classify lane types. Accompanying techniques for indexing and retrieval of anomalous trajectories are also derived. The predictive trajectory merge-and-split algorithm is used to detect partial or complete occlusions during object motion and incorporates a Kalman filter that is used to perform vehicle tracking. The resulting motion trajectories are modeled using variable low-degree polynomials. A K-means clustering technique on the coefficient space can be used to obtain approximate lane centers. Estimation bias due to vehicle lane changes can be removed using robust estimation techniques based on Random Sample Consensus (RANSAC). Through the use of nonmetric distance functions and a simple directional indicator, highway lanes can be classified into one of the following categories: entry, exit, primary, or secondary. Experimental results are presented to show the real-time application of this approach to multiple views obtained by an uncalibrated pan-tilt-zoom traffic camera monitoring the junction of two busy intersecting highways. José Melo, Andrew Naftel, Alexandre Bernardino, José Santos-Victor |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2005 | A strategy for building topological maps through scene observation
Roger Freitas, Mário Sarcinelli Filho, Teodiano Freire Bastos-Filho, José Santos-Victor |
ICINCO | 4 |
| 2005 | Cooperative localization by fusing vision-based bearing measurements and motionabstractThis paper presents a method to cooperatively localize pairs of robots fusing bearing-only information provided by cameras and the motion of the vehicles. The algorithm uses the robots as landmarks to estimate their relative location. Bearings are the simplest measurements directly obtained from the cameras, as opposed to measuring depths which would require knowledge or reconstruction of the world structure. We present the general recursive Bayes estimator and three different implementations based on an extended Kalman filter, a particle filter and a combination of both techniques. We have compared the performance of the different implementations using real data acquired with two platforms equipped with omnidirectional cameras and simulated data. Luis Montesano, José António Gaspar, José Santos-Victor, Luis Montano |
IROS | 3 |
| 2005 | Least-squares 3D reconstruction from one or more views and geometric clues
Etienne Grossmann, José Santos-Victor |
Comput. Vis. Image Underst. | 2 |
| 2005 | Visual learning by imitation with motor representationsabstractWe propose a general architecture for action (mimicking) and program (gesture) level visual imitation. Action-level imitation involves two modules. The viewpoint Transformation (VPT) performs a "rotation" to align the demonstrator's body to that of the learner. The Visuo-Motor Map (VMM) maps this visual information to motor data. For program-level (gesture) imitation, there is an additional module that allows the system to recognize and generate its own interpretation of observed gestures to produce similar gestures/goals at a later stage. Besides the holistic approach to the problem, our approach differs from traditional work in i) the use of motor information for gesture recognition; ii) usage of context (e.g., object affordances) to focus the attention of the recognition system and reduce ambiguities, and iii) use iconic image representations for the hand, as opposed to fitting kinematic models to the video sequence. This approach is motivated by the finding of visuomotor neurons in the F5 area of the macaque brain that suggest that gesture recognition/imitation is performed in motor terms (mirror) and rely on the use of object affordances (canonical) to handle ambiguous actions. Our results show that this approach can outperform more conventional (e.g., pure visual) methods. Manuel Lopes 0001, José Santos-Victor |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2004 | An anthropomorphic robot torso for imitation: design and experimentsabstractWe describe the design of an anthropomorphic robot, combining a binocular head, an arm and a hand, for research in visuomotor coordination and learning by imitation. Our goal was to produce a system resembling the human arm-hand kinematics as closely as possible, while keeping it simple and relatively low-cost. We present mechanical details, kinematics and sensors together with a discussion of the main design options. We present results with human-arm coordination, as well as imitation of a human demonstrator, in real time. Manuel Lopes 0001, Ricardo Beira, Miguel Praça, José Santos-Victor |
IROS | 4 |
| 2003 | Visual transformations in gesture imitation: what you see is what you doabstractWe propose an approach for a robot to imitate the gestures of a human demonstrator. Our framework consists solely of two components: a Sensory-Motor Map (SMM) and a View-Point Transformation (VPT). The SMM establishes an association between an arm image and the corresponding joint angles and it is learned by the system during a period of observation of its own gestures. The VPT is widely discussed in the psychology of visual perception and is used to transform the image of the demonstrator's arm to the so-called ego-centric image, as if the robot were observing its own arm. Different structures of the SMM and VPT are proposed in accordance with observations in human imitation. The whole system relies on monocular visual information and leads to a parsimonious architecture for learning by imitation. Real-time results are presented and discussed. Manuel Lopes 0001, José Santos-Victor |
ICRA | 2 |
| 2002 | Maximum Likelihood 3D Reconstruction from One or More Images under Geometric ConstraintsabstractWe address the 3D reconstruction of scenes in which some planarity, collinearity, symmetry and other geometric properties are known-priori. Our main contribution is a reconstruction method that has advantages of both constraintbased and model-based methods. Etienne Grossmann, José Santos-Victor |
BMVC | 2 |
| 2002 | Multiple Plane Segmentation Using Optical FlowabstractIn this paper we present a motion based segmentation algorithm to automatically detect multiple planes from sparse optical ¤ow information. An optimal estimate for planar motion in the presence of additive Gaussian noise is £rst proposed, in-cluding directional uncertainty of the measurements (thus coping with the aper-ture problem) and a multi-frame (n> 2) setting (adding overall robustness). In the presence of multiple planes in motion, the residuals of the motion estimation model are used in a clustering algorithm to segment the different planes. The image motion parameters are used to £nd an initial cluster of features be-longing to a surface, which is then grown towards the surface borders. Initialization is random and only robust statistics and continuity constraints are used. There is no need for using and tuning thresholds. Since the exact parametric planar ¤ow model is used, the algorithm is able to cope ef£ciently with projective distortions and 3D motion and structure can be directly estimated. 1 Marco Zucchelli, José Santos-Victor, Henrik I. Christensen |
BMVC | 2 |
| 2002 | Reactive Navigation for Non-Holonomic Robots using the Ego-Kinematic SpaceabstractWe address the problem of applying reactive navigation methods to non-holonomic robots. Rather than embedding the motion constraints when designing a navigation method, we propose to introduce the robot's kinematic constraints directly in the spatial representation. In this space - the ego-kinematic space - the robot moves as a "free-flying object". Hence, standard reactive navigation methods applied to this space will automatically take into account the robot's kinematic constraints, without additional modifications. This methodology can be used with a large class of constrained mobile platforms (e.g. differential-driven robots, car-like robots, tri-cycle robots). We show experiments involving non-holonomic robots with two reactive navigation methods whose original formulation does not take the robot kinematic constraints into account (the Nearness Diagram Navigation and a Potential Field method). Javier Minguez, Luis Montano, José Santos-Victor |
ICRA | 3 |
| 2002 | Using motor representations for topological mapping and navigationabstractWe propose the use of motor vocabulary, that express a robot's specific motor capabilities, for topological mapbuilding and navigation. First, the motor vocabulary is created automatically through an imitation behaviour where the robot learns about its own motor repertoire, by following a tutor and associating its own motion perception to motor words. The learnt motor representation is then used for building the topological map. The robot is guided through the environment and automatically captures relevant (omnidirectional) images and associates motor words to links between places in the topological map. Finally, the created map is used for navigation, by invoking sequences of motor words that represent the actions for reaching a desired goal. In addition, a reflex-type behaviour based on optical flow extracted from omnidirectional images is used to avoid lateral collisions during navigation. The relation between motor vocabulary and imitation is stressed by the recent findings in neurophysiology, of visuomotor (mirror) neurons that may represent an internal motor representation related to the animal's capacity of imitation. This approach provides a natural adaptation between the robot's motion capabilities, the environment representations (maps) and navigation processes. Encouraging results are presented and discussed. Raquel Frizera Vassallo, José Santos-Victor, Hans J. Schneebeli |
IROS | 2 |
| 2001 | Algebraic Aspects of Reconstruction of 3D Scenes from One or More Views
Etienne Grossmann, Diego Ortin, José Santos-Victor |
BMVC | 3 |
| 2001 | Visual attention-based robot navigation using information samplingabstractPresents a method whereby an autonomous mobile robot automatically selects the most informative data from a set of images acquired a priori, using a statistical method termed information sampling. These data could be a single pixel or a number scattered throughout an image. This information is then used to build a topological map of the environment. Our sole input data are omnidirectional images obtained from a catadioptric panoramic camera. Experimental results show that by using only the best data the topological position of a robot, visually maneuvering through a simple indoor environment, can easily be determined. Niall Winters, José Santos-Victor |
IROS | 2 |
| 2001 | Vision-based Navigation, Environmental Representations and Imaging Geometries
José Santos-Victor, Alexandre Bernardino |
ISRR | 1 |
| 2000 | Dual Representations for Vision-Based 3D ReconstructionabstractWe consider the problem of representing sets of 3D points in the context of 3D reconstruction from point matches. We present a new representation for sets of 3D points, which is general, compact and expressive : any set of points can be represented; geometric relations that are often present in manmade scenes, such as coplanarity, alignment and orthogonality, are explicitly expressed. In essence, we propose to define each 3D point by three independent linear constraints that it verifies, and exploit the fact that coplanar points verify a common constraint. We show how to use the dual representation in Maximum Likelihood estimation, and that it substantially improves the precision of 3D reconstruction. Etienne Grossmann, José Santos-Victor |
BMVC | 2 |
| 2000 | Intrinsic Images for Dense Stereo Matching with Occlusions
César Silva, José Santos-Victor |
ECCV (1) | 2 |
| 2000 | A Closed-Form Solution for Paraperspective ReconstructionabstractWe address the problem of 3D reconstruction from image features tracked along a sequence. The most precise algorithms compute the maximum likelihood (ML) estimate and are iterative. They need an approximate 3D reconstruction as starting position. For that purpose, we propose a closed-form expression of paraperspective reconstruction. A matrix that approximately verifies the properties of a paraperspective projection matrix is first built, as in Christy and Horaud (1994) or Poelman and Kanade (1997). Our contribution lies in showing how to transform this matrix so that it exactly verifies the properties of paraperspective projection matrices. This is done by a closed form expression, in which the depth of the camera is also retrieved. The camera position is then found directly, instead of being obtained as the solution of a non-linear optimization problem, like in Poelman and Kanade (1997). Etienne Grossmann, José Santos-Victor |
ICPR | 2 |
| 2000 | Vision based station keeping and docking for an aerial blimpabstractThis paper describes a method for station keeping and docking of a lighter-than-air vehicle based on visual input. Due to the motion disturbances in the environment (currents), these tasks are important to keep the vehicle stabilized relative to an external reference frame. The main difficulties to achieve station keeping and docking are related to the nonholonomic constraints of the blimp moving in 3D, having a limited number of controllable degrees of freedom. The relative position of the vehicle with respect to a docking station is tracked using vision. A planar surface is chosen as a reference plane which allows visual tracking of an environmental region, based on planar projective transformations. An image-based control law is proposed together with a dynamic model for the vehicle. Experiments and results are described and discussed. Sjoerd van der Zwaan, Alexandre Bernardino, José Santos-Victor |
IROS | 3 |
| 2000 | Underwater Video Mosaics as Visual Navigation Maps
Nuno Gracias, José Santos-Victor |
Comput. Vis. Image Underst. | 2 |
| 2000 | Uncertainty analysis of 3D reconstruction from uncalibrated views
Etienne Grossmann, José Santos-Victor |
Image Vis. Comput. | 2 |
| 2000 | Vision-based navigation and environmental representations with an omnidirectional cameraabstractProposes a method for the visual-based navigation of a mobile robot in indoor environments, using a single omnidirectional (catadioptric) camera. The geometry of the catadioptric sensor and the method used to obtain a bird's eye (orthographic) view of the ground plane are presented. This representation significantly simplifies the solution to navigation problems, by eliminating any perspective effects. The nature of each navigation task is taken into account when designing the required navigation skills and environmental representations. We propose two main navigation modalities: topological navigation and visual path following. Topological navigation is used for traveling long distances and does not require knowledge of the exact position of the robot but rather, a qualitative position on the topological map. The navigation process combines appearance based methods and visual servoing upon some environmental features. Visual path following is required for local, very precise navigation, e.g., door traversal, docking. The robot is controlled to follow a prespecified path accurately, by tracking visual landmarks in bird's eye views of the ground plane. By clearly separating the nature of these navigation tasks, a simple and yet powerful navigation system is obtained. José António Gaspar, Niall Winters, José Santos-Victor |
IEEE Trans. Robotics Autom. | 3 |
| 1999 | Topological Maps for Visual Navigation
José Santos-Victor, Raquel Frizera Vassallo, Hans J. Schneebeli |
ICVS | 1 |
| 1999 | Binocular tracking: integrating perception and controlabstractPresents an active binocular tracking system using log-polar images with contributions in both the perceptual and control aspects. The control part is based on the visual servoing framework, including kinematics and dynamics. We introduce a fixation constraint that simplifies the tracking problem by decoupling the visual kinematics and allowing us to express system dynamics in image coordinates. Simple dynamic controllers are designed for each degree of freedom directly from image features. In the perceptual part, we use a space variant sensor that emphasizes the center of the visual field (log-polar geometry). We present a disparity estimation algorithm for log-polar images and provide a theoretical analysis to illustrate the advantages of using space variant images. The overall system is implemented in the Medusa binocular head without any specific processing hardware. The use of log-polar images allows real-time performance (50 Hz). Tracking experiments are presented to illustrate system performance with different control strategies and objects of different shapes and motions. Alexandre Bernardino, José Santos-Victor |
IEEE Trans. Robotics Autom. | 2 |
| 1998 | The Precision of 3D Reconstruction from Uncalibrated ViewsabstractWe consider reconstruction algorithms using points tracked over a sequence of (at least three) images, to estimate the positions of the cameras (motion parameters), the 3D coordinates (structure parameters), and the calibration matrix of the cameras (calibration parameters). Many algorithms have been reported in literature, and there is a need to know how well they may perform. We show how the choice of assumptions on the camera intrinsic parameters (either fixed, or with a probabilistic prior) influences the precision of the estimator. We associate a Maximum Likelihood estimator to each type of assumptions, and derive analytically their covariance matrices, independently of any specific implementation. We verify that the obtained covariance matrices are realistic, and compare the relative performance of each type of estimator. 1 Introduction The problem of 3D reconstruction from images has drawn considerable attention. We focus on the problem of reconstruction from matched points (co... Etienne Grossmann, José Santos-Victor |
BMVC | 2 |
| 1998 | Egomotion Estimation Using Log-Polar ImagesabstractWe address the problem of egomotion estimation of a monocular observer moving with arbitrary translation and rotation in an unknown environment, using log-polar images. The method we propose is uniquely based on the spatio-temporal image derivatives, or the normal flow. Thus, we avoid computing the complete optical flow field, which is an ill-posed problem due to the aperture problem. We use a search paradigm based on geometric properties of the normal flow field, and consider a family of search subspaces to estimate the egomotion parameters. These algorithms are particularly well-suited for the log-polar image geometry, as we use a selection of special normal flow, vectors with simple representation in log-polar coordinates. This approach highlights the close coupling between algorithmic aspects and the sensor geometry (retina physiology), often, found in nature. Finally, we present and discuss a set of experiments, for various kinds of camera motions, which show encouraging results. César Silva, José Santos-Victor |
ICCV | 2 |
| 1998 | Egomotion estimation on a topological spaceabstractWe present an egomotion estimation method for a monocular observer moving with arbitrary motion. The method is uniquely based on the normal flow field and the estimation problem is defined in a particular topological space-the /spl Lscr/-space-to allow the use of global data for the final estimates. Additionally, robust statistics methods are used for motion model selection, and provide improved robustness to image noise and data outliers. César Silva, José Santos-Victor |
ICPR | 2 |
| 1997 | Robust visual tracking by an active observerabstractIn this paper we address the problem of tracking a moving target by a monocular observer. The ability to track a moving object has many applications in robotics, teleoperation, surveillance systems, human-machine interfaces, etc. Our goal was the development of a robust tracking system for practical (industrial) applications and therefore based on inexpensive hardware. The strategy we present is based on the integration of correlation based techniques together with active contours, using a Kalman filtering approach. The overall operating frequency is about 6 Hz. The system is robust both to changes in the illumination and multiple moving objects in cluttered environments. Results are presented and discussed. Artur M. Arsénio, José Santos-Victor |
IROS | 2 |
| 1997 | Visual Behaviors for Docking
José Santos-Victor, Giulio Sandini |
Comput. Vis. Image Underst. | 1 |
| 1997 | Robust Egomotion Estimation From the Normal Flow Using Search SubspacesabstractWe address the problem of egomotion estimation for a monocular observer moving under arbitrary translation and rotation, in an unknown environment. The method we propose is uniquely based on the spatio-temporal image derivatives, or the normal flow. We introduce a search paradigm which is based on geometric properties of the normal flow field, and consists in considering a family of search subspaces to estimate the egomotion parameters. Various algorithms are proposed within this framework. In order to decrease the noise sensitivity of the estimation methods, we use statistical tools, based on robust regression theory. Finally, we present and discuss a wide variety of experiments with synthetic and real images, for various kinds of camera motion. César Silva, José Santos-Victor |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1996 | Direct egomotion estimationabstractIn this paper, we address the problem of egomotion estimation for a monocular moving observer under arbitrary translation and rotation. This is a key step for the navigation of any autonomous system. The method proposed is uniquely based on the spatio-temporal image derivatives, and has O(n/sup 2/) complexity. We search the image for particular geometric properties of the normal flow, tightly connected to the egomotion parameters, in two successive steps. The /spl Psi/-line search algorithm determines the direction of the focus of expansion (FOE) and computes some information on the rotation parameters. Finally, the /spl Phi/-line search algorithm determines the FOE, and the individual values of the rotation. Various experiments with synthetic and real images are presented and discussed. César Silva, José Santos-Victor |
ICPR | 2 |
| 1996 | Vergence control for robotic heads using log-polar imagesabstractThis paper describes a real-time vergence control mechanism based on, log-polar images, developed for a robot head. The real-time control of active vision systems imposes strong constraints on the computational complexity of the vision algorithms. In this paper, we illustrate that vergence of a stereo head can be achieved at reduced computational cost using log-polar images. These images have higher resolution at the center, where the attention is focused on, and the rest of the visual field is covered at a coarser resolution, still enabling the detection of events at the image periphery. The main advantages of using a non-uniform image sampling mechanism, such as the log-polar images, are related both to perceptual and algorithm complexity issues. We show that, when using correlation measures to control vergence, log-polar images give better results than cartesian images. Additionally, as log-polar images are smaller, the computation time is reduced. Two algorithms for closed loop vergence control, using correlation measures over log-polar images, are proposed. In the test examples described in the paper, we compare the two algorithms and address the problem of designing an adequate log-polar sensor, by introducing a performance analysis to evaluate the various log-polar sensor layouts. Alexandre Bernardino, José Santos-Victor |
IROS | 2 |
| 1996 | Uncalibrated obstacle detection using normal flow
José Santos-Victor, Giulio Sandini |
Mach. Vis. Appl. | 1 |
| 1995 | Divergent stereo in autonomous navigation: From bees to robots
José Santos-Victor, Giulio Sandini, Francesca Curotto, Stefano Garibaldi |
Int. J. Comput. Vis. | 1 |
| 1993 | Divergent stereo for robot navigation: learning from beesabstractA qualitative approach to visually guided navigation based on the computation of optical flow field is presented. The approach is based on the use of two cameras mounted on a mobile robot and with the optical axis directed in opposite directions, such that the two visual fields do not overlap (divergent stereo). Range computation is based on the computation of the apparent image speed on images acquired during the robot's motion. An example of reflex-type control of motion, driven by differential estimation of the flow field measured by the two eyes, is presented. It is shown how a difficult task like navigating through a funneled corridor with obstacles is possible without the need for metric depth estimation.> José Santos-Victor, Giulio Sandini, Francesca Curotto, Stefano Garibaldi |
CVPR | 1 |
| 1993 | Robotic beesabstractThis work presents some experiments of a real-time navigation system driven by two cameras pointed laterally to the navigation direction (divergent stereo). The approach is based on the observation that the stereo set-up traditionally used in vision (i.e. with the optical axis pointing forward) may not be the best one for navigation, and particularly for continuous control of a mobile actor moving in unconstrained environment. For navigation purposes, the driving information is not distance (as it is obtainable by a stereo set-up) but motion and, more precisely, by optical flow information computed over different areas of the visual field. Following this idea, a mobile vehicle has been equipped with a pair of cameras looking laterally (much like honeybees) and a controller based on fast, real-time computation of optical flow, has been implemented. The control of the mobile robot (ROBEE) is based on the comparison between the apparent image speed of the left and the right eye. A detailed description of the control structure is presented to demonstrate the feasibility of the approach in driving the mobile robot within a very cluttered environment. A discussion about the potentialities of the approach and the implications in terms of sensor's structure is also presented. Giulio Sandini, José Santos-Victor, Francesca Curotto, Stefano Garibaldi |
IROS | 2 |
| 1992 | Generation of 3D Dense Depth Maps by Dynamic Vision
José Santos-Victor, João Sentieiro |
BMVC | 1 |