VLDB 2026 Research / reviewers in the wild / expert
Alexandre Bernardino
dblp:53/5306
· DBLP profile ↗
98ranked-venue papers
3as first author
24since 2021 · last 2026
0000-0003-3991-1269ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 70 · 1 first-author · 14 since 2021Systems, architecture and hardware · 33 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 15 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Values Across Contexts: Understanding How Older Adults Enact What Matters Through TechnologyabstractAs populations age and technology becomes more pervasive, understanding the alignment between older adults’ values and technology design is paramount. More research is needed to understand how older adults’ living contexts shape their values and the use of technology. To address this, through a multi-context study, we explored how values differ for older adults and how their context of living might influence the adoption and use of technology. We conducted 22 semi-structured interviews with older adults in various residential contexts. We show that older adults tend to prioritize the same core values across living contexts, yet how they express values in each context differs. Technology can amplify or inhibit key values. We describe implications for context-responsive technology and design for continuity, to allow older adults to continually uphold important values through technology use. Hugo Simão, David Gonçalves, Neeta M. Khanuja, Valentina Nisi, Alexandre Bernardino, Tiago João Vieira Guerreiro, Jodi Forlizzi |
CHI | 5 |
| 2026 | SemBA-FAST: Semantic-based Bayesian attention applied to foveal active visual search tasksabstractBoth robots and humans have visual sensors with limited fields of view that need to be controlled to explore the environment and search for objects. To make this process efficient, visual attention methods actively select the information that contributes the most to the success of the task. Two key factors are characteristic of human vision. First, sensors can have space-varying resolution to process only certain parts of the scene with high resolution. Second, the attentional focus is deployed at highly informative regions, e.g. highly conspicuous regions. In this paper, we propose the use of semantic information, readily available in state-of-the-art deep object detectors, as an effective method to guide visual target search tasks using foveal sensors, which we refer to as SemBA -FAST. Because state-of-the-art object detectors are trained in conventional Cartesian images, we propose methods to calibrate detections in foveated images without requiring retraining the deep models. The information collected across multiple saccades is fused using Bayesian filters that keep a semantic representation of the world with associated uncertainty, on which the next gaze direction is actively determined. The proposed model is compared with state-of-the-art saliency-based methods. Our results demonstrate that semantic information positively influences the performance of target-present visual search in static scenes, highlighting its importance in designing visual attention systems for robots. • Semantic information available on current deep learning models can be exploited in active visual search of known objects and brings advantages with respect to saliency-based models. • Foveal vision effectively reduces the amount of visual information to be processed. • Deep-learning pre-trained object detectors can be calibrated to foveal images with low computational effort. • Biologically inspired computational models provide better insights into human visual cognition. • Probabilistic framework for integrating information across multiple views and best next view planning that enhances interpretability and mathematical explainability. João Luzio, Alexandre Bernardino, Plinio Moreno |
Neurocomputing | 2 |
| 2025 | Principal Direction 2-Gaussian Fit
Nicola Greggio, Alexandre Bernardino |
ICPRAM | 2 |
| 2024 | Non-Verbal Cues on Robot-Group PersuasionabstractWhen integrating robots into human daily life, persuasive power can be essential. However, there are often group dynamics which can complicate persuasion. This study focuses on how non-verbal cues, specifically gaze and hand gestures, affect the persuasiveness of a social robot. We have designed a protocol to include non-verbal cues in the social robot Vizzy (head and eye gaze, hand gestures) and test them in a series of experiments using the paradigm of the "Desert Survival Challenge". The goal of the robot is to persuade the participants of the game into changing their answers whilst avoiding negative feelings. It is hypothesized that the nonverbal cues will help avoid psychological reactance without diminishing compliance to the verbal requests issued by the robot. This phenomenon has been verified before for single person persuasion, but it is yet to be tested on groups. Thus, the goal of this project is to verify the effect of non-verbal cues in group persuasion by a robot and comparing it to single person persuasion. The results showed that the robot’s gestures increased compliance by the group and the gaze behaviour decreased psychological reactance. Alexandra Gonçalves, Plinio Moreno, Jodi Forlizzi, Leonel Garcia-Marques, Alexandre Bernardino |
ICRA | 5 |
| 2024 | Physics-Informed Neural Network for Multirotor Slung Load Systems ModelingabstractRecent advances in aerial robotics have enabled the use of multirotor vehicles for autonomous payload transportation. Resorting only to classical methods to reliably model a quadrotor carrying a cable-slung load poses significant challenges. On the other hand, purely data-driven learning methods do not comply by design with the problem’s physical constraints, especially in states that are not densely represented in training data. In this work, we explore the use of physics-informed neural networks to learn an end-to-end model of the multirotor-slung-load system and, at a given time, estimate a sequence of the future system states. An LSTM encoder-decoder with an attention mechanism is used to capture the dynamics of the system. To guarantee the cohesiveness between the multiple predicted states of the system, we propose the use of a physics-based term in the loss function, which includes a discretized physical model derived from first principles together with slack variables that allow for a small mismatch between expected and predicted values. To train the model, a dataset using a real-world quadrotor carrying a slung load was curated and is made available. Prediction results are presented and corroborate the feasibility of the approach. The proposed method outperforms both the first principles physical model and a comparable neural network model trained without the physics regularization proposed. Gil Serrano, Marcelo Jacinto, José Ribeiro-Gomes, João Pinto, Bruno J. Guerreiro, Alexandre Bernardino, Rita Cunha |
ICRA | 6 |
| 2024 | MotionGPT: Human Motion Synthesis with Improved Diversity and Realism via GPT-3 PromptingabstractThere are numerous applications for human motion synthesis, including animation, gaming, robotics, or sports science. In recent years, human motion generation from natural language has emerged as a promising alternative to costly and labor-intensive data collection methods relying on motion capture or wearable sensors (e.g., suits). Despite this, generating human motion from textual descriptions remains a challenging and intricate task, primarily due to the scarcity of large-scale supervised datasets capable of capturing the full diversity of human activity.This study proposes a new approach, called MotionGPT, to address the limitations of previous text-based human motion generation methods by utilizing the extensive semantic information available in large language models (LLMs). We first pretrain a doubly text-conditional motion diffusion model on both coarse ("high-level") and detailed ("low-level") ground truth text data. Then during inference, we improve motion diversity and alignment with the training set, by zero-shot prompting GPT-3 for additional "low-level" details. Our method achieves new state-of-the-art quantitative results in terms of Fréchet Inception Distance (FID) and motion diversity metrics, and improves all considered metrics. Furthermore, it has strong qualitative performance, producing natural results. Code is available at https://github.com/humansensinglab/MotionGPT José Ribeiro-Gomes, Tianhui Cai, Zoltán Ádám Milacski, Aayush Prakash, Shingo Takagi 0001, Amaury Aubel, Daeil Kim, Alexandre Bernardino, Fernando De la Torre |
WACV | 9 |
| 2024 | Unsupervised incremental estimation of Gaussian mixture models with 1D split moves
Nicola Greggio, Alexandre Bernardino |
Pattern Recognit. | 2 |
| 2023 | Designing a Human-Centered Intelligent System to Monitor & Explain Abnormal Patterns of Older AdultsabstractOlder adult care technologies are increasingly explored to support the independent living of older adults by monitoring their abnormal activities and informing caregivers to provide intervention if necessary. However, the adoption of these technologies remains challenging due to several factors (e.g. lack of usability). In this work, we present a human-centered, intelligent system for older adult care. Our proposed designs of the system were created based on the findings from a focus group session with caregivers. This system monitors the abnormal activities of an older adult using wireless motion sensors and machine learning models. In addition, unlike previous work that only notifies an outcome of activity recognition and abnormal detection models to a caregiver, the system supports interactive dialogue responses to explain the abnormal activities of an older adult to a caregiver and allow the caregiver to elicit additional information about the older adult and the older adult to proactively share his/her status with the caregiver for an adequate intervention. Min Hun Lee, Daniel P. Siewiorek, Alexandre Bernardino |
ASSETS | 3 |
| 2023 | Adapt-FuseNet: Context-aware Multimodal Adaptive Fusion of Face and Gait Features using Attention Techniques for Human IdentificationabstractBiometrics plays a significant role in vision-based surveillance applications. Soft biometrics such as gait is widely used with face in surveillance tasks like person recognition and re-identification. Nevertheless, in practical scenarios, classical fusion techniques respond poorly to the changes in individual users, external environment and varying contexts such as viewpoints. To this end, we propose a novel context-aware adaptive multi-biometric fusion strategy viz., ‘Adapt-FuseNet’ for the dynamic incorporation of gait and face biometric cues leveraging attention techniques. In particular, we investigate the impact of attention models such as parallel co-attention & keyless attention, along with various fusion strategies such as naïve fusion & adaptive fusion for human identification. Extensive experiments are carried out on two publically available large gait datasets i.e. CASIA-A and CASIA-B. Results show the superior performance of our proposed context-aware adaptive fusion model compared with the state-of-the-art models. Ashwin Prakash, Thejaswin S, Athira Nambiar, Alexandre Bernardino |
IJCB | 4 |
| 2023 | 3DSGrasp: 3D Shape-Completion for Robotic GraspabstractReal-world robotic grasping can be done robustly if a complete 3D Point Cloud Data (PCD) of an object is available. However, in practice, PCDs are often incomplete when objects are viewed from few and sparse viewpoints before the grasping action, leading to the generation of wrong or inaccurate grasp poses. We propose a novel grasping strategy, named 3DSGrasp, that predicts the missing geometry from the partial PCD to produce reliable grasp poses. Our proposed PCD completion network is a Transformer-based encoder-decoder network with an Offset-Attention layer. Our network is inherently invariant to the object pose and point's permutation, which generates PCDs that are geometrically consistent and completed properly. Experiments on a wide range of partial PCD show that 3DSGrasp outperforms the best state-of-the-art method on PCD completion tasks and largely improves the grasping success rate in real-world scenarios. The code and dataset are available at: https://github.com/NunoDuarte/3DSGrasp. Seyed Saber Mohammadi, Nuno Ferreira Duarte, Dimitrios Dimou, Yiming Wang 0002, Matteo Taiana, Pietro Morerio, Atabak Dehban, Plinio Moreno, Alexandre Bernardino, Alessio Del Bue, José Santos-Victor |
ICRA | 9 |
| 2023 | Learning Open-Loop Saccadic Control of a 3D Biomimetic Eye Using the Actor-Critic AlgorithmabstractThe application of reinforcement learning algorithms to robotics has increased over the last decade, especially for the control of robots with non-linear dynamics and a redundant number of degrees of freedom using classic control techniques. Here we study the control of a biomimetic robotic eye with three extraocular muscle pairs as a prime example. Using an actor-critic algorithm, this paper aims to link reinforcement learning to this control problem, and create a framework that will learn the open-loop control of saccadic movements of the robotic eye. The basis for the implemented control is inspired by the primate physiological pulsed control signal, which is generated, integrated and sent to the appropriate muscles to perform the saccade. The metric that evaluates the saccadic output is also inspired by the primate oculomotor system and is used to shape the reward function. This methodology was applied to a simplified 3D physical model of the human eye as a proof of concept. The algorithm managed to learn a saccadic control strategy in 3D. The trajectories obtained, have similar non-linear dynamics as those recorded in humans and their 3D rotational kinematics are constrained by Listing's law. Henrique Granado, Reza Javanmard Alitappeh, Akhil John, A. John van Opstal, Alexandre Bernardino |
IROS | 5 |
| 2023 | Quantifying Object Detection Uncertainty in Autonomous Driving with Test-Time AugmentationabstractIn this paper, we propose the first Test-Time Augmentation (TTA) method to estimate uncertainty and improve accuracy of pre-trained object detectors in the autonomous driving domain. We show, in an autonomous driving dataset, that even simple color-based augmentations are able to improve mean Average Precision (mAP) performance with respect to other state-of-the-art methods with the same purpose (Monte Carlo (MC) Dropout and Output Redundancy). Furthermore, we show that the quality of the estimated uncertainty and distributions can be improved, both in our method and in the state-of-the-art, if some of the parameters of the methods (bounding box selection and clustering criteria) are independently tuned for the classification and the localization subtasks of object detection. Rui Magalhães, Alexandre Bernardino |
IV | 2 |
| 2023 | Fire images classification based on a handcraft approach
Houda Harkat, José M. P. Nascimento, Alexandre Bernardino, Hasmath Farhana Thariq Ahmed |
Expert Syst. Appl. | 3 |
| 2023 | Learning Performance Models of Distributed Computer Vision Methods for Decision Making in Detection and Tracking Algorithms in UAVsabstractUnmanned Aerial Vehicles (UAVs) are getting more and more uses in recent times. However, low-cost commercial UAVs may not possess enough computational power to run state of the art algorithms in order to perform certain tasks, negatively affecting performance. Remote computational systems, where heavy processing tasks can be offloaded emerge as a solution. However, they introduce latency, which can be undesirable for real-time tasks. Furthermore, if the task is simple, using a local algorithm with worse performance may be acceptable to avoid latency. As such, a method to decide which algorithm to use is of great importance. We consider the use case of computer vision tasks, in particular detection and tracking. In these tasks, image properties such as brightness, contrast, motion blur and clutter affect the algorithm performance. Our proposed methods use a combination of neural networks and kernel machines to estimate the performance of the algorithm given the input image. An appropriate cost function is then used to identify the best algorithm for the task given the input image, the task deadline, and the uncertainty in the variables of the algorithm, in particular computing time and error rate. Results show that our method matches or outperforms similar state of the art methods, complying with time restrictions while delivering increased performance. João Correia 0003, Alexandre Bernardino, Ricardo A. Ribeiro |
IEEE Internet Things J. | 2 |
| 2023 | Design, development, and evaluation of an interactive personalized social robot to monitor and coach post-stroke rehabilitation exercises
Min Hun Lee, Daniel P. Siewiorek, Asim Smailagic, Alexandre Bernardino, Sergi Bermúdez i Badia |
User Model. User Adapt. Interact. | 4 |
| 2022 | Comparison of Methodologies for Detecting Feeding Activity in Aquatic EnvironmentabstractIn this paper, several methodologies were applied to automatically detect feeding activity in the Oceanário de Lisboa Main Aquarium, using videos acquired on the outside of the aquarium by a static camera. We propose three methods. The first one is based on Convolutional Neural Networks (CNN) and learns to detect the feeding patterns at each frame based on training data. The second one is based on motion variability, computed either from frame difference or optical flow. The third one uses the analysis of spatial patterns (aggregation) formed by fish present in each frame, assuming their prior detection. For the development of the methods and quantitative evaluation, several videos were filmed at Oceanário de Lisboa. To evaluate each of the approaches, several metrics are extracted such as accuracy, precision, and recall. We analysed videos of the feeding patterns of rays (bottom feeding) and sharks (surface feeding). These different patterns are quite distinct in terms of motion and aggregation of fish. For bottom feeding, we concluded that the frame difference approach was the best performing, but the CNN and aggregation methods showed competitive results. For surface feeding, the CNN was the only method able to perform with reasonable accuracy. Gonçalo Adolfo, Alexandre Bernardino, Núria Baylina, H. Sofia Pinto |
ICPR | 2 |
| 2022 | Weakly Supervised Fire and Smoke Segmentation in Forest Images with CAM and CRFabstractThe number of publicly available datasets with annotated fire and smoke regions in wildfire scenarios is very scarce. To develop flexible models that can help firefighters protect the forest and nearby populations, we propose a method for segmenting fire and smoke regions in images using only image-level annotations, i.e. simple labels that just indicate the presence or absence of fire and smoke. The method uses Class Activation Mapping (CAM) on multi-label classifiers of fire and smoke, followed by Conditional Random Fields (CRF) to accurately detect fire/smoke masks at the pixel-level. Due to the high correlation of fire and smoke labels, we found that a single multi-label classifier is unable to provide simultaneously good segmentation for fire and smoke. Instead, we trained two classifiers of different complexities, one to support the segmentation of fire and the other for smoke. Compared with fully-supervised methods, the proposed weakly-supervised method is quite competitive, while requiring much less labeling effort. Bernardo Amaral, Milad Niknejad, Catarina Barata, Alexandre Bernardino |
ICPR | 4 |
| 2022 | Towards Efficient Annotations for a Human-AI Collaborative, Clinical Decision Support System: A Case Study on Physical Stroke Rehabilitation AssessmentabstractArtificial intelligence (AI) and machine learning (ML) algorithms are increasingly being explored to support various decision-making tasks in health (e.g. rehabilitation assessment). However, the development of such AI/ML-based decision support systems is challenging due to the expensive process to collect an annotated dataset. In this paper, we describe the development process of a human-AI collaborative, clinical decision support system that augments an ML model with a rule-based (RB) model from domain experts. We conducted its empirical evaluation in the context of assessing physical stroke rehabilitation with the dataset of three exercises from 15 post-stroke survivors and therapists. Our results bring new insights on the efficient development and annotations of a decision support system: when an annotated dataset is not available initially, the RB model can be used to assess post-stroke survivor’s quality of motion and identify samples with low confidence scores to support efficient annotations for training an ML model. Specifically, our system requires only 22 - 33% of annotations from therapists to train an ML model that achieves equally good performance with an ML model with all annotations from a therapist. Our work discusses the values of a human-AI collaborative approach for effectively collecting an annotated dataset and supporting a complex decision-making task. Min Hun Lee, Daniel P. Siewiorek, Asim Smailagic, Alexandre Bernardino, Sergi Bermúdez i Badia |
IUI | 4 |
| 2021 | A Human-AI Collaborative Approach for Clinical Decision Making on Rehabilitation AssessmentabstractAdvances in artificial intelligence (AI) have made it increasingly applicable to supplement expert’s decision-making in the form of a decision support system on various tasks. For instance, an AI-based system can provide therapists quantitative analysis on patient’s status to improve practices of rehabilitation assessment. However, there is limited knowledge on the potential of these systems. In this paper, we present the development and evaluation of an interactive AI-based system that supports collaborative decision making with therapists for rehabilitation assessment. This system automatically identifies salient features of assessment to generate patient-specific analysis for therapists, and tunes with their feedback. In two evaluations with therapists, we found that our system supports therapists significantly higher agreement on assessment (0.71 average F1-score) than a traditional system without analysis (0.66 average F1-score, p < 0.05). After tuning with therapist’s feedback, our system significantly improves its performance from 0.8377 to 0.9116 average F1-scores (p < 0.01). This work discusses the potential of a human-AI collaborative system to support more accurate decision making while learning from each other’s strengths. Min Hun Lee, Daniel P. Siewiorek, Asim Smailagic, Alexandre Bernardino, Sergi Bermúdez i Badia |
CHI | 4 |
| 2021 | Attention on Classification for Fire SegmentationabstractDetection and localization of fire in images and videos are important in tackling fire incidents. Although semantic segmentation methods can be used to indicate the location of pixels with fire in the images, their predictions are localized, and they often fail to consider global information of the existence of fire in the image which is implicit in the image labels. We propose a Convolutional Neural Network (CNN) for joint classification and segmentation of fire in images which improves the performance of the fire segmentation. We use a spatial self-attention mechanism to capture long-range dependency between pixels, and a new channel attention module which uses the classification probability as an attention weight. The network is jointly trained for both segmentation and classification, leading to improvement in the performance of the single-task image segmentation methods, and the previous methods proposed for fire segmentation. Milad Niknejad, Alexandre Bernardino |
ICMLA | 2 |
| 2021 | Fire Detection using Deeplabv3+ with Mobilenetv2abstractFire detection is high priority task in the current decade, due to the high occurrences of fire in urban and forest area. Every year, millions of hectares of forests are burned and destroyed. The cost of dislocation could be optimized by implementing an accurate detection system. In this paper a Deeplabv3+ model with a Mobilenetv2 backbone is implemented and tested over R GB and Infrared pictures of the Corsican french dataset. Three different types of loss function were used to overcome the problem of unbalanced dataset. The results obtained with the model herein presented are very encouraging. Houda Harkat, José M. P. Nascimento, Alexandre Bernardino |
IGARSS | 3 |
| 2021 | Human-Robot greeting: tracking human greeting mental states and acting accordinglyabstractMobile social robots should be able to engage in interaction with people effectively. However, greeting someone is a complex task since it implies an exchange of social signals. Adam Kendon modeled human greetings as a set of six phases: initiation of approach, distance salutation, head dip, approach, final approach, and close salutation. Based on Kendon’s model, we propose a system for mobile social robots that manages the greeting process through the exchange of social signals. A Hidden Markov Model keeps track of the greeting stage through the observation of the human gestures, while a behavior tree generates appropriate robot actions. We used publicly available datasets to train the Hidden Markov Model. Evaluation on test sets showed an average greeting phase estimation accuracy of 80.9%. We tested the full system (Hidden Markov Model + Behavior Tree) in simulation and in a real world pilot experiment using the Vizzy robot, and it recognized and replicated the correct phase with an accuracy of 91.8% and 53.8%, respectively. Manuel Carvalho, João Avelino, Alexandre Bernardino, Rodrigo M. M. Ventura, Plinio Moreno |
IROS | 3 |
| 2021 | LiDAR Data Noise Models and Methodology for Sim-to-Real Domain Generalization and Adaptation in Autonomous Driving PerceptionabstractIn autonomous driving, object detection and semantic segmentation are critical tasks for path planning and control of an autonomous vehicle. Recent approaches are based on supervised learning methods, with large datasets sampled in the target domain. However, annotating training data for supervised learning methods is a high resource and time-consuming task. In this work, we propose to exploit artificial LiDAR data for object detection and semantic segmentation. We use the CARLA simulator [1] to generate artificial data of autonomous driving scenarios and propose ways to mitigate the differences between artificial and real-world data (domain generalization). We modeled both the noise and the missed reflections (denoted point dropout) that occur in real-world data collection, and show their effects in the detection and segmentation tasks. We assess the potential benefits of using pre-trained models on artificial data when fine-tuning with all, or a fraction, of the available real-world data (domain adaptation). We find clear improvements when using artificial data to pretrain a network, which allows to use a reduced amount of realworld data, and boost the performance of the trained models. João Espadinha, Ivan Lebedev, Luka Lukic, Alexandre Bernardino |
IV | 4 |
| 2021 | Modelling 3D saccade generation by feedforward optimal controlabstractAn interesting problem for the human saccadic eye-movement system is how to deal with the degrees-of-freedom problem: the six extra-ocular muscles provide three rotational degrees of freedom, while only two are needed to point gaze at any direction. Measurements show that 3D eye orientations during head-fixed saccades in far-viewing conditions lie in Listing's plane (LP), in which the eye's cyclotorsion is zero (Listing's law, LL). Moreover, while saccades are executed as single-axis rotations around a stable eye-angular velocity axis, they follow straight trajectories in LP. Another distinctive saccade property is their nonlinear main-sequence dynamics: the affine relationship between saccade size and movement duration, and the saturation of peak velocity with amplitude. To explain all these properties, we developed a computational model, based on a simplified and upscaled robotic prototype of an eye with 3 degrees of freedom, driven by three independent motor commands, coupled to three antagonistic elastic muscle pairs. As the robotic prototype was not intended to faithfully mimic the detailed biomechanics of the human eye, we did not impose specific prior mechanical constraints on the ocular plant that could, by themselves, generate Listing's law and the main-sequence. Instead, our goal was to study how these properties can emerge from the application of optimal control principles to simplified eye models. We performed a numerical linearization of the nonlinear system dynamics around the origin using system identification techniques, and developed open-loop controllers for 3D saccade generation. Applying optimal control to the simulated model, could reproduce both Listing's law and and the main-sequence. We verified the contribution of different terms in the cost optimization functional to realistic 3D saccade behavior, and identified four essential terms: total energy expenditure by the motors, movement duration, gaze accuracy, and the total static force exerted by the muscles during fixation. Our findings suggest that Listing's law, as well as the saccade dynamics and their trajectories, may all emerge from the same common mechanism that aims to optimize speed-accuracy trade-off for saccades, while minimizing the total muscle force during eccentric fixation. Akhil John, Carlos Aleluia, A. John van Opstal, Alexandre Bernardino |
PLoS Comput. Biol. | 4 |
| 2020 | Highly sensitive bio-inspired sensor for fine surface exploration and characterizationabstractTexture sensing is one of the types of information sensed by humans through touch, and is thus of interest to robotics that this type of information can be acquired and processed. In this work we present a texture topography sensor based on a ciliary structure, a biological structure found in many organisms. The device consists of up to 9 elastic cilia with permanent magnetization assembled on top of a highly sensitive tunneling magnetoresistance (TMR) sensor, within a compact footprint of 6×6 mm2. When these cilia brush against some textured surface, their movement and vibrations give rise to a signal that can be correlated to the characteristics of the texture being measured. We also present an electronic signal acquisition board, used in this work. Various configurations of cilia sizes are tested, with the most precise being capable of differentiating different types of sandpaper from 9.2 μm to 213 μm average surface roughness with a 7 μm resolution. As a topography scanner the sensor was able to scan a 20 μm high step in a flat surface. Pedro Ribeiro 0006, Susana Cardoso, Alexandre Bernardino, Lorenzo Jamone |
ICRA | 3 |
| 2020 | Fruit quality control by surface analysis using a bio-inspired soft tactile sensorabstractThe growing consumer demand for large volumes of high quality fruit has generated an increasing need for auto-mated fruit quality control during production. Optical methods have been proved successful in a few cases, but with limitations related to the variability of fruit colors and lighting conditions during tests. Tactile sensing provides a valuable alternative, although it comes with the need of a physical interaction that could damage the fruit. To overcome these limitations, we propose the usage of a recently developed soft tactile sensor for non-invasive fruit quality control. The ability of the sensor to detect very small forces and to finely analyze surfaces allows the collection of relevant information about the fruit by performing a very delicate physical interaction, that does not cause any damage. We report experiments in which such information is used to determine whether apples and strawberries are ripe or senescent. We test different configurations of the sensor and different classification algorithms, achieving very good accuracy for both apples (96%) and strawberries (83%). Pedro Ribeiro 0006, Susana Cardoso, Alexandre Bernardino, Lorenzo Jamone |
IROS | 3 |
| 2020 | Towards Personalized Interaction and Corrective Feedback of a Socially Assistive Robot for Post-Stroke Rehabilitation TherapyabstractA robotic exercise coaching system requires the capability of automatically assessing a patient's exercise to interact with a patient and generate corrective feedback. However, even if patients have various physical conditions, most prior work on robotic exercise coaching systems has utilized generic, pre-defined feedback.This paper presents an interactive approach that combines machine learning and rule-based models to automatically assess a patient's rehabilitation exercise and tunes with patient's data to generate personalized corrective feedback. To generate feedback when an erroneous motion occurs, our approach applies an ensemble voting method that leverages predictions from multiple frames for frame-level assessment. According to the evaluation with the dataset of three stroke rehabilitation exercises from 15 post-stroke subjects, our interactive approach with an ensemble voting method supports more accurate frame-level assessment (p <; 0.01), but also can be tuned with held-out user's unaffected motions to significantly improve the performance of assessment from 0.7447 to 0.8235 average F1-scores over all exercises (p <; 0.01). This paper discusses the value of an interactive approach with an ensemble voting method for personalized interaction of a robotic exercise coaching system. Min Hun Lee, Daniel P. Siewiorek, Asim Smailagic, Alexandre Bernardino, Sergi Bermúdez i Badia |
RO-MAN | 4 |
| 2020 | An Exploratory Study on Techniques for Quantitative Assessment of Stroke Rehabilitation ExercisesabstractTechnology-assisted systems to monitor and assess rehabilitation exercises have an opportunity of enhancing rehabilitation practices by automatically collecting patient's quantitative performance data. However, even if a complex algorithm (e.g. Neural Network) is applied, it is still challenging to develop such a system due to patients with various physical conditions. The system with a complex algorithm is limited to be a black-box system that cannot provide explanations on its predictions. To address these challenges, this paper presents a hybrid model that integrates a machine learning (ML) model with a rule-based (RB) model as an explainable artificial intelligence (AI) technique for quantitative assessment of stroke rehabilitation exercises. For evaluation, we collected therapist's knowledge on assessment as 15 rules from interviews with therapists and the dataset of three upper-limb stroke rehabilitation exercises from 15 post-stroke and 11 healthy subjects using a Kinect sensor. Experimental results show that a hybrid model can achieve comparable performance with a ML model using Neural Network, but also provide explanations on a model prediction with a RB model. The results indicate the potential of a hybrid model as an explainable AI technique to support the interpretation of a model and fine-tune a model with user-specific rules for personalization. Min Hun Lee, Daniel P. Siewiorek, Asim Smailagic, Alexandre Bernardino, Sergi Bermúdez i Badia |
UMAP | 4 |
| 2020 | Co-Design and Evaluation of an Intelligent Decision Support System for Stroke Rehabilitation AssessmentabstractClinical decision support systems have the potential to improve work flows of experts in practice (e.g. therapist's evidence-based rehabilitation assessment). However, the adoption of these systems is challenging, and the gains of these systems have not fully demonstrated yet. In this paper, we identified the needs of therapists to assess patient's functional abilities (e.g. alternative perspectives with quantitative information on patient's exercise motions). As a result, we co-designed and developed an intelligent decision support system that automatically identifies salient features of assessment using reinforcement learning to assess the quality of motion and generate patient-specific analysis. We evaluated this system with seven therapists using the dataset from 15 patients performing three exercises. The results show that therapists have higher usage intent on our system than a traditional system without patient-specific analysis ($p < 0.05$). While presenting richer information ($p < 0.10$), our system significantly reduces therapists' effort on assessment ($p < 0.10$) and improves their agreement on assessment from 0.66 to 0.71 F1-scores ($p < 0.01$). This work discusses the importance of human centered design and development of a machine learning-based decision support system that presents contextually relevant information and salient explanations on its prediction for better adoption in practice. Min Hun Lee, Daniel P. Siewiorek, Asim Smailagic, Alexandre Bernardino, Sergi Bermúdez i Badia |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2019 | The Impact of Domain Randomization on Object Detection: A Case Study on Parametric Shapes and Synthetic Textures*abstractRecent advances in deep learning-based object detection techniques have revolutionized their applicability in several fields. However, since these methods rely on unwieldy and large amounts of data, a common practice is to download models pre-trained on standard datasets and fine-tune them for specific application domains with a small set of domain-relevant images. In this work, we show that using synthetic datasets that are not necessarily photo-realistic can be a better alternative to simply fine-tune pre-trained networks. Specifically, our results show an impressive 25%improvement in the mAP metric over a fine-tuning baseline when only about 200 labelled images are available to train. Finally, an ablation study of our results is presented to delineate the individual contribution of different components in the randomization pipeline. Atabak Dehban, João Borrego, Rui Pimentel de Figueiredo, Plinio Moreno, Alexandre Bernardino, José Santos-Victor |
IROS | 5 |
| 2019 | Learning to assess the quality of stroke rehabilitation exercisesabstractDue to the limited number of therapists, task-oriented exercises are often prescribed for post-stroke survivors as in-home rehabilitation. During in-home rehabilitation, a patient may become unmotivated or confused to comply prescriptions without the feedback of a therapist. To address this challenge, this paper proposes an automated method that can achieve not only qualitative, but also quantitative assessment of stroke rehabilitation exercises. Specifically, we explored a threshold model that utilizes the outputs of binary classifiers to quantify the correctness of a movements into a performance score. We collected movements of 11 healthy subjects and 15 post-stroke survivors using a Kinect sensor and ground truth scores from primary and secondary therapists. The proposed method achieves the following agreement with the primary therapist: 0.8436, 0.8264, and 0.7976 F1-scores on three task-oriented exercises. Experimental results show that our approach performs equally well or better than multi-class classification, regression, or the evaluation of the secondary therapist. Furthermore, we found a strong correlation (R2 = 0.95) between the sum of computed exercise scores and the Fugl-Meyer Assessment scores, clinically validated motor impairment index of post-stroke survivors. Our results demonstrate a feasibility of automatically assessing stroke rehabilitation exercises with the decent agreement levels and clinical relevance. Min Hun Lee, Daniel P. Siewiorek, Asim Smailagic, Alexandre Bernardino, Sergi Bermúdez i Badia |
IUI | 4 |
| 2019 | A Data Set for Airborne Maritime Surveillance EnvironmentsabstractThis paper presents a data set with surveillance imagery over the sea captured by a small size UAV. This data set presents the object examples ranging from cargo ships, small boats, life rafts to hydrocarbon slick. The video sequences were captured using different types of cameras, at different heights, and different perspectives. The data set also contains thousands of labels with positions of objects of interest. This was only possible to achieve with the labeling tool also described in this paper. Additionally, using standard evaluation frameworks, we establish a baseline of results using algorithms developed by the authors, which are better adapted to the maritime environment. Ricardo A. Ribeiro, Gonçalo Cruz, Jorge Matos, Alexandre Bernardino |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Learning Temporal Features for Detection on Maritime Airborne Video Sequences Using Convolutional LSTMabstractIn this paper, we study the effectiveness of learning temporal features to improve detection performance in videos captured by small aircraft. To implement this learning process, we use a convolutional long short-term memory (LSTM) associated with a pretrained convolutional neural network (CNN). To improve the training process, we incorporate domain-specific knowledge about the expected size and number of boats. We carry out three tests. The first searches the best sequence length and subsampling rate for training and the second compares the proposed method with a traditional CNN, a traditional LSTM, and a gated recurrent unit (GRU). The final test evaluates our method with the already published detectors in two data sets. Results show that in favorable conditions, our method's performance is comparable to other detectors but, on more challenging environments, it stands out from other techniques. Gonçalo Cruz, Alexandre Bernardino |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | AHA-3D: A Labelled Dataset for Senior Fitness Exercise Recognition and Segmentation from 3D Skeletal Data
João Antunes, Alexandre Bernardino, Asim Smailagic, Daniel P. Siewiorek |
BMVC | 2 |
| 2018 | Playdough to Roombots: Towards a Novel Tangible User Interface for Self-reconfigurable Modular RobotsabstractOne of the main strengths of self-reconfigurable modular robots (SRMR) is their ability to shape-shift and dynamically change their morphology. In the case of our SRMR system “Roombots”, these shapes can be quite arbitrary for a great variety of tasks while the major utility is envisioned to be self-reconfigurable furniture. As such, the ideas and inspirations from users quickly need to be translated into the final Roombots shape. This involves a multitude of separate processes and - most importantly - requires an intuitive user interface. Our current approach led to the development of a tangible user interface (TUI) which involves 3D-scanning of a shape formed by modeling clay and the necessary steps to prepare the digitized model to be formed by Roombots. The system is able to generate a solution in less than two minutes for our target use as demonstrated with various examples. Mehmet Mutlu, Simon Hauser, Alexandre Bernardino, Auke Jan Ijspeert |
ICRA | 3 |
| 2018 | The Power of a Hand-shake in Human-Robot InteractionsabstractIn this paper, we study the influence of a handshake in the human emotional bond to a robot. In particular, we evaluate the human willingness to help a robot whether the robot first introduces itself to the human with or without a handshake. In the tested paradigm the robot and the human have to perform a joint task, but at a certain stage, the robot needs help to navigate through an obstacle. Without requesting explicit help from the human, the robot performs some attempts to navigate through the obstacle, suggesting to the human that it requires help. In a study with 45 participants, we measure the human's perceptions of the social robot Vizzy, comparing the handshake vs non-handshake conditions. In addition, we evaluate the influence of a handshake in the pro-social behaviour of helping it and the willingness to help it in the future. The results show that a handshake increases the perception of Warmth, Animacy, Likeability, and the tendency to help the robot more, by removing the obstacle. João Avelino, Plinio Moreno, Alexandre Bernardino, Filipa Correia, Ana Paiva 0001, João Catarino, Pedro Ribeiro 0006 |
IROS | 3 |
| 2018 | Finding safe 3D robot grasps through efficient haptic exploration with unscented Bayesian optimization and collision penaltyabstractRobust grasping is a major, and still unsolved, problem in robotics. Information about the 3D shape of an object can be obtained either from prior knowledge (e.g., accurate models of known objects or approximate models of familiar objects) or real-time sensing (e.g., partial point clouds of unknown objects) and can be used to identify good potential grasps. However, due to modeling and sensing inaccuracies, local exploration is often needed to refine such grasps and successfully apply them in the real world. The recently proposed unscented Bayesian optimization technique can make such exploration safer by selecting grasps that are robust to uncertainty in the input space (e.g., inaccuracies in the grasp execution). Extending our previous work on 2D optimization, in this paper we propose a 3D haptic exploration strategy that combines unscented Bayesian optimization with a novel collision penalty heuristic to find safe grasps in a very efficient way: while by augmenting the search-space to 3D we are able to find better grasps, the collision penalty heuristic allows us to do so without increasing the number of exploration steps. João Castanheira, Pedro Vicente, Ruben Martinez-Cantin, Lorenzo Jamone, Alexandre Bernardino |
IROS | 5 |
| 2017 | Context-Aware Person Re-Identification in the Wild Via Fusion of Gait and Anthropometric FeaturesabstractIn this work, we present a context-aware ensemble fusion framework based on soft-biometric features, for long term person re-identification (Re-ID) in wild surveillance scenarios. The characteristics of a person that best correlate to its identity depend strongly on the view point. For instance, a person with a short stride gait is better perceived from a lateral view, whereas a person with a large chest is more distinct from a frontal view. Thus we associate context to the viewing direction of walking people in a surveillance scenario and choose the best features for each case. Using the MS KinectTM sensor v.2, we collect data from walking subjects and extract associated anthropometric and gait features. Each context is analysed with a Feature selection technique (Sequential Forward Selection) so that only the most relevant features for the context are retained. Then, individual context-specific classifiers are trained leveraging those selected features. Finally, we propose a contextaware ensemble fusion strategy, which we term as 'Contextspecific score-level fusion', based on the adaptive weighted sum of the results of individual classifiers. The proposed contextaware Re-ID framework demonstrate significant performance improvement both in terms of speed (up to 4.5 times faster) and accuracy (up to 17% rank-1 Re-ID rate) compared to the Context-unaware systems. From the study, we show that gait features are better for lateral views and anthropometric features are better for frontal views, confirming the results of previous studies. Athira Nambiar, Alexandre Bernardino, Jacinto C. Nascimento, Ana Fred |
FG | 2 |
| 2017 | Low-cost 3-axis soft tactile sensors for the human-friendly robot VizzyabstractIn this paper we present a low-cost and easy to fabricate 3-axis tactile sensor based on magnetic technology. The sensor consists in a small magnet immersed in a silicone body with an Hall-effect sensor placed below to detect changes in the magnetic field caused by displacements of the magnet, generated by an external force applied to the silicone body. The use of a 3-axis Hall-effect sensor allows to detect the three components of the force vector, and the proposed design assures high sensitivity, low hysteresis and good repeatability of the measurement: notably, the minimum sensed force is about 0.007N. All components are cheap and easy to retrieve and to assemble; the fabrication process is described in detail and it can be easily replicated by other researchers. Sensors with different geometries have been fabricated, calibrated and successfully integrated in the hand of the human-friendly robot Vizzy. In addition to the sensor characterization and validation, real world experiments of object manipulation are reported, showing proper detection of both normal and shear forces. Tiago Paulino, Pedro Ribeiro 0006, Susana Cardoso, Alexander Schmitz, José Santos-Victor, Alexandre Bernardino, Lorenzo Jamone |
ICRA | 7 |
| 2017 | Towards markerless visual servoing of grasping tasks for humanoid robotsabstractVision-based grasping for humanoid robots is a challenging problem due to a multitude of factors. First, humanoid robots use an “eye-to-hand” kinematics configuration that, on the contrary to the more common “eye-in-hand” configuration, demands a precise estimate of the position of the robot's hand. Second, humanoid robots have a long kinematic chain from the eyes to the hands, prone to accumulate the calibration errors of the kinematics model, which offsets the measured hand-to-object relative pose from the real one. In this paper, we propose a method able to solve these two issues jointly. A robust pose estimation of the robot's hand is achieved via a 3D model-based stereo-vision algorithm, using an edge-based distance transform metric and synthetically generated images of a robot's arm-hand internal computer-graphics model (kinematics and appearance). Then, a particle-based optimization method adapts on-line the robot's internal model to match the real and the synthetically generated images, effectively compensating the kinematics calibration errors. We evaluate the proposed approach using a position-based visual-servoing method on the iCub robot, showing the importance of the continuous visual feedback in humanoid grasping tasks. Pedro Vicente, Lorenzo Jamone, Alexandre Bernardino |
ICRA | 3 |
| 2017 | Self-reconfigurable modular robot interface using virtual reality: Arrangement of furniture made out of roombots modulesabstractSelf-reconfigurable modular robots (SRMR) offer high flexibility in task space by adopting different morphologies for different tasks. Using the same simple module, complex and more capable morphologies can be built. However, increasing the number of modules increases the degrees of freedom (DOF) of the system. Thus, controlling the system as a whole becomes harder. Indeed, even a 10 DOFs system is difficult to consider and manipulate. Intuitive and easy to use interfaces are needed, particularly when modular robots need to interact with humans. In this study we present an interface to assemble desired structures and placement of such structures, with a focus on the assembly process. Roombots modules, a particular SRMR design, are used for the demonstration of the proposed interface. Two non-conventional input/output devices - a head mounted display and hand tracking system - are added to the system to enhance the user experience. Finally, a user study was conducted to evaluate the interface. The results show that most users enjoyed their experience. However, they were not necessarily convinced by the gesture control, most likely for technical reasons. Valentin Z. Nigolian, Mehmet Mutlu, Simon Hauser, Alexandre Bernardino, Auke Jan Ijspeert |
RO-MAN | 4 |
| 2017 | Improving the performance of pedestrian detectors using convolutional learning
David Ribeiro, Jacinto C. Nascimento, Alexandre Bernardino, Gustavo Carneiro 0001 |
Pattern Recognit. | 3 |
| 2016 | Aerial Detection in Maritime Scenarios Using Convolutional Neural Networks
Gonçalo Cruz, Alexandre Bernardino |
ACIVS | 2 |
| 2016 | Person Re-identification in Frontal Gait Sequences via Histogram of Optic Flow Energy Image
Athira Nambiar, Jacinto C. Nascimento, Alexandre Bernardino, José Santos-Victor |
ACIVS | 3 |
| 2016 | From human instructions to robot actions: Formulation of goals, affordances and probabilistic planningabstractThis paper addresses the problem of having a robot executing motor tasks requested by a human through spoken language. Verbal instructions do not typically have a one-to-one mapping to robot actions, due to various reasons: economy of spoken language, e.g., one short instruction might indeed correspond to a complex sequence of robot actions, and details about action execution might be omitted; grounding, e.g., some actions might need to be added or adapted due to environmental contingencies; embodiment, e.g., a robot might have different means than the human ones to obtain the goals that the instruction refers to. We propose a general cognitive architecture to deal with these issues, based on three steps: i) language-based semantic reasoning on the instruction (high-level), ii) formulation of goals in robot symbols and probabilistic planning to achieve them (mid-level), iii) action execution (low-level). The description of the mid-level is the main focus of this paper. The robot plans are adapted to the current scenario, perceived in real-time and continuously updated, taking in consideration the robot capabilities, modeled through the concept of affordances: this allows for flexibility and creativity in the task execution. We showcase the performance of the proposed architecture with real world experiments using the iCub humanoid robot, also in the presence of unexpected events and action failures. Alexandre Antunes, Lorenzo Jamone, Giovanni Saponaro, Alexandre Bernardino, Rodrigo M. M. Ventura |
ICRA | 4 |
| 2016 | Unscented Bayesian optimization for safe robot graspingabstractSafe and robust grasping of unknown objects is a major challenge in robotics, which has no general solution yet. A promising approach relies on haptic exploration, where active optimization strategies can be employed to reduce the number of exploration trials. One critical problem is that certain optimal grasps discoverd by the optimization procedure may be very sensitive to small deviations of the parameters from their nominal values: we call these unsafe grasps because small errors during motor execution may turn optimal grasps into bad grasps. To reduce the risk of grasp failure, safe grasps should be favoured. Therefore, we propose a new algorithm, unscented Bayesian optimization, that performs efficient optimization while considering uncertainty in the input space, leading to the discovery of safe optima. The results highlight how our method outperforms the classical Bayesian optimization both in synthetic problems and in realistic robot grasp simulations, finding robust and safe grasps after a few exploration trials. José Nogueira, Ruben Martinez-Cantin, Alexandre Bernardino, Lorenzo Jamone |
IROS | 3 |
| 2016 | Natural user interface for lighting control: Case study on desktop lighting using modular robotsabstractRoombots (RB) are self-reconfigurable modular robots designed to explore physical structure change by robotic reconfiguration and adaptive locomotion on structured grid environments or unstructured environments. The primary goal of RB is to create adaptive furniture. In this study, we propose a novel and user-friendly interface to control position and intensity of a mobile desk light using RB modules. In the proposed method, the user interacts with the RB with only hand/arm gestures. The user's arm is tracked with a single Kinect having bird's eye view. We demonstrate the effectiveness of the proposed interface in real hardware setup and discuss contributions of it. Mehmet Mutlu, Stéphane Bonardi, Massimo Vespignani, Simon Hauser, Alexandre Bernardino, Auke Jan Ijspeert |
RO-MAN | 5 |
| 2016 | A Window-Based Classifier for Automatic Video-Based ReidentificationabstractThe vast quantity of visual data generated by the rapid expansion of large scale distributed multicamera networks, makes automated person detection and reidentification (RE-ID) essential components of modern surveillance systems. However, the integration of automated person detection and RE-ID algorithms is not without problems, and the errors arising in this integration must be measured (e.g., detection failures that may hamper the RE-ID performance). In this paper, we present a window-based classifier based on a recently proposed architecture for the integration of pedestrian detectors and RE-ID algorithms, that takes the output of any bounding-box RE-ID classifier and exploits the temporal continuity of persons in video streams. We evaluate our contributions on a recently proposed dataset featuring 13 high-definition cameras and over 80 people, acquired during 30 min at rush hour in an office space scenario. We expect our contributions to drive research in integrated pedestrian detection and RE-ID systems, bringing them closer to practical applications. Dario Figueira, Matteo Taiana, Jacinto C. Nascimento, Alexandre Bernardino |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2015 | A novel approach to dynamic movement imitation based on quadratic programmingabstractThis paper proposes a novel approach to generate trajectories that generalize given demonstrations according to optimality criteria. By formulating the problem as a quadratic program we can efficiently incorporate constraints to adapt to new desired motion requirements while achieving the main goal of matching the acceleration profile of the demonstration. This makes our method particularly suited for the imitation and generalization of trajectories such as hitting movements, where it is crucial to maintain the dynamic traits of the demonstration while respecting strict requirements for the goals position, velocity and time. Our method draws inspiration from the Dynamical Movement Primitives (DMPs) framework, preserving its desirable properties of flexibility and rejection of disturbances during execution. Moreover, it offers an higher degree of control on the generated solution, allowing for example i) to limit the instantaneous positions, velocities and accelerations during the whole trajectory, and ii) to add intermediate way points that were not present in the demonstration. With current state-of-the-art solvers of quadratic programs, a problem with hundreds of parameters can be solved in tens of milliseconds in a standard computer, allowing practical applications. Our methodology results in trajectories with a very good approximation of the shape traits of the demonstration, with additional flexibility in specifying constraints of the generated trajectory. Carlos Cardoso, Lorenzo Jamone, Alexandre Bernardino |
ICRA | 3 |
| 2015 | Image Saliency Applied to Infrared Images for Unmanned Maritime Monitoring
Gonçalo Cruz, Alexandre Bernardino |
ICVS | 2 |
| 2015 | Low-rank forward models: A path to the self-organization of visuo-motor systemsabstractSensorimotor coupling is ubiquitous in living organisms. Sensory and motor systems are utterly useless if left without the presence of the other. One crucial faculty that organisms have developed with tremendous ecological advantages is the ability to discern between the origins of perceptual input as being originated by the environment or the organism itself, provided by resource efficient sensor and motor systems. This ability has been shown to be implemented through a specialized circuit (forward model) receiving a copy of the motor command (corollary discharge). We propose a fast method to derive a resource constrained forward model by framing sensorimotor coupling as a low-rank approximation of an overly detailed forward model. By framing the problem as a factorization approach we can resort to currently available off-the-shelf solvers for matrix factorization. We experimentally show that by solving the problem as a low-rank approximation we obtain more than an order of magnitude speed up relatively to minimizing the objective function with gradient descent methods. The development of resource constrained and ecologically adapted sensorimotor systems is essential for the deployment of low-cost energy efficient autonomous robots for the execution of specific tasks in particular environments. Ângelo Cardoso, Ricardo Ferreira 0002, Ricardo Santos 0003, Alexandre Bernardino |
IROS | 4 |
| 2015 | Uncertainty analysis of the DLT-Lines calibration algorithm for cameras with radial distortion
Ricardo Galego, Agustin Alberto Ortega Jimenez, Ricardo Ferreira 0002, Alexandre Bernardino, Juan Andrade-Cetto, José António Gaspar |
Comput. Vis. Image Underst. | 4 |
| 2015 | Efficient pose estimation of rotationally symmetric objects
Rui Pimentel de Figueiredo, Plinio Moreno, Alexandre Bernardino |
Neurocomputing | 3 |
| 2015 | On the purity of training and testing data for learning: The case of pedestrian detection
Matteo Taiana, Jacinto C. Nascimento, Alexandre Bernardino |
Neurocomputing | 3 |
| 2015 | People and Mobile Robot Classification Through Spatio-Temporal Analysis of Optical FlowabstractThe goal of this work is to distinguish between humans and robots in a mixed human-robot environment. We analyze the spatio-temporal patterns of optical flow-based features along several frames. We consider the Histogram of Optical Flow (HOF) and the Motion Boundary Histogram (MBH) features, which have shown good results on people detection. The spatio-temporal patterns are composed of groups of feature components that have similar values on previous frames. The groups of features are fed into the FuzzyBoost algorithm, which at each round selects the spatio-temporal pattern (i.e. feature set) having the lowest classification error. The search for patterns is guided by grouping feature dimensions, considering three algorithms: (a) similarity of weights from dimensionality reduction matrices, (b) Boost Feature Subset Selection (BFSS) and (c) Sequential Floating Feature Selection (SFSS), which avoid the brute force approach. The similarity weights are computed by the Multiple Metric Learning for large Margin Nearest Neighbor (MMLMNN), a linear dimensionality algorithm that provides a type of Mahalanobis metric Weinberger and Saul, J. MaCh. Learn. Res.10 (2009) 207–244. The experiments show that FuzzyBoost brings good generalization properties, better than the GentleBoost, the Support Vector Machines (SVM) with linear kernels and SVM with Radial Basis Function (RBF) kernels. The classifier was implemented and tested in a real-time, multi-camera dynamic setting. Plinio Moreno, Dario Figueira, Alexandre Bernardino, José Santos-Victor |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2015 | Matrix Completion for Weakly-Supervised Multi-Label Image ClassificationabstractIn the last few years, image classification has become an incredibly active research topic, with widespread applications. Most methods for visual recognition are fully supervised, as they make use of bounding boxes or pixelwise segmentations to locate objects of interest. However, this type of manual labeling is time consuming, error prone and it has been shown that manual segmentations are not necessarily the optimal spatial enclosure for object classifiers. This paper proposes a weakly-supervised system for multi-label image classification. In this setting, training images are annotated with a set of keywords describing their contents, but the visual concepts are not explicitly segmented in the images. We formulate the weakly-supervised image classification as a low-rank matrix completion problem. Compared to previous work, our proposed framework has three advantages: (1) Unlike existing solutions based on multiple-instance learning methods, our model is convex. We propose two alternative algorithms for matrix completion specifically tailored to visual data, and prove their convergence. (2) Unlike existing discriminative methods, our algorithm is robust to labeling errors, background noise and partial occlusions. (3) Our method can potentially be used for semantic segmentation. Experimental validation on several data sets shows that our method outperforms state-of-the-art classification algorithms, while effectively capturing each class appearance. Ricardo Silveira Cabral, Fernando De la Torre, João Paulo Costeira, Alexandre Bernardino |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2015 | Shape Context for soft biometrics in person re-identification and database retrieval
Athira Nambiar, Alexandre Bernardino, Jacinto C. Nascimento |
Pattern Recognit. Lett. | 2 |
| 2014 | An algorithm for the detection of vessels in aerial imagesabstractIn this paper we present a sea vessel detection algorithm in aerial image sequences acquired by an unmanned aerial vehicle. The proposed method is robust to variable background lighting, highlights due to sun reflections, vehicle self motion and scale changes. By relying in simple blob analysis rules, based on both spatial and temporal constraints, the algorithm is capable of real-time operation onboard the vehicle, even with non optimized code. We evaluate our method on three sequences labeled with ground truth vessel position, with more that 2900 frames. Overall we are able to achieve very low false positive rates even in heavy sun reflection conditions. Jorge S. Marques, Alexandre Bernardino, Gonçalo Cruz, Maria Bento |
AVSS | 2 |
| 2014 | Optimal no-intersection multi-label binary localization for time series using totally unimodular linear programmingabstractWe propose a new model for simultaneously localizing different classes in the same media, casting it as an integer optimization problem. Our model subsumes into a single formulation previous single and multi-class localization methods, as well as allows us to exploit optimal relaxations to the linear domain. We apply our model to the problem of multi-label multiple instance learning for tagging video collections. Given weakly labeled training samples, where tags for actions in video and objects in images are known but not their locations, our aim is to train classifiers for both detection and localization of said classes on new data. Experimental results demonstrate our approach obtains similar performances when compared to fully supervised methods. Ricardo Silveira Cabral, João Paulo Costeira, Alexandre Bernardino, Fernando De la Torre |
ICIP | 3 |
| 2014 | Probabilistic stereo egomotion transformabstractIn this paper we propose a novel fully probabilistic solution to the stereo egomotion estimation problem. We extend the notion of probabilistic correspondence to the stereo case which allow us to compute the whole 6D motion information in a probabilistic way. We compare the developed approach against other known state-of-the-art methods for stereo egomotion estimation, and the obtained results compare favorably both for the linear and angular velocities estimation. Hugo Silva 0003, Eduardo P. da Silva, Alexandre Bernardino |
ICRA | 3 |
| 2013 | Semi-supervised multi-feature learning for person re-identificationabstractPerson re-identification is probably the open challenge for low-level video surveillance in the presence of a camera network with non-overlapped fields of view. A large number of direct approaches has emerged in the last five years, often proposing novel visual features specifically designed to highlight the most discriminant aspects of people, which are invariant to pose, scale and illumination. On the other hand, learning-based methods are usually based on simpler features, and are trained on pairs of cameras to discriminate between individuals. In this paper, we present a method that joins these two ideas: given an arbitrary state-of-the-art set of features, no matter their number, dimensionality or descriptor, the proposed multi-class learning approach learns how to fuse them, ensuring that the features agree on the classification result. The approach consists of a semi-supervised multi-feature learning strategy, that requires at least a single image per person as training data. To validate our approach, we present results on different datasets, using several heterogeneous features, that set a new level of performance in the person re-identification problem. Dario Figueira, Loris Bazzani, Hà Quang Minh, Marco Cristani, Alexandre Bernardino, Vittorio Murino |
AVSS | 5 |
| 2013 | Unifying Nuclear Norm and Bilinear Factorization Approaches for Low-Rank Matrix DecompositionabstractLow rank models have been widely used for the representation of shape, appearance or motion in computer vision problems. Traditional approaches to fit low rank models make use of an explicit bilinear factorization. These approaches benefit from fast numerical methods for optimization and easy kernelization. However, they suffer from serious local minima problems depending on the loss function and the amount/type of missing data. Recently, these low-rank models have alternatively been formulated as convex problems using the nuclear norm regularizer, unlike factorization methods, their numerical solvers are slow and it is unclear how to kernelize them or to impose a rank a priori. This paper proposes a unified approach to bilinear factorization and nuclear norm regularization, that inherits the benefits of both. We analyze the conditions under which these approaches are equivalent. Moreover, based on this analysis, we propose a new optimization algorithm and a "rank continuation'' strategy that outperform state-of-the-art approaches for Robust PCA, Structure from Motion and Photometric Stereo with outliers and missing data. Ricardo Silveira Cabral, Fernando De la Torre, João Paulo Costeira, Alexandre Bernardino |
ICCV | 4 |
| 2013 | Multi-object detection and pose estimation in 3D point clouds: A fast grid-based Bayesian FilterabstractWe address the problem of object detection and pose estimation using 3D dense data in a multiple object library scenario. State-of-the-art object detection and pose estimation methods are able cope with background clutter and occlusion with acceptable noise levels in the single object scenario. However, with multiple object libraries, even moderate amount of noise lead to frequent object identity switches and serious pose estimation errors. To attenuate these effects, we propose a joint object-id and pose filtering approach using grid-based Recursive Bayesian Filters (RBF). The grid method considers as state variables the object label and its pose, and models the dynamics of the filter with two “inertia” parameters: one for the object label and the other for the object pose. Sensor noise characteristics are taken into account with an observation noise parameter. To allow real-time functionality we propose a selective update approach that dynamically reduces the set of hypotheses evaluated at run time. We present results in realistic scenarios and compare our approach with state-of-the-art approaches in a three object problem, with significant performance improvements. Rui Pimentel de Figueiredo, Plinio Moreno, Alexandre Bernardino, José Santos-Victor |
ICRA | 3 |
| 2012 | Online calibration of a humanoid robot head from relative encoders, IMU readings and visual dataabstractHumanoid robots are complex sensorimotor systems where the existence of internal models are of utmost importance both for control purposes and for predicting the changes in the world arising from the system's own actions. This so-called expected perception relies on the existence of accurate internal models of the robot's sensorimotor chains. Nuno Moutinho, Martim Brandão, Ricardo Ferreira 0002, José António Gaspar, Alexandre Bernardino, Atsuo Takanishi, José Santos-Victor |
IROS | 5 |
| 2012 | Modeling and planning high-level in-hand manipulation actions from human knowledge and active learning from demonstrationabstractWe propose a method to plan in-hand manipulation actions with a robotic anthropomorphic hand. We consider in-hand manipulation actions as sequences between canonical grasp types identified in the humans. Our work concerns the generation of this sequence, which should be autonomous and fast enough to be performed on-line. We use a Markov Decision Process (MDP) governing the transitions between grasp types, depending on the object and on the goal grasp. The policy is learnt directly from human behavior, after an initialization using an empirical estimation of the state action probabilities of the MDP. Then, the policy is finely learnt from samples of human in-hand manipulation records. These samples are chosen using active learning, in order to maximize the useful information of every record, and speed up the learning process. For planning, the policy gives the sequence with highest probability of success. We show a serie of realistic human-like grasp transition sequences derived from the proposed method. Urbain Prieur, Véronique Perdereau, Alexandre Bernardino |
IROS | 3 |
| 2012 | Fast estimation of Gaussian mixture models for image segmentation
Nicola Greggio, Alexandre Bernardino, Cecilia Laschi, Paolo Dario, José Santos-Victor |
Mach. Vis. Appl. | 2 |
| 2012 | Language Bootstrapping: Learning Word Meanings From Perception-Action AssociationabstractWe address the problem of bootstrapping language acquisition for an artificial system similarly to what is observed in experiments with human infants. Our method works by associating meanings to words in manipulation tasks, as a robot interacts with objects and listens to verbal descriptions of the interactions. The model is based on an affordance network, i.e., a mapping between robot actions, robot perceptions, and the perceived effects of these actions upon objects. We extend the affordance model to incorporate spoken words, which allows us to ground the verbal symbols to the execution of actions and the perception of the environment. The model takes verbal descriptions of a task as the input and uses temporal co-occurrence to create links between speech utterances and the involved objects, actions, and effects. We show that the robot is able form useful word-to-meaning associations, even without considering grammatical structure in the learning process and in the presence of recognition errors. These word-to-meaning associations are embedded in the robot's own understanding of its actions. Thus, they can be directly used to instruct the robot to perform tasks and also allow to incorporate context in the speech recognition task. We believe that the encouraging results with our approach may afford robots with a capacity to acquire language descriptors in their operation's environment as well as to shed some light as to how this challenging process develops with human infants. Giampiero Salvi, Luis Montesano, Alexandre Bernardino, José Santos-Victor |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2011 | Generation of meaningful robot expressions with active learningabstractWe propose a mechanism to communicate emotions to humans by using head, torso and arm movements of a humanoid robot, without exploiting its facial features. To this end, we build a library of pre-programmed robot movements and we ask people to attribute emotional scores to these initial movements. The answers are then used to fine-tune motion parameters with an active learning approach. Giovanni Saponaro, Alexandre Bernardino |
HRI | 2 |
| 2011 | Real-time Ellipse Fitting, 3D Spherical Object Localization, and Tracking for the iCub Simulator
Nicola Greggio, Alexandre Bernardino, Cecilia Laschi, Paolo Dario, José Santos-Victor |
ICINCO (2) | 2 |
| 2011 | Monocular Vs Binocular 3D Real-time Ball Tracking from 2D Ellipses
Nicola Greggio, José António Gaspar, Alexandre Bernardino, José Santos-Victor |
ICINCO (2) | 3 |
| 2011 | Fast incremental method for matrix completion: An application to trajectory correctionabstractWe address the problem of incrementally recovering a matrix of tracked image points, based on partial observations of their trajectories. Besides partial observability, we assume the existence of gross, but sparse, noise on the known entries. This problem has obvious applications in real-time tracking and structure from motion, where observations are plagued by self-occlusion and outliers. Recently, work in the optimization community has spun optimal methods for matrix completion when this matrix is known to be low rank by minimizing the nuclear norm, the sum of its singular values. Despite exhibiting several optimality properties, no available algorithms perform this minimization incrementally. In this paper, we build upon the Nuclear Norm Robust PCA method and SPectrally Optimal Completion to propose a fast and incremental algorithm which is able to cope with outliers. We present experiments showing the competitive speed of our method while maintaining performance comparable to the state-of-the-art. Ricardo Silveira Cabral, João Paulo Costeira, Fernando De la Torre, Alexandre Bernardino |
ICIP | 4 |
| 2011 | An expected perception architecture using visual 3D reconstruction for a humanoid robotabstractThe maintenance of a stable and coherent representation of the surrounding environment is an essential capability in cognitive robotic systems. Most systems employ some form of 3D perception to create internal representations of space (maps) to support tasks such as navigation, manipulation and interaction. The creation and update of such representations may represent a significant effort in the overall computation performed by the robot. In this paper we propose an architecture based on the concept of Expected Perception that allows lightweight map updates whenever the course of action happens according to the robot's expectations. It is only when the robot's predictions and the real world outcomes differ, that corrections must be done at its full extent. We performed experiments and show results in a real robotic platform with stereo (3D) perception where map corrections are proposed by simple image level (2D) comparisons. Nuno Moutinho, Nino Cauli, Egidio Falotico, Ricardo Ferreira 0002, José António Gaspar, Alexandre Bernardino, José Santos-Victor, Paolo Dario, Cecilia Laschi |
IROS | 6 |
| 2011 | Matrix Completion for Multi-label Image ClassificationabstractRecently, image categorization has been an active research topic due to the urgent need to retrieve and browse digital images via semantic keywords. This paper formulates image categorization as a multi-label classification problem using recent advances in matrix completion. Under this setting, classification of testing data is posed as a problem of completing unknown label entries on a data matrix that concatenates training and testing features with training labels. We propose two convex algorithms for matrix completion based on a Rank Minimization criterion specifically tailored to visual data, and prove its convergence properties. A major advantage of our approach w.r.t. standard discriminative classification methods for image categorization is its robustness to outliers, background noise and partial occlusions both in the feature and label space. Experimental validation on several datasets shows how our method outperforms state-of-the-art algorithms, while effectively capturing semantic concepts of classes. Ricardo Silveira Cabral, Fernando De la Torre, João Paulo Costeira, Alexandre Bernardino |
NIPS | 4 |
| 2010 | A Practical Method for Self-adapting Gaussian Expectation Maximization
Nicola Greggio, Alexandre Bernardino, José Santos-Victor |
ICINCO (1) | 2 |
| 2010 | Unsupervised Greedy Learning of Finite Mixture ModelsabstractThis work deals with a new technique for the estimation of the parameters and number of components in a finite mixture model. The learning procedure is performed by means of a expectation maximization (EM) methodology. The key feature of our approach is related to a top-down hierarchical search for the number of components, together with the integration of the model selection criterion within a modified EM procedure, used for the learning the mixture parameters. We start with a single component covering the whole data set. Then new components are added and optimized to best cover the data. The process is recursive and builds a binary tree like structure that effectively explores the search space. We show that our approach is faster that state-of-the- art alternatives, is insensitive to initialization, and has better data fits in average. We elucidate this through a series of experiments, both with synthetic and real data. Nicola Greggio, Alexandre Bernardino, Cecilia Laschi, Paolo Dario, José Santos-Victor |
ICTAI (2) | 2 |
| 2010 | An Algorithm for the Least Square-Fitting of EllipsesabstractIn this paper we propose a new algorithm for the least square fitting of ellipses from scattered data. Originally based on the one proposed by Fitzgibbon et Al in 1999, our procedure is able to overcome the numerical instability of that algorithm. We test our approach versus the latter and another approach with different ellipses. Then, we present and discuss our results. Nicola Greggio, Alexandre Bernardino, Cecilia Laschi, Paolo Dario, José Santos-Victor |
ICTAI (2) | 2 |
| 2010 | Gaussian mixture models for affordance learning using Bayesian NetworksabstractAffordances are fundamental descriptors of relationships between actions, objects and effects. They provide the means whereby a robot can predict effects, recognize actions, select objects and plan its behavior according to desired goals. This paper approaches the problem of an embodied agent exploring the world and learning these affordances autonomously from its sensory experiences. Models exist for learning the structure and the parameters of a Bayesian Network encoding this knowledge. Although Bayesian Networks are capable of dealing with uncertainty and redundancy, previous work considered complete observability of the discrete sensory data, which may lead to hard errors in the presence of noise. In this paper we consider a probabilistic representation of the sensors by Gaussian Mixture Models (GMMs) and explicitly taking into account the probability distribution contained in each discrete affordance concept, which can lead to a more correct learning. Pedro Osório, Alexandre Bernardino, Ruben Martinez-Cantin, José Santos-Victor |
IROS | 2 |
| 2010 | Sensor-based self-calibration of the iCub's headabstractIn this paper we propose techniques for the calibration of the iCub's stereo head using vision and inertial measurements. Given that wear and tear can change the geometrical relationship between the different elements in the kinematic chain, new calibrations must be performed periodically. We propose methods that allow automatic calibration without the need for using external sensors or specially designed calibration objects. The methods can be applied at any time during the operation of the system, thus being an alternative for systems whose calibrations are imprecise or that require frequent recalibration. Results are shown both in simulations and on the iCub's stereo head. José Fragoso Santos, Alexandre Bernardino, José Santos-Victor |
IROS | 2 |
| 2010 | Self-adaptive Gaussian mixture models for real-time video segmentation and background subtractionabstractThe usage of Gaussian mixture models for video segmentation has been widely adopted. However, the main difficulty arises in choosing the best model complexity. High complex models can describe the scene accurately, but they come with a high computational requirements, too. Low complex models promote segmentation speed, with the drawback of a less exhaustive description. In this paper we propose an algorithm that first learns a description mixture for the first video frames, and then it uses these results as a starting point for the analysis of the further frames. Then, we apply it to a video sequence and show its effectiveness for real-time tracking multiple moving objects. Moreover, we integrated this procedure into a foreground/background subtraction statistical framework. We compare our procedure against the state-of-the-art alternatives, and we show both its initialization efficacy and its improved segmentation performance. Nicola Greggio, Alexandre Bernardino, Cecilia Laschi, Paolo Dario, José Santos-Victor |
ISDA | 2 |
| 2010 | The iCub humanoid robot: An open-systems platform for research in cognitive development
Giorgio Metta, Lorenzo Natale, Francesco Nori, Giulio Sandini, David Vernon, Luciano Fadiga, Claes von Hofsten, Kerstin Rosander, Manuel Lopes 0001, José Santos-Victor, Alexandre Bernardino, Luis Montesano |
Neural Networks | 11 |
| 2009 | Affordance based word-to-meaning associationabstractThis paper presents a method to associate meanings to words in manipulation tasks. We base our model on an affordance network, i.e., a mapping between robot actions, robot perceptions and the perceived effects of these actions upon objects. We extend the affordance model to incorporate words. Using verbal descriptions of a task, the model uses temporal co-occurrence to create links between speech utterances and the involved objects, actions and effects. We show that the robot is able form useful word-to-meaning associations, even without considering grammatical structure in the learning process and in the presence of recognition errors. These word-to-meaning associations are embedded in the robot's own understanding of its actions. Thus they can be directly used to instruct the robot to perform tasks and also allow to incorporate context in the speech recognition task. Verica Krunic, Giampiero Salvi, Alexandre Bernardino, Luis Montesano, José Santos-Victor |
ICRA | 3 |
| 2009 | ISROBOTNET: A testbed for sensor and robot network systemsabstractThis paper introduces a testbed for sensor and robot network systems, currently composed of 10 cameras and 5 mobile wheeled robots equipped with several sensors for self-localization, obstacle avoidance and vision cameras, and wireless communications. The testbed includes a service-oriented middleware to enable fast prototyping and implementation of algorithms previously tested in simulation, as well as to simplify integration of subsystems developed by different partners. We survey an integrated approach to human-robot interaction that has been developed supported by the testbed under an European research project. The application integrates innovative methods and algorithms for people tracking and waving detection, cooperative perception among static and mobile cameras to improve people tracking accuracy, as well as decision-theoretical approaches to sensor selection and task allocation within the sensor network. Marco Barbosa, Alexandre Bernardino, Dario Figueira, José António Gaspar, Nelson Gonçalves, Pedro U. Lima, Plinio Moreno, Abdolkarim Pahliani, José Santos-Victor, Matthijs T. J. Spaan, João Sequeira 0001 |
IROS | 2 |
| 2009 | Calibrating an outdoor distributed camera network using Laser Range Finder dataabstractOutdoor camera networks are becoming ubiquitous in critical urban areas of large cities around the world. Although current applications of camera networks are mostly limited to video surveillance, recent research projects are exploiting advances on outdoor robotics technology to develop systems that put together networks of cameras and mobile robots in people assisting tasks. Such systems require the creation of robot navigation systems in urban areas with a precise calibration of the distributed camera network. Despite camera calibration has been an extensively studied topic, the calibration (intrinsic and extrinsic) of large outdoor camera networks with no overlapping view fields, and likely to suffer frequent recalibration, poses novel challenges in the development of practical methods for user-assisted calibration that minimize intervention times and maximize precision. In this paper we propose the utilization of Laser Range Finder (LRF) data covering the area of the camera network to support the calibration process and develop a semi-automated methodology allowing quick and precise calibration of large camera networks. The proposed methods have been tested in a real urban environment and have been applied to create direct mappings (homographies) between image coordinates and world points in the ground plane (walking areas) to support person and robot detection and localization algorithms. Agustin Alberto Ortega Jimenez, Ernesto Homar Teniente Avilés, Alexandre Bernardino, José António Gaspar, Juan Andrade-Cetto |
IROS | 4 |
| 2009 | Improving the SIFT descriptor with smooth derivative filters
Plinio Moreno, Alexandre Bernardino, José Santos-Victor |
Pattern Recognit. Lett. | 2 |
| 2008 | Sample-Based 3D Tracking of Colored Objects : A Flexible ArchitectureabstractThis paper presents a method for 3D model-based tracking of colored objects using a sampling methodology. The problem is formulated in a Monte Carlo filtering approach, whereby the state of an object is represented by a set of hypotheses. The main originality of this work is an observation model consisting in the comparison of the color information in some sampling points around the target’s hypothetical edges. On the contrary to existing approaches the method does not need to explicitly compute edges in the video stream, thus dealing well with optical or motion blur. The method does not require the projection of the full 3D object on the image, but just of some selected points around the target’s boundaries. This allows a flexible and modular architecture illustrated by experiments performed with different objects (balls and boxes), camera models (perspective, catadioptric, dioptric) and tracking methodologies (particle and Kalman filtering). 1 Matteo Taiana, Jacinto C. Nascimento, José António Gaspar, Alexandre Bernardino |
BMVC | 4 |
| 2008 | Multimodal saliency-based bottom-up attention a framework for the humanoid robot iCubabstractThis work presents a multimodal bottom-up attention system for the humanoid robot iCub where the robot's decisions to move eyes and neck are based on visual and acoustic saliency maps. We introduce a modular and distributed software architecture which is capable of fusing visual and acoustic saliency maps into one egocentric frame of reference. This system endows the iCub with an emergent exploratory behavior reacting to combined visual and auditory saliency. The developed software modules provide a flexible foundation for the open iCub platform and for further experiments and developments, including higher levels of attention and representation of the peripersonal space. Jonas Ruesch, Manuel Lopes 0001, Alexandre Bernardino, Jonas Hörnstein, José Santos-Victor, Rolf Pfeifer |
ICRA | 3 |
| 2008 | Learning Object Affordances: From Sensory-Motor Coordination to ImitationabstractAffordances encode relationships between actions, objects, and effects. They play an important role on basic cognitive capabilities such as prediction and planning. We address the problem of learning affordances through the interaction of a robot with the environment, a key step to understand the world properties and develop social skills. We present a general model for learning object affordances using Bayesian networks integrated within a general developmental architecture for social robots. Since learning is based on a probabilistic model, the approach is able to deal with uncertainty, redundancy, and irrelevant information. We demonstrate successful learning in the real world by having an humanoid robot interacting with objects. We illustrate the benefits of the acquired knowledge in imitation games. Luis Montesano, Manuel Lopes 0001, Alexandre Bernardino, José Santos-Victor |
IEEE Trans. Robotics | 3 |
| 2007 | Modeling affordances using Bayesian networksabstractAffordances represent the behavior of objects in terms of the robot's motor and perceptual skills. This type of knowledge plays a crucial role in developmental robotic systems, since it is at the core of many higher level skills such as imitation. In this paper, we propose a general affordance model based on Bayesian networks linking actions, object features and action effects. The network is learnt by the robot through interaction with the surrounding objects. The resulting probabilistic model is able to deal with uncertainty, redundancy and irrelevant information. We evaluate the approach using a real humanoid robot that interacts with objects. Luis Montesano, Manuel Lopes 0001, Alexandre Bernardino, José Santos-Victor |
IROS | 3 |
| 2007 | On the use of perspective catadioptric sensors for 3D model-based tracking with particle filtersabstractWe present a model-based 3D tracking system, using wide angle perspective catadioptric sensors. These sensors acquire 360deg views of the environment and the projection from 3D world points to the image plane is approximated by a perspective model. This is a major advantage in structured environments because straight lines on specific surfaces are not deformed by the sensor, allowing the application of standard computer vision algorithms. Objects off the surface are distorted according to a complex projection model, but can be approximated by a simple wide angle perspective mapping. This is exploited here to develop a robust tracking system for autonomous robots using a 3D shape and color-based object model. The use of particle filters allows tracking to be done with 3D realistic motion models and tackling object occlusion, overlap and ambiguities. We show that the use of the perspective model is advantageous over more standard catadioptric projection models, since it renders a very good approximation to the true model, being simpler and more efficient to use, in particular with 3D particle filtering methods. Matteo Taiana, José António Gaspar, Jacinto C. Nascimento, Alexandre Bernardino, Pedro U. Lima |
IROS | 4 |
| 2007 | 3D Tracking by Catadioptric Vision Based on Particle Filters
Matteo Taiana, José António Gaspar, Jacinto C. Nascimento, Alexandre Bernardino, Pedro U. Lima |
RoboCup | 4 |
| 2006 | Design of the Robot-cub (iCub) HeadabstractThis paper describes the design of a robot head, developed in the framework of the RobotCub project. This project goals consists on the design and construction of a humanoid robotic platform, the iCub, for studying human cognition. The final platform would be approximately 90 cm tall, with 23 kg and with a total number of 53 degrees of freedom. For its size, the iCub is the most complete humanoid robot currently being designed, in terms of kinematic complexity. The eyes can also move, as opposed to similarly sized humanoid platforms. Specifications are made based on biological anatomical and behavioral data, as well as tasks constraints. Different concepts for the neck design (flexible, parallel and serial solutions) are analyzed and compared with respect to the specifications. The eye structure and the proprioceptive sensors are presented, together with some discussion of preliminary work on the face design Ricardo Beira, Manuel Lopes 0001, Miguel Praça, José Santos-Victor, Alexandre Bernardino, Giorgio Metta, Francesco Becchi, Roque J. Saltarén |
ICRA | 5 |
| 2006 | Fast IIR Isotropic 2-D Complex Gabor Filters With Boundary InitializationabstractGabor filters are widely applied in image analysis and computer vision applications. This paper describes a fast algorithm for isotropic complex Gabor filtering that outperforms existing implementations. The main computational improvement arises from the decomposition of Gabor filtering into more efficient Gaussian filtering and sinusoidal modulations. Appropriate filter initial conditions are derived to avoid boundary transients, without requiring explicit image border extension. Our proposal reduces up to 39% the number of required operations with respect to state-of-the-art approaches. A full C++ implementation of the method is publicly available. Alexandre Bernardino, José Santos-Victor |
IEEE Trans. Image Process. | 1 |
| 2006 | Detection and classification of highway lanes using vehicle motion trajectoriesabstractIntelligent vision-based traffic surveillance systems are assuming an increasingly important role in highway monitoring and road management schemes. This paper describes a low-level object tracking system that produces accurate vehicle motion trajectories that can be further analyzed to detect lane centers and classify lane types. Accompanying techniques for indexing and retrieval of anomalous trajectories are also derived. The predictive trajectory merge-and-split algorithm is used to detect partial or complete occlusions during object motion and incorporates a Kalman filter that is used to perform vehicle tracking. The resulting motion trajectories are modeled using variable low-degree polynomials. A K-means clustering technique on the coefficient space can be used to obtain approximate lane centers. Estimation bias due to vehicle lane changes can be removed using robust estimation techniques based on Random Sample Consensus (RANSAC). Through the use of nonmetric distance functions and a simple directional indicator, highway lanes can be classified into one of the following categories: entry, exit, primary, or secondary. Experimental results are presented to show the real-time application of this approach to multiple views obtained by an uncalibrated pan-tilt-zoom traffic camera monitoring the junction of two busy intersecting highways. José Melo, Andrew Naftel, Alexandre Bernardino, José Santos-Victor |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2001 | Vision-based Navigation, Environmental Representations and Imaging Geometries
José Santos-Victor, Alexandre Bernardino |
ISRR | 2 |
| 2000 | Vision based station keeping and docking for an aerial blimpabstractThis paper describes a method for station keeping and docking of a lighter-than-air vehicle based on visual input. Due to the motion disturbances in the environment (currents), these tasks are important to keep the vehicle stabilized relative to an external reference frame. The main difficulties to achieve station keeping and docking are related to the nonholonomic constraints of the blimp moving in 3D, having a limited number of controllable degrees of freedom. The relative position of the vehicle with respect to a docking station is tracked using vision. A planar surface is chosen as a reference plane which allows visual tracking of an environmental region, based on planar projective transformations. An image-based control law is proposed together with a dynamic model for the vehicle. Experiments and results are described and discussed. Sjoerd van der Zwaan, Alexandre Bernardino, José Santos-Victor |
IROS | 2 |
| 1999 | Binocular tracking: integrating perception and controlabstractPresents an active binocular tracking system using log-polar images with contributions in both the perceptual and control aspects. The control part is based on the visual servoing framework, including kinematics and dynamics. We introduce a fixation constraint that simplifies the tracking problem by decoupling the visual kinematics and allowing us to express system dynamics in image coordinates. Simple dynamic controllers are designed for each degree of freedom directly from image features. In the perceptual part, we use a space variant sensor that emphasizes the center of the visual field (log-polar geometry). We present a disparity estimation algorithm for log-polar images and provide a theoretical analysis to illustrate the advantages of using space variant images. The overall system is implemented in the Medusa binocular head without any specific processing hardware. The use of log-polar images allows real-time performance (50 Hz). Tracking experiments are presented to illustrate system performance with different control strategies and objects of different shapes and motions. Alexandre Bernardino, José Santos-Victor |
IEEE Trans. Robotics Autom. | 1 |
| 1996 | Vergence control for robotic heads using log-polar imagesabstractThis paper describes a real-time vergence control mechanism based on, log-polar images, developed for a robot head. The real-time control of active vision systems imposes strong constraints on the computational complexity of the vision algorithms. In this paper, we illustrate that vergence of a stereo head can be achieved at reduced computational cost using log-polar images. These images have higher resolution at the center, where the attention is focused on, and the rest of the visual field is covered at a coarser resolution, still enabling the detection of events at the image periphery. The main advantages of using a non-uniform image sampling mechanism, such as the log-polar images, are related both to perceptual and algorithm complexity issues. We show that, when using correlation measures to control vergence, log-polar images give better results than cartesian images. Additionally, as log-polar images are smaller, the computation time is reduced. Two algorithms for closed loop vergence control, using correlation measures over log-polar images, are proposed. In the test examples described in the paper, we compare the two algorithms and address the problem of designing an adequate log-polar sensor, by introducing a performance analysis to evaluate the various log-polar sensor layouts. Alexandre Bernardino, José Santos-Victor |
IROS | 1 |
| 1994 | Interleaving Real-Time Multi-Agent Planning and Execution: An ApplicationabstractWhen faced with real-world planning problems-such as generating plans for simultaneous execution by several interacting agents in a changing environment-traditional approaches to planning fail to provide direct answers. Moreover, if optimal or near optimal solutions are demanded, the complexity involved increases drastically. However certain well-known AI planning techniques (e.g., hierarchical planning, interleaving planning and execution) provide an adequate framework for developing successful applications. This paper is about such an application-how to efficiently plan tasks related with the simultaneous movements of five grippers and a carousel in a robotic system-ensuring a near optimal performance of the overall system.> César Santos Silva, Alexandre Bernardino, Carlos A. Pinto-Ferreira |
ICTAI | 2 |