Josechu J. Guerrero

dblp:117/8257 · also J. J. Guerrero, Josechu Guerrero, José J. Guerrero, José Jesús Guerrero · DBLP profile ↗
← Back
60ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0001-5209-2267ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 54 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 6 since 2021Systems, architecture and hardware · 17 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Temporal video segmentation with natural language using text-video cross attention and Bayesian order-priors
abstract
Video is a crucial perception component in both robotics and wearable devices, two key technologies to enable innovative assistive applications, such as navigation and procedure execution assistance tools. Video understanding tasks are essential to enable these systems to interpret and execute complex instructions in real-world environments. One such task is step grounding, which involves identifying the temporal boundaries of activities based on natural language descriptions in long, untrimmed videos. This paper introduces Bayesian-VSLNet, a probabilistic formulation of step grounding that predicts a likelihood distribution over segments and refines it through Bayesian inference with temporal-order priors. These priors disambiguate cyclic and repeated actions that frequently appear in procedural tasks, enabling precise step localization in long videos. Our evaluations demonstrate superior performance over existing methods, achieving state-of-the-art results in the Ego4D Goal-Step dataset, winning the Goal Step challenge at the EgoVis 2024 CVPR. Furthermore, experiments on additional benchmarks confirm the generality of our approach beyond Ego4D. In addition, we present qualitative results in a real-world robotics scenario, illustrating the potential of this task to improve human–robot interaction in practical applications. Code is released at https://github.com/cplou99/BayesianVSLNet .
Carlos Plou, Lorenzo Mur-Labadia, Josechu J. Guerrero, Ruben Martinez-Cantin, Ana Cristina Murillo
Comput. Vis. Image Underst.3
2026 Integrating Affordances and Attention Models for Short-Term Object Interaction Anticipation
abstract
Short-Term object-interaction Anticipation (STA) consists in detecting the location of the next-active objects, the noun and verb categories of the interaction, as well as the time to contact from the observation of egocentric video. This ability is fundamental for wearable assistants to understand user's goals and provide timely assistance, or to enable human-robot interaction. In this work, we present a method to improve the performance of STA predictions. Our contributions are two-fold: 1) We propose STAformer and STAformer++, two novel attention-based architectures integrating frame-guided temporal pooling, dual image-video attention, and multiscale feature fusion to support STA predictions from an image-input video pair; 2) We introduce two novel modules to ground STA predictions on human behavior by modeling affordances. First, we integrate an environment affordance model which acts as a persistent memory of interactions that can take place in a given physical scene. We explore how to integrate environment affordances via simple late fusion and with an approach which adaptively learns how to best fuse affordances with end-to-end predictions. Second, we predict interaction hotspots from the observation of hands and object trajectories, increasing confidence in STA predictions localized around the hotspot. Our results show significant improvements on Overall Top-5 mAP, with gain up to $+23\%$+23% on Ego4D and $+31\%$+31% on a novel set of curated EPIC-Kitchens STA labels. We released the https://github.com/lmur98/AFFttention code, annotations, and pre-extracted affordances on Ego4D and EPIC-Kitchens to encourage future research in this area.
Lorenzo Mur-Labadia, Ruben Martinez-Cantin, Josechu J. Guerrero, Giovanni Maria Farinella, Antonino Furnari
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Uncertainty estimation in instance segmentation of affordances via Bayesian visual transformers
Lorenzo Mur-Labadia, Ruben Martinez-Cantin, Josechu J. Guerrero
Pattern Recognit.3
2025 DIV-FF: Dynamic Image-Video Feature Fields For Environment Understanding in Egocentric Videos
abstract
Environment understanding in egocentric videos is an important step for applications like robotics, augmented reality and assistive technologies. These videos are characterized by dynamic interactions and a strong dependence on the wearer’s engagement with the environment. Traditional approaches often focus on isolated clips or fail to integrate rich semantic and geometric information, limiting scene comprehension. We introduce Dynamic Image-Video Feature Fields (DIV-FF), a framework that decomposes the egocentric scene into persistent, dynamic, and actor-based components while integrating both image and video-language features. Our model enables detailed segmentation, captures affordances, understands the surroundings and maintains consistent understanding over time. DIV-FF outperforms state-of-the-art methods, particularly in dynamically evolving scenarios, demonstrating its potential to advance long-term, spatiotemporal scene understanding.
Lorenzo Mur-Labadia, Josechu J. Guerrero, Ruben Martinez-Cantin
CVPR2
2025 O-MaMa: Learning Object Mask Matching Between Egocentric and Exocentric Views
abstract
Understanding the world from multiple perspectives is essential for intelligent systems operating together, where segmenting common objects across different views remains an open problem. We introduce a new approach that re-defines cross-image segmentation by treating it as a mask matching task. Our method consists of: (1) A Mask-Context Encoder that pools dense DINOv2 semantic features to obtain discriminative object-level representations from FastSAM mask candidates, (2) an Ego$\leftrightarrow$Exo Cross-Attention that fuses multi-perspective observations, (3) a Mask Matching contrastive loss that aligns cross-view features in a shared latent space, and (4) a Hard Negative Adjacent Mining strategy to encourage the model to better differentiate between nearby objects. O-MaMa achieves the state of the art in the Ego-Exo4D Correspondences benchmark, obtaining relative gains of +22% and +76% in the Ego2Exo and Exo2Ego IoU against the official challenge baselines, and a +13% and +6% compared with the SOTA with 1% of the training parameters.
Lorenzo Mur-Labadia, Maria Santos-Villafranca, Jesus Bermudez-Cameo, Alejandro Pérez-Yus, Ruben Martinez-Cantin, Josechu J. Guerrero
ICCV6
2025 Panoramic Depth and Semantic Estimation With Frequency and Distortion Aware Convolutions
abstract
ABSTRACT Omnidirectional images reveal advantages when addressing the understanding of the environment due to the 360‐degree contextual information. However, the inherent characteristics of the omnidirectional images add additional problems to obtain an accurate detection and segmentation of objects or a good depth estimation. To overcome these problems, we exploit convolutions in the frequency domain, obtaining a wider receptive field in each convolutional layer, and convolutions in the equirectangular projection, to cope with the image distortion. Both convolutions allow to leverage the whole context information from omnidirectional images. Our experiments show that our proposal has better performance on non‐gravity‐oriented panoramas than state‐of‐the‐art methods and similar performance on oriented panoramas as specific state‐of‐the‐art methods for semantic segmentation and for monocular depth estimation, outperforming the sole other method which provides both tasks.
Bruno Berenguel-Baeta, Jesus Bermudez-Cameo, Josechu J. Guerrero
IET Image Process.3
2024 AFF-ttention! Affordances and Attention Models for Short-Term Object Interaction Anticipation
Lorenzo Mur-Labadia, Ruben Martinez-Cantin, Josechu J. Guerrero, Giovanni Maria Farinella, Antonino Furnari
ECCV (23)3
2023 Convolution kernel adaptation to calibrated fisheye
Bruno Berenguel-Baeta, Maria Santos-Villafranca, Jesus Bermudez-Cameo, Alejandro Pérez-Yus, Josechu J. Guerrero
BMVC5
2023 Multi-label affordance mapping from egocentric vision
abstract
Accurate affordance detection and segmentation with pixel precision is an important piece in many complex systems based on interactions, such as robots and assitive devices. We present a new approach to affordance perception which enables accurate multi-label segmentation. Our approach can be used to automatically extract grounded affordances from first person videos of interactions using a 3D map of the environment providing pixel level precision for the affordance location. We use this method to build the largest and most complete dataset on affordances based on the EPIC-Kitchen dataset, EPIC-Aff, which provides interaction-grounded, multi-label, metric and spatial affordance annotations. Then, we propose a new approach to affordance segmentation based on multi-label detection which enables multiple affordances to co-exists in the same space, for example if they are associated with the same object. We present several strategies of multi-label detection using several segmentation architectures. The experimental results highlight the importance of the multi-label detection. Finally, we show how our metric representation can be exploited for build a map of interaction hotspots in spatial action-centric zones and use that representation to perform a task-oriented navigation.
Lorenzo Mur-Labadia, Josechu J. Guerrero, Ruben Martinez-Cantin
ICCV2
2023 FreDSNet: Joint Monocular Depth and Semantic Segmentation with Fast Fourier Convolutions from Single Panoramas
abstract
In this work we present FreDSNet, a deep learning solution which obtains semantic 3D understanding of indoor environments from single panoramas. Omnidirectional images reveal task-specific advantages when addressing scene understanding problems due to the 360-degree contextual information about the entire environment they provide. However, the inherent characteristics of the omnidirectional images add additional problems to obtain an accurate detection and segmentation of objects or a good depth estimation. To overcome these problems, we exploit convolutions in the frequential domain obtaining a wider receptive field in each convolutional layer. These convolutions allow to leverage the whole context information from omnidirectional images. FreDSNet is the first network that jointly provides monocular depth estimation and semantic segmentation from a single panoramic image exploiting fast Fourier convolutions. Our experiments show that FreDSNet has slight better performance than the sole state-of-the-art method that obtains both semantic segmentation and depth estimation from panoramas. FreDSNet code is publicly available in https://github.com/Sbrunoberenguel/FreDSNet
Bruno Berenguel-Baeta, Jesus Bermudez-Cameo, Josechu J. Guerrero
ICRA3
2023 Bayesian deep learning for affordance segmentation in images
abstract
Affordances are a fundamental concept in robotics since they relate available actions for an agent depending on its sensory-motor capabilities and the environment. We present a novel Bayesian deep network to detect affordances in images, at the same time that we quantify the distribution of the aleatoric and epistemic variance at the spatial level. We adapt the Mask-RCNN architecture to learn a probabilistic representation using Monte Carlo dropout. Our results outperform the state-of-the-art of deterministic networks. We attribute this improvement to a better probabilistic feature space representation on the encoder and the Bayesian variability induced at the mask generation, which adapts better to the object contours. We also introduce the new Probability-based Mask Quality measure that reveals the semantic and spatial differences on a probabilistic instance segmentation model. We modify the existing Probabilistic Detection Quality metric by comparing the binary masks rather than the predicted bounding boxes, achieving a finer-grained evaluation of the probabilistic segmentation. We find aleatoric variance in the contours of the objects due to the camera noise, while epistemic variance appears in visual challenging pixels.
Lorenzo Mur-Labadia, Ruben Martinez-Cantin, Josechu J. Guerrero
ICRA3
2022 Atlanta scaled layouts from non-central panoramas
abstract
In this work we present a novel approach for 3D layout recovery of indoor environments using a non-central acquisition system. From a single non-central panorama, full and scaled 3D lines can be independently recovered by geometry reasoning without additional nor scale assumptions. However, their sensitivity to noise and complex geometric modeling has led these panoramas and required algorithms being little investigated. Our new pipeline aims to extract the boundaries of the structural lines of an indoor environment with a neural network and exploit the properties of non-central projection systems in a new geometrical processing to recover scaled 3D layouts. The results of our experiments show that we improve state-of-the-art methods for layout recovery and line extraction in non-central projection systems. We completely solve the problem both in Manhattan and Atlanta environments, handling occlusions and retrieving the metric scale of the room without extra measurements. As far as the authors’ knowledge goes, our approach is the first work using deep learning on non-central panoramas and recovering scaled layouts from single panoramas.
Bruno Berenguel-Baeta, Jesus Bermudez-Cameo, Josechu J. Guerrero
Pattern Recognit.3
2020 Unsupervised Learning of Category-Specific Symmetric 3D Keypoints from Point Sets
Clara Fernandez-Labrador, Ajad Chhatkuli, Danda Pani Paudel, Josechu J. Guerrero, Cédric Demonceaux, Luc Van Gool
ECCV (25)4
2020 Floor Extraction and Door Detection for Visually Impaired Guidance
abstract
Finding obstacle-free paths in unknown environments is a big navigation issue for visually impaired people and autonomous robots. Previous works focus on obstacle avoidance, however they do not have a general view of the environment they are moving in. New devices based on computer vision systems can help impaired people to overcome the difficulties of navigating in unknown environments in safe conditions. In this work it is proposed a combination of sensors and algorithms that can lead to the building of a navigation system for visually impaired people. Based on traditional systems that use RGB-D cameras for obstacle avoidance, it is included and combined the information of a fish-eye camera, which will give a better understanding of the user's surroundings. The combination gives robustness and reliability to the system as well as a wide field of view that allows to obtain many information from the environment. This combination of sensors is inspired by human vision where the center of the retina (fovea) provides more accurate information than the periphery, where humans have a wider field of view. The proposed system is mounted on a wearable device that provides the obstacle-free zones of the scene, allowing the planning of trajectories for people guidance.
Bruno Berenguel-Baeta, Manuel Guerrero-Viu, A. Nova, Jesus Bermudez-Cameo, Alejandro Pérez-Yus, Josechu J. Guerrero
ICARCV6
2020 What's in my Room? Object Recognition on Indoor Panoramic Images
abstract
In the last few years, there has been a growing interest in taking advantage of the 360°panoramic images potential, while managing the new challenges they imply. While several tasks have been improved thanks to the contextual information these images offer, object recognition in indoor scenes still remains a challenging problem that has not been deeply investigated. This paper provides an object recognition system that performs object detection and semantic segmentation tasks by using a deep learning model adapted to match the nature of equirectangular images. From these results, instance segmentation masks are recovered, refined and transformed into 3D bounding boxes that are placed into the 3D model of the room. Quantitative and qualitative results support that our method outperforms the state of the art by a large margin and show a complete understanding of the main objects in indoor scenes.
Julia Guerrero-Viu, Clara Fernandez-Labrador, Cédric Demonceaux, Josechu J. Guerrero
ICRA4
2019 Scaled layout recovery with wide field of view RGB-D
Alejandro Pérez-Yus, Gonzalo López-Nicolás, Josechu J. Guerrero
Image Vis. Comput.3
2018 Fitting line projections in non-central catadioptric cameras with revolution symmetry
abstract
Line-images in non-central cameras contain much richer information of the original 3D line than line projections in central cameras. The projection surface of a 3D line in most catadioptric non-central cameras is a ruled surface, encapsulating the complete information of the 3D line. The resulting line-image is a curve which contains the 4 degrees of freedom of the 3D line. That means a qualitative advantage with respect to the central case, although extracting this curve is quite difficult. In this paper, we focus on the analytical description of the line-images in non-central catadioptric systems with symmetry of revolution. As a direct application we present a method for automatic line-image extraction for conical and spherical calibrated catadioptric cameras. For designing this method we have analytically solved the metric distance from point to line-image for non-central catadioptric systems. We also propose a distance we call effective baseline measuring the quality of the reconstruction of a 3D line from the minimum number of rays. This measure is used to evaluate the different random attempts of a robust scheme allowing to reduce the number of trials in the process. The proposal is tested and evaluated in simulations and with both synthetic and real images.
Jesus Bermudez-Cameo, Gonzalo López-Nicolás, Josechu J. Guerrero
Comput. Vis. Image Underst.3
2017 Stairs detection with odometry-aided traversal from a wearable RGB-D camera
Alejandro Pérez-Yus, Daniel Gutiérrez-Gómez, Gonzalo López-Nicolás, Josechu J. Guerrero
Comput. Vis. Image Underst.4
2017 Exploiting line metric reconstruction from non-central circular panoramas
Jesus Bermudez-Cameo, Olivier Saurer, Gonzalo López-Nicolás, Josechu J. Guerrero, Marc Pollefeys
Pattern Recognit. Lett.4
2016 Line reconstruction using prior knowledge in single non-central view
Jesus Bermudez-Cameo, Cédric Demonceaux, Gonzalo López-Nicolás, Josechu J. Guerrero
BMVC4
2016 Dense Labeling with User Interaction: an Example for Depth-Of-Field Simulation
Ana B. Cambra, Adolfo Muñoz 0001, Josechu J. Guerrero, Ana Cristina Murillo
BMVC3
2016 Peripheral Expansion of Depth Information via Layout Estimation with Fisheye Camera
Alejandro Pérez-Yus, Gonzalo López-Nicolás, Josechu J. Guerrero
ECCV (8)3
2016 A novel hybrid camera system with depth and fisheye cameras
abstract
We introduce a novel hybrid camera configuration composed by a fisheye camera attached to an RGB-D system. Current RGB-D sensors provide the 3D information and scale of the scene, but they are limited by a small field of view. In contrast, wide field of view cameras capture a larger portion of the scene, but providing highly distorted images that require specific algorithms. By coupling a fisheye camera to an RGB-D system we take advantage of both types of cameras overcoming their drawbacks. The system provides a portion of the fisheye image with depth data and we use this seed information to perform scaled operations in the complete image. We also present a calibration procedure of the system to map depth information to the wide angle image. With this purpose, we propose a depth-fisheye calibration algorithm nurturing from state of the art camera models and methods. Several experiments test the accuracy of the system with real images.
Alejandro Pérez-Yus, Gonzalo López-Nicolás, Josechu J. Guerrero
ICPR3
2016 True scaled 6 DoF egocentric localisation with monocular wearable systems
Daniel Gutiérrez-Gómez, Josechu J. Guerrero
Image Vis. Comput.2
2015 Inverse depth for accurate photometric and geometric error minimisation in RGB-D dense visual odometry
abstract
In this paper we present a dense visual odometry system for RGB-D cameras performing both photometric and geometric error minimisation to estimate the camera motion between frames. Contrary to most works in the literature, we parametrise the geometric error by the inverse depth instead of the depth, which translates into a better fit of the distribution of the geometric error to the used robust cost functions. We also provide a unified evaluation under the same framework of different estimators and ways of computing the scale of the residuals which can be found spread along the related literature. For the comparison of our approach with state-of-the-art approaches we use the popular dataset from the TUM for RGB-D benchmarking. Our approach shows to be competitive with state-of-the-art methods in terms of drift in meters per second, even compared to methods performing loop closure too. When comparing to approaches performing pure odometry like ours, our method outperforms them in the majority of the tested datasets. Additionally we show that our approach is able to work in real time and we provide a qualitative evaluation on our own sequences showing a low drift in the 3D reconstructions.
Daniel Gutiérrez-Gómez, Walterio W. Mayol-Cuevas, Josechu J. Guerrero
ICRA3
2015 What should I landmark? Entropy of normals in depth juts for place recognition in changing environments using RGB-D data
abstract
One open problem in the fields of place recognition and mapping is to be able to recognise a revisited place when its appearance and layout have changed between visits. In this paper, we investigate this problem in the context of RGB-D mapping in indoor environments. We propose to segment the scene in juts (neighbourhood of 3D points with normals that stick out from the surroundings) and look at low-level features, like textureness or entropy of the normals. These could differentiate those zones of the scene that change or move along time from those that are likely to remain static. We also present a method which improves the matching between images of the same place taken at different times by pruning details basing on these features. We evaluate on a number of communal areas and also on some scenes captured 6 months apart. Experiments with our approach, show an increase up to 70% in inlier matching ratio at the cost of pruning only less than 20% of correct matches, without the need of performing geometric verification.
Daniel Gutiérrez-Gómez, Walterio W. Mayol-Cuevas, Josechu J. Guerrero
ICRA3
2015 Automatic Line Extraction in Uncalibrated Omnidirectional Cameras with Revolution Symmetry
Jesus Bermudez-Cameo, Gonzalo López-Nicolás, Josechu J. Guerrero
Int. J. Comput. Vis.3
2014 Minimal Solution for Computing Pairs of Lines in Non-central Cameras
Jesus Bermudez-Cameo, João Pedro Barreto 0001, Gonzalo López-Nicolás, Josechu J. Guerrero
ACCV (1)4
2014 Line-based global descriptor for omnidirectional vision
abstract
Scene understanding is a widely studied problem in computer vision. Many works approach this problem in indoor environments assuming constraints about the scene, such as the typical Manhattan World assumption. The goal of this work is to design and evaluate a global descriptor for indoor panoramic images that encloses information about the 3D structure. This descriptor is based on the detection of representative lines of the scene, which encode the scene structure. Our work focuses on omnidirectional imagery, where observed lines are longer than in conventional images and the whole scene is captured in a single image. Experiments using two public datasets analyze the performance of the descriptor for scene categorization. We also analyze the influence of different parameters and show sample results for a navigation assistance application.
Alejandro Rituerto, Ana Cristina Murillo, Josechu J. Guerrero
ICIP3
2014 Line-Images in Cone Mirror Catadioptric Systems
abstract
The projection surface of a 3D line in a non-central camera is a ruled surface, containing the complete information of the 3D line. The resulting line-image is a curve which contains the 4 degrees of freedom of the 3D line. In this paper we investigate the properties of the line-image in conical catadioptric systems. This curve is a particular quartic that can be described by only six homogeneous parameters. We present the relation between the line-image description and the geometry of the mirror. This result reveals the coupling between the depth of the line and the distance from the camera to the mirror. If this distance is unknown the 3D information of a projected line can be recovered up to scale. Knowing this distance allows obtaining the 3D metric reconstruction. The proposed parametrization also allows to simultaneously reconstruct the 3D line and computing the aperture angle of the mirror from five projected points on the line-image. We analytically solve the metric distance from a point to a line-image and we evaluate the proposal with real images.
Jesus Bermudez-Cameo, Gonzalo López-Nicolás, Josechu J. Guerrero
ICPR3
2014 Exploiting projective geometry for view-invariant monocular human motion analysis in man-made environments
Grégory Rogez, Carlos Orrite-Uruñuela, Josechu J. Guerrero, Philip Torr 0001
Comput. Vis. Image Underst.3
2014 Scale Space for Camera Invariant Features
abstract
In this paper we propose a new approach to compute the scale space of any central projection system, such as catadioptric, fisheye or conventional cameras. Since these systems can be explained using a unified model, the single parameter that defines each type of system is used to automatically compute the corresponding Riemannian metric. This metric, is combined with the partial differential equations framework on manifolds, allows us to compute the Laplace-Beltrami (LB) operator, enabling the computation of the scale space of any central projection system. Scale space is essential for the intrinsic scale selection and neighborhood description in features like SIFT. We perform experiments with synthetic and real images to validate the generalization of our approach to any central projection system. We compare our approach with the best-existing methods showing competitive results in all type of cameras: catadioptric, fisheye, and perspective.
Luis Puig, Josechu J. Guerrero, Kostas Daniilidis
IEEE Trans. Pattern Anal. Mach. Intell.2
2014 Monocular 3-D Gait Tracking in Surveillance Scenes
abstract
Gait recognition can potentially provide a noninvasive and effective biometric authentication from a distance. However, the performance of gait recognition systems will suffer in real surveillance scenarios with multiple interacting individuals and where the camera is usually placed at a significant angle and distance from the floor. We present a methodology for view-invariant monocular 3-D human pose tracking in man-made environments in which we assume that observed people move on a known ground plane. First, we model 3-D body poses and camera viewpoints with a low dimensional manifold and learn a generative model of the silhouette from this manifold to a reduced set of training views. During the online stage, 3-D body poses are tracked using recursive Bayesian sampling conducted jointly over the scene's ground plane and the pose-viewpoint manifold. For each sample, the homography that relates the corresponding training plane to the image points is calculated using the dominant 3-D directions of the scene, the sampled location on the ground plane and the sampled camera view. Each regressed silhouette shape is projected using this homographic transformation and is matched in the image to estimate its likelihood. Our framework is able to track 3-D human walking poses in a 3-D environment exploring only a 4-D state space with success. In our experimental evaluation, we demonstrate the significant improvements of the homographic alignment over a commonly used similarity transformation and provide quantitative pose tracking results for the monocular sequences with a high perspective effect from the CAVIAR dataset.
Grégory Rogez, Jonathan Rihan, Josechu J. Guerrero, Carlos Orrite-Uruñuela
IEEE Trans. Cybern.3
2013 Line extraction in uncalibrated central images with revolution symmetry
abstract
In omnidirectional cameras, straight lines in the scene are projected onto curves called line-images.The shape of these curves is strongly dependent of the particular camera configuration.The great diversity of omnidirectional camera systems makes harder the line-image extraction in a general way.Therefore, it is difficult to design uncalibrated general approaches, and existing methods to extract lines in omnidirectional images require the camera calibration.In this paper, we present a novel method to extract lineimages in uncalibrated images which is valid for radially symmetric central systems.In our proposal, the distortion function is analytically solved for different types of camera systems, dioptric or catadioptric.We present the unified line-image constraints to extract the projection plane of each line and main calibration parameter of the camera from a single line-image.The use of gradient-based information allows computing both from a minimum of two image points.This scheme is used in a line-image extraction algorithm to obtain lines from uncalibrated omnidirectional images without any assumption about the scene.The algorithm is evaluated with synthetic and real images showing good performance.
Jesus Bermudez-Cameo, Gonzalo López-Nicolás, Josechu J. Guerrero
BMVC3
2013 Hybrid homographies and fundamental matrices mixing uncalibrated omnidirectional and conventional cameras
Luis Puig, Peter F. Sturm, Josechu J. Guerrero
Mach. Vis. Appl.3
2013 Localization in Urban Environments Using a Panoramic Gist Descriptor
abstract
Vision-based topological localization and mapping for autonomous robotic systems have received increased research interest in recent years. The need to map larger environments requires models at different levels of abstraction and additional abilities to deal with large amounts of data efficiently. Most successful approaches for appearance-based localization and mapping with large datasets typically represent locations using local image features. We study the feasibility of performing these tasks in urban environments using global descriptors instead and taking advantage of the increasingly common panoramic datasets. This paper describes how to represent a panorama using the global gist descriptor, while maintaining desirable invariance properties for location recognition and loop detection. We propose different gist similarity measures and algorithms for appearance-based localization and an online loop-closure detection method, where the probability of loop closure is determined in a Bayesian filtering framework using the proposed image representation. The extensive experimental validation in this paper shows that their performance in urban environments is comparable with local-feature-based approaches when using wide field-of-view images.
Ana Cristina Murillo, Gautam Singh, Jana Kosecka, Josechu J. Guerrero
IEEE Trans. Robotics4
2012 A Unified Framework for Line Extraction in Dioptric and Catadioptric Cameras
Jesus Bermudez-Cameo, Gonzalo López-Nicolás, Josechu J. Guerrero
ACCV (4)3
2012 Full scaled 3D visual odometry from a single wearable omnidirectional camera
abstract
In the last years monocular SLAM has been widely used to obtain highly accurate maps and trajectory estimations of a moving camera. However, one of the issues of this approach is that, due to the impossibility of the depth being measured in a single image, global scale is not observable and scene and camera motion can only be recovered up to scale. This problem gets aggravated as we deal with larger scenes since it is more likely that scale drift arises between different map portions and their corresponding motion estimates. To compute the absolute scale we need to know some kind of dimension of the scene (e.g., actual size of an element of the scene, velocity of the camera or baseline between two frames) and somehow integrate it in the SLAM estimation. In this paper, we present a method to recover the scale of the scene using an omnidirectional camera mounted on a helmet. The high precision of visual SLAM allows the head vertical oscillation during walking to be perceived in the trajectory estimation. By performing a spectral analysis on the camera vertical displacement, we can measure the step frequency. We relate the step frequency to the speed of the camera by an empirical formula based on biomedical experiments on human walking. This speed measurement is integrated in a particle filter to estimate the current scale factor and the 3D motion estimation with its true scale. We evaluated our approach using image sequences acquired while a person walks. Our experiments show that the proposed approach is able to cope with scale drift.
Daniel Gutiérrez-Gómez, Luis Puig, Josechu J. Guerrero
IROS3
2012 Calibration of omnidirectional cameras in practice: A comparison of methods
Luis Puig, Jesús Bermúdez, Peter F. Sturm, Josechu J. Guerrero
Comput. Vis. Image Underst.4
2011 Scale space for central catadioptric systems: Towards a generic camera feature extractor
abstract
In this paper we propose a new approach to compute the scale space of any omnidirectional image acquired with a central catadioptric system. When these cameras are central they are explained using the sphere camera model, which unifies in a single model, conventional, paracatadioptric and hypercatadioptric systems. Scale space is essential in the detection and matching of interest points, in particular scale invariant points based on Laplacian of Gaussians, like the well known SIFT. We combine the sphere camera model and the partial differential equations framework on manifolds, to compute the Laplace-Beltrami (LB) operator which is a second order differential operator required to perform the Gaussian smoothing on catadioptric images. We perform experiments with synthetic and real images to validate the generalization of our approach to any central catadioptric system.
Luis Puig, Josechu J. Guerrero
ICCV2
2011 Calibration of Central Catadioptric Cameras Using a DLT-Like Approach
Luis Puig, Yalin Bastanlar, Peter F. Sturm, Josechu J. Guerrero, João Pedro Barreto 0001
Int. J. Comput. Vis.4
2010 Visual SLAM with an Omnidirectional Camera
abstract
In this work we integrate the Spherical Camera Model for catadioptric systems in a Visual-SLAM application. The Spherical Camera Model is a projection model that unifies central catadioptric and conventional cameras. To integrate this model into the Extended Kalman Filter-based SLAM we require to linearize the direct and the inverse projection. We have performed an initial experimentation with omni directional and conventional real sequences including challenging trajectories. The results confirm that the omni directional camera gives much better orientation accuracy improving the estimated camera trajectory.
Alejandro Rituerto, Luis Puig, Josechu J. Guerrero
ICPR3
2010 Self-orientation of a hand-held catadioptric system in man-made environments
abstract
In central catadioptric systems the 3D lines are projected into conics, actually degenerate conics. In this paper we present a new approach to extract the projected lines corresponding to straight lines in the scene and to compute vanishing points from them. Using the internal calibration and two image points we are able to compute the catadioptric image lines analytically. We exploit the presence of parallel lines in man-made environments to compute the dominant vanishing points in the omnidirectional image. In order to obtain the intersection of two of these conics to compute vanishing points we analyze the self-polar triangle common to this pair. With the information contained in the vanishing points we are able to obtain the self-orientation of a hand-held catadioptric system. This system can be used in a vertical stabilization system required by autonomous navigation or to rectify images required in applications where the vertical orientation of the catadioptric system is assumed. We test our approach performing vertical and full rectifications in real sequences of images.
Luis Puig, Jesús Bermúdez, Josechu J. Guerrero
ICRA3
2010 Projective active shape models for pose-variant image analysis of quasi-planar objects: Application to facial analysis
Federico Sukno, Josechu J. Guerrero, Alejandro F. Frangi
Pattern Recognit.2
2010 Homography-Based Control Scheme for Mobile Robots With Nonholonomic and Field-of-View Constraints
abstract
In this paper, we present a visual servo controller that effects optimal paths for a nonholonomic differential drive robot with field-of-view constraints imposed by the vision system. The control scheme relies on the computation of homographies between current and goal images, but unlike previous homography-based methods, it does not use the homography to compute estimates of pose parameters. Instead, the control laws are directly expressed in terms of individual entries in the homography matrix. In particular, we develop individual control laws for the three path classes that define the language of optimal paths: rotations, straight-line segments, and logarithmic spirals. These control laws, as well as the switching conditions that define how to sequence path segments, are defined in terms of the entries of homography matrices. The selection of the corresponding control law requires the homography decomposition before starting the navigation. We provide a controllability and stability analysis for our system and give experimental results.
Gonzalo López-Nicolás, Nicholas R. Gans, Sourabh Bhattacharya, Carlos Sagüés, Josechu J. Guerrero, Seth Hutchinson 0001
IEEE Trans. Syst. Man Cybern. Part B5
2009 Parking with the essential matrix without short baseline degeneracies
abstract
This paper addresses the problem of visual control of a mobile robot. The system consists of a calibrated camera fixed onboard a robot with nonholonomic motion constraints. The parking task is defined by a reference image taken at the target location. The proposed control law is based on the essential matrix, but unlike traditional methods, it is not used to compute pose parameters. Instead, the control law is defined directly in terms of individual entries of the essential matrix by means of the input-output linearization of the system. Here we solve the problem of degeneracies due to short baseline by taking advantage of the planar motion constraint of the robot. Thus, a virtual target is defined providing a stable estimation of the essential matrix without degeneracies despite short baseline.
Gonzalo López-Nicolás, Carlos Sagüés, Josechu J. Guerrero
ICRA3
2009 Visual homing for undulatory robotic locomotion
abstract
This paper addresses the problem of vision-based closed-loop control for undulatory robots. We present an image-based visual servoing scheme, which drives the robot to a desired location specified by a target image, without explicitly estimating its pose. Instead, the control relies on the computation of the epipolar geometry between the current and target images. We analyze controllability and stability of the proposed control scheme, which is validated by simulation studies using the SIMUUN computational tools. Preliminary experiments, involving the Nereisbot undulatory robotic prototype, are also presented.
Gonzalo López-Nicolás, Michael Sfakiotakis, Dimitris P. Tsakiris, Antonis A. Argyros, Carlos Sagüés, Josechu J. Guerrero
ICRA6
2009 Improving topological maps for safer and robust navigation
abstract
Nowadays we frequently find big amounts of data to work with, what facilitates many robotic tasks and helps to solve perception problems. At the same time, this fact origins an interesting ongoing research problem: how to organize and arrange big sets of information to be useful in later uses. Topological mapping is a very useful tool to arrange and deal with big amounts of reference images for robotic tasks. There are many previous works on topological mapping and many others use this kind of maps for topological localization, planning and navigation. This work is focused on the problem of carefully design topological map building processes that facilitate the posterior robot tasks that use them and make them safer. We propose a new hierarchy of topological maps focused on this aspect. The experiments included in this paper were run outdoors using omnidirectional images and GPS information, and show the good topological maps obtained and how they allow robust and safer localization and navigation tasks.
Ana Cristina Murillo, Pablo Abad, Josechu J. Guerrero, Carlos Sagüés
IROS3
2009 Self-location from monocular uncalibrated vision using reference omniviews
abstract
In this paper we present a novel approach to perform indoor self-localization using reference omnidirectional images. We only need one omnidirectional image of the whole scene stored in the robot memory and a conventional uncalibrated on-board camera. We match the omnidirectional image and the conventional images captured by the on-board camera and compute the hybrid epipolar geometry using lifted coordinates and robust techniques. We map the epipole in the reference omnidirectional image to a ground plane through a homography in lifted coordinates also, giving the position of the robot in the planar ground, and its uncertainty. We perform experiments with simulated and real data to show the feasibility of this new self-localization approach.
Luis Puig, Josechu J. Guerrero
IROS2
2008 Localization and Matching Using the Planar Trifocal Tensor With Bearing-Only Data
abstract
This paper addresses the robot and landmark localization problem from bearing-only data in three views, simultaneously to the robust association of this data. The localization algorithm is based on the 1-D trifocal tensor, which relates linearly the observed data and the robot localization parameters. The aim of this work is to bring this useful geometric construction from computer vision closer to robotic applications. One contribution is the evaluation of two linear approaches of estimating the 1-D tensor: the commonly used approach that needs seven bearing-only correspondences and another one that uses only five correspondences plus two calibration constraints. The results in this paper show that the inclusion of these constraints provides a simpler and faster solution and better estimation of robot and landmark locations in the presence of noise. Moreover, a new method that makes use of scene planes and requires only four correspondences is presented. This proposal improves the performance of the two previously mentioned methods in typical man-made scenarios with dominant planes, while it gives similar results in other cases. The three methods are evaluated with simulation tests as well as with experiments that perform automatic real data matching in conventional and omnidirectional images. The results show sufficient accuracy and stability to be used in robotic tasks such as navigation, global localization or initialization of simultaneous localization and mapping (SLAM) algorithms.
Josechu J. Guerrero, Ana Cristina Murillo, Carlos Sagüés
IEEE Trans. Robotics1
2007 View-invariant human feature extraction for video-surveillance applications
abstract
We present a view-invariant human feature extractor (shape+pose) for pedestrian monitoring in man-made environments. Our approach can be divided into 2 steps: firstly, a series of view-based models is built by discretizing the viewpoint with respect to the camera into several training views. During the online stage, the Homography that relates the image points to the closest and most adequate training plane is calculated using the dominant 3D directions. The input image is then warped to this training view and processed using the corresponding view-based model. After model fitting, the inverse transformation is performed on the resulting human features obtaining a segmented silhouette and a 2D pose estimation in the original input image. Experimental results demonstrate our system performs well, independently of the direction of motion, when it is applied to monocular sequences with high perspective effect.
Grégory Rogez, Josechu J. Guerrero, Carlos Orrite-Uruñuela
AVSS2
2007 Switched Homography-Based Visual Control of Differential Drive Vehicles with Field-of-View Constraints
abstract
This paper presents a switched homography-based visual control for differential drive vehicles. The goal is defined by an image taken at the desired position, which is the only previous information needed from the scene. The control takes into account the field-of-view constraints of the vision system through the specific design of the paths with optimality criteria. The optimal paths consist of straight lines and curves that saturate the sensor viewing angle. We present the controls that move the robot along these paths based on the convergence of the elements of the homography matrix. Our contribution is the design of the switched homography-based control, following optimal paths guaranteeing the visibility of the target.
Gonzalo López-Nicolás, Sourabh Bhattacharya, Josechu J. Guerrero, Carlos Sagüés, Seth Hutchinson 0001
ICRA3
2007 Homography-Based Visual Control of Nonholonomic Vehicles
abstract
This paper presents a new visual control approach based on homography. The method is intended for nonholonomic vehicles with a fixed monocular system on board. The idea of visual control used here is the usual approach where the desired position of the robot is given by a target image taken at that position. This target image is the only previous information needed by the control law to perform the navigation from the initial position to the target. The control law is designed by the input-output linearization of the system using elements of the homography as output. The contribution is a controller that deals with the nonholonomic constraints of the mobile platform needing neither decomposition of the homography nor depth estimation to the target.
Gonzalo López-Nicolás, Carlos Sagüés, Josechu J. Guerrero
ICRA3
2007 SURF features for efficient robot localization with omnidirectional images
abstract
Many robotic applications work with visual reference maps, which usually consist of sets of more or less organized images. In these applications, there is a compromise between the density of reference data stored and the capacity to identify later the robot localization, when it is not exactly in the same position as one of the reference views. Here we propose the use of a recently developed feature, SURF, to improve the performance of appearance-based localization methods that perform image retrieval in large data sets. This feature is integrated with a vision-based algorithm that allows both topological and metric localization using omnidirectional images in a hierarchical approach. It uses pyramidal kernels for the topological localization and three-view geometric constraints for the metric one. Experiments with several omnidirectional images sets are shown, including comparisons with other typically used features (radial lines and SIFT). The advantages of this approach are proved, showing the use of SURF as the best compromise between efficiency and accuracy in the results.
Ana Cristina Murillo, Josechu J. Guerrero, Carlos Sagüés
ICRA2
2006 Viewpoint Independent Human Motion Analysis in Man-made Environments
abstract
This work addresses the problem of human motion analysis in video sequences of a scene observed by a single fixed camera with high perspective effect. The goal of this work is to make a 2D-Model (made of Shape and Stick figure) viewpoint-insensitive and preprocess the input image for removing the perspective effect. We focus our methodology on using the 3D principal directions of man-made environments and also the direction of motion to transform both 2D-Model and input images to a common frontal view (parallel or orthogonal to the direction of motion) before the fitting process. The inverse transformation is then performed on the resulting human features obtaining a segmented silhouette and a pose estimation in the original input image. Preliminary results are very promising since the proposed algorithm is able to locate head and feet with a better precision than previous one. 1
Grégory Rogez, Josechu J. Guerrero, Jesús Martínez del Rincón, Carlos Orrite-Uruñuela
BMVC2
2006 Nonholonomic Epipolar Visual Servoing
abstract
A significant amount of work has been reported in the area of visual servoing during the last decade. However, most of the contributions are applied in cases of holonomic robots. More recently, the use of visual feedback for control of nonholonomic vehicles has been reported. Some of the examples are docking and parallel parking maneuvers of cars or vision-based stabilization of a mobile manipulator to a desired pose with respect to a target of interest. Still, many of the approaches are mostly interested in the control part of visual servoing loop considering very simple vision algorithms based on artificial markers. In this paper, we present an approach for nonholonomic visual servoing based on epipolar geometry. The method facilitates a classical teach-by-showing approach where a reference image is used to define the desired pose (position and orientation) of the robot. The major contribution of the paper is the design of the control law that considers nonholonomic constraints of the robot as well as the robust feature detection and matching process based on scale and rotation invariant image features. An extensive experimental evaluation has been performed in a realistic indoor setting and the results are summarized in the paper
Gonzalo López-Nicolás, Carlos Sagüés, Josechu J. Guerrero, Danica Kragic, Patric Jensfelt
ICRA3
2006 Localization with Omnidirectional Images using the Radial Trifocal Tensor
abstract
In this paper we present a technique to linearly recover 2D structure and motion in man made environments from three uncalibrated omnidirectional views. We use vertical lines from the scene which are projected as radial lines in the images and are automatically matched. The algorithm is based on a 1D radial trifocal tensor which encodes the relations of the three views and the projected lines. We include experiments with real images, which demonstrate the good performance of the method and its application to robotic tasks, such as robot localization based in a database of reference images or to obtain the initial values of robot and landmarks localization in SLAM algorithms
Carlos Sagüés, Ana Cristina Murillo, Josechu J. Guerrero, Toon Goedemé, Tinne Tuytelaars, Luc Van Gool
ICRA3
2006 Robot and Landmark Localization using Scene Planes and the 1D Trifocal Tensor
abstract
This paper presents a method for robot and landmarks 2D localization, in man made environments, taking profit of scene planes. The method uses bearing-only measurements that are robustly matched in three views. In our experiments we obtain them from vertical lines corresponding to natural landmarks. With these three view line-matches a trifocal tensor can be computed. This tensor contains the three views geometry and is used to estimate the aforementioned localization. As it is very usual to find a planar surface, we use the homography corresponding to that plane to obtain the tensor with one match less than the general case method. This implies lower computational complexity, mainly when trying a robust estimation, where we see a reduction in the number of iterations needed. Another advantage of obtaining an homography during the process is that it can help to automatically detect singular situations, such us totally planar scenes. It is shown that our proposal performs similarly to the general case method in a general scenario and better in case that we have some dominant plane in the scene. This paper includes simulated results proving this, as well as examples with real images, both with conventional and omnidirectional cameras
Ana Cristina Murillo, Josechu J. Guerrero, Carlos Sagüés
IROS2
2006 From lines to epipoles through planes in two views
Carlos Sagüés, Ana Cristina Murillo, F. Escudero, Josechu J. Guerrero
Pattern Recognit.4
1999 Camera motion from brightness on lines. Combination of features and normal flow
Josechu J. Guerrero, Carlos Sagüés
Pattern Recognit.1