EDBT 2026 Demo / reviewers in the wild / expert
J. M. M. Montiel
dblp:22/5371 · also José M. M. Montiel, José María Martínez Montiel
· DBLP profile ↗
51ranked-venue papers
3as first author
14since 2021 · last 2025
0000-0002-3627-7306ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 3 first-author · 6 since 2021Systems, architecture and hardware · 23 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EndoMetric: Near-Light Monocular Metric Scale Estimation in Endoscopy
Raúl Iranzo, Victor M. Batlle, Juan D. Tardós, J. M. M. Montiel |
MICCAI (10) | 4 |
| 2024 | ColonMapper: topological mapping and localization for colonoscopyabstractWe propose a topological mapping and localization system able to operate on real human colonoscopies, despite significant shape and illumination changes. The map is a graph where each node codes a colon location by a set of real images, while edges represent traversability between nodes. For close-in-time images, where scene changes are minor, place recognition can be successfully managed with the recent transformers-based local feature matching algorithms. However, under long-term changes –such as different colonoscopies of the same patient– feature-based matching fails. To address this, we train on real colonoscopies a deep global descriptor achieving high recall with significant changes in the scene. The addition of a Bayesian filter boosts the accuracy of long-term place recognition, enabling relocalization in a previously built map. Our experiments show that ColonMapper is able to autonomously build a map and localize against it in two important use cases: localization within the same colonoscopy or within different colonoscopies of the same patient. Code: github.com/jmorlana/ColonMapper. Javier Morlana, Juan D. Tardós, J. M. M. Montiel |
ICRA | 3 |
| 2024 | Topological SLAM in Colonoscopies Leveraging Deep Features and Topological Priors
Javier Morlana, Juan D. Tardós, J. M. M. Montiel |
MICCAI (11) | 3 |
| 2024 | SimCol3D - 3D reconstruction during colonoscopy challengeabstractColorectal cancer is one of the most common cancers in the world. While colonoscopy is an effective screening technique, navigating an endoscope through the colon to detect polyps is challenging. A 3D map of the observed surfaces could enhance the identification of unscreened colon tissue and serve as a training platform. However, reconstructing the colon from video footage remains difficult. Learning-based approaches hold promise as robust alternatives, but necessitate extensive datasets. Establishing a benchmark dataset, the 2022 EndoVis sub-challenge SimCol3D aimed to facilitate data-driven depth and pose prediction during colonoscopy. The challenge was hosted as part of MICCAI 2022 in Singapore. Six teams from around the world and representatives from academia and industry participated in the three sub-challenges: synthetic depth prediction, synthetic pose prediction, and real pose prediction. This paper describes the challenge, the submitted methods, and their results. We show that depth prediction from synthetic colonoscopy images is robustly solvable, while pose estimation remains an open research question. Anita Rau, Sophia Bano, Yueming Jin, Pablo Azagra, Javier Morlana, Rawen Kader, Edward Sanderson, Bogdan J. Matuszewski, Erez Posner, Netanel Frank, Varshini Elangovan, Sista Raviteja, Zhengwen Li, Jiquan Liu, Seenivasan Lalithkumar, Mobarakol Islam, Hongliang Ren 0001, Laurence B. Lovat, J. M. M. Montiel, Danail Stoyanov |
Medical Image Anal. | 21 |
| 2024 | NR-SLAM: Nonrigid Monocular SLAMabstractThis article presents NR-SLAM, a novel nonrigid monocular simultaneous localization and mapping (SLAM) system founded on the combination of a dynamic deformation graph with a visco-elastic deformation model. The former enables our system to represent the dynamics of the deforming environment as the camera explores, while the later allows us to model general deformations in a simple way. The presented system is able to automatically initialize and extend a map modeled by a sparse point cloud in deforming environments, that is refined with a sliding-window deformable bundle adjustment. This map serves as base for the estimation of the camera motion and deformation and enables us to represent arbitrary surface topologies, overcoming the limitations of previous methods. To assess the performance of our system in challenging deforming scenarios, we evaluate it in several representative medical datasets. In our experiments, NR-SLAM outperforms previous deformable SLAM systems, achieving millimeter reconstruction accuracy and bringing automated medical intervention closer. For the benefit of the community, we make the source code public. Juan J. Gómez Rodríguez, J. M. M. Montiel, Juan D. Tardós |
IEEE Trans. Robotics | 2 |
| 2023 | LightDepth: Single-View Depth Self-Supervision from Illumination DeclineabstractSingle-view depth estimation can be remarkably effective if there is enough ground-truth depth data for supervised training. However, there are scenarios, especially in medicine in the case of endoscopies, where such data cannot be obtained. In such cases, multi-view self-supervision and synthetic-to-real transfer serve as alternative approaches, however, with a considerable performance reduction in comparison to supervised case. Instead, we propose a single-view self-supervised method that achieves a performance similar to the supervised case. In some medical devices, such as endoscopes, the camera and light sources are co-located at a small distance from the target surfaces. Thus, we can exploit that, for any given albedo and surface orientation, pixel brightness is inversely proportional to the square of the distance to the surface, providing a strong single-view self-supervisory signal. In our experiments, our self-supervised models deliver accuracies comparable to those of fully supervised ones, while being applicable without depth ground-truth data. Javier Rodriguez Puigvert, Victor M. Batlle, J. M. M. Montiel, Ruben Martinez-Cantin, Pascal Fua, Juan D. Tardós, Javier Civera 0001 |
ICCV | 3 |
| 2023 | Reuse your features: unifying retrieval and feature-metric alignmentabstractWe propose a compact pipeline to unify all the steps of Visual Localization: image retrieval, candidate re-ranking and initial pose estimation, and camera pose refinement. Our key assumption is that the deep features used for these individual tasks share common characteristics, so we should reuse them in all the procedures of the pipeline. Our DRAN (Deep Retrieval and image Alignment Network) is able to extract global descriptors for efficient image retrieval, use intermediate hierarchical features to re-rank the retrieval list and produce an initial pose guess, which is finally refined by means of a feature-metric optimization based on learned deep multi-scale dense features. DRAN is the first single network able to produce the features for the three steps of visual localization. DRAN achieves competitive performance in terms of robustness and accuracy under challenging conditions in public benchmarks, outperforming other unified approaches and consuming lower computational and memory cost than its counterparts using multiple networks. Code and models will be publicly available at github.com/jmorlana/DRAN. Javier Morlana, J. M. M. Montiel |
ICRA | 2 |
| 2023 | Tracking Adaptation to Improve SuperPoint for 3D Reconstruction in Endoscopy
Oscar León Barbed, J. M. M. Montiel, Pascal Fua, Ana Cristina Murillo |
MICCAI (1) | 2 |
| 2023 | LightNeuS: Neural Surface Reconstruction in Endoscopy Using Illumination Decline
Victor M. Batlle, J. M. M. Montiel, Pascal Fua, Juan D. Tardós |
MICCAI (10) | 2 |
| 2022 | Photometric single-view dense 3D reconstruction in endoscopyabstractVisual SLAM inside the human body will open the way to computer-assisted navigation in endoscopy. However, due to space limitations, medical endoscopes only provide monocular images, leading to systems lacking true scale. In this paper, we exploit the controlled lighting in colonoscopy to achieve the first in-vivo 3D reconstruction of the human colon using photometric stereo on a calibrated monocular endoscope. Our method works in a real medical environment, providing both a suitable in-place calibration procedure and a depth estimation technique adapted to the colon's tubular geometry. We validate our method on simulated colonoscopies, obtaining a mean error of 7% on depth estimation, which is below 3 mm on average. Our qualitative results on the EndoMapper dataset show that the method is able to correctly estimate the colon shape in real human colonoscopies, paving the ground for truescale monocular SLAM in endoscopy. Victor M. Batlle, J. M. M. Montiel, Juan D. Tardós |
IROS | 2 |
| 2022 | Tracking monocular camera pose and deformation for SLAM inside the human bodyabstractMonocular SLAM in deformable scenes will open the way to multiple medical applications like computer-assisted navigation in endoscopy, automatic drug delivery or autonomous robotic surgery. In this paper we propose a novel method to simultaneously track the camera pose and the 3D scene deformation, without any assumption about environment topology or shape. The method uses an illumination-invariant photometric method to track image features and estimates camera motion and deformation combining reprojection error with spatial and temporal regularization of deformations. Our results in simulated colonoscopies show the method's accuracy and robustness in complex scenes under increasing levels of deformation. Our qualitative results in human colonoscopies from Endomapper dataset show that the method is able to successfully cope with the challenges of real endoscopies: deformations, low texture and strong illumination changes. We also compare with previous tracking methods in simpler scenarios from Hamlyn dataset where we obtain competitive performance, without needing any topological assumption. Juan J. Gómez Rodríguez, J. M. M. Montiel, Juan D. Tardós |
IROS | 2 |
| 2021 | SD-DefSLAM: Semi-Direct Monocular SLAM for Deformable and Intracorporeal ScenesabstractConventional SLAM techniques strongly rely on scene rigidity to solve data association, ignoring dynamic parts of the scene. In this work we present Semi-Direct DefSLAM (SD-DefSLAM), a novel monocular deformable SLAM method able to map highly deforming environments, built on top of DefSLAM [1]. To robustly solve data association in challenging deforming scenes, SD-DefSLAM combines direct and indirect methods: an enhanced illumination-invariant Lucas-Kanade tracker for data association, geometric Bundle Adjustment for pose and deformable map estimation, and bag-of-words based on feature descriptors for camera relocalization. Dynamic objects are detected and segmented-out using a CNN trained for the specific application domain.We thoroughly evaluate our system in two public datasets. The mandala dataset is a SLAM benchmark with increasingly aggressive deformations. The Hamlyn dataset contains intracorporeal sequences that pose serious real-life challenges beyond deformation like weak texture, specular reflections, surgical tools and occlusions. Our results show that SD-DefSLAM outperforms DefSLAM in point tracking, reconstruction accuracy and scale drift thanks to the improvement in all the data association steps, being the first system able to robustly perform SLAM inside the human body. Juan J. Gómez Rodríguez, José Lamarca, Javier Morlana, Juan D. Tardós, J. M. M. Montiel |
ICRA | 5 |
| 2021 | ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual-Inertial, and Multimap SLAMabstractThis article presents ORB-SLAM3, the first system able to perform visual, visual-inertial and multimap SLAM with monocular, stereo and RGB-D cameras, using pin-hole and fisheye lens models. The first main novelty is a tightly integrated visual-inertial SLAM system that fully relies on maximuma posteriori(MAP) estimation, even during IMU initialization, resulting in real-time robust operation in small and large, indoor and outdoor environments, being two to ten times more accurate than previous approaches. The second main novelty is a multiple map system relying on a new place recognition method with improved recall that lets ORB-SLAM3 survive to long periods of poor visual information: when it gets lost, it starts a new map that will be seamlessly merged with previous maps when revisiting them. Compared with visual odometry systems that only use information from the last few seconds, ORB-SLAM3 is the first system able to reuse in all the algorithm stages all previous information from high parallax co-visible keyframes, even if they are widely separated in time or come from previous mapping sessions, boosting accuracy. Our experiments show that, in all sensor configurations, ORB-SLAM3 is as robust as the best systems available in the literature and significantly more accurate. Notably, our stereo-inertial SLAM achieves an average accuracy of 3.5 cm in the EuRoC drone and 9 mm under quick hand-held motions in the room of TUM-VI dataset, representative of AR/VR scenarios. For the benefit of the community we make public the source code. Carlos Campos 0001, Richard Elvira, Juan J. Gómez Rodríguez, J. M. M. Montiel, Juan D. Tardós |
IEEE Trans. Robotics | 4 |
| 2021 | DefSLAM: Tracking and Mapping of Deforming Scenes From Monocular SequencesabstractMonocular simultaneous localization and mapping (SLAM) algorithms perform robustly when observing rigid scenes; however, they fail when the observed scene deforms, for example, in medical endoscopy applications. In this article, we present DefSLAM, the first monocular SLAM capable of operating in deforming scenes in real time. Our approach intertwines Shape-from-Template (SfT) and Non-Rigid Structure-from-Motion (NRSfM) techniques to deal with the exploratory sequences typical of SLAM. A deformation tracking thread recovers the pose of the camera and the deformation of the observed map, at frame rate, by means of SfT processing a template that models the scene shape-at-rest. A deformation mapping thread runs in parallel with the tracking to update the template, at keyframe rate, by means of an isometric NRSfM processing a batch of full perspective keyframes. In our experiments, DefSLAM processes close-up sequences of deforming scenes, both in a laboratory-controlled experiment and in medical endoscopy sequences, producing accurate 3-D models of the scene with respect to the moving camera. José Lamarca, Shaifali Parashar, Adrien Bartoli, J. M. M. Montiel |
IEEE Trans. Robotics | 4 |
| 2020 | Inertial-Only Optimization for Visual-Inertial InitializationabstractWe formulate for the first time visual-inertial initialization as an optimal estimation problem, in the sense of maximum-a-posteriori (MAP) estimation. This allows us to properly take into account IMU measurement uncertainty, which was neglected in previous methods that either solved sets of algebraic equations, or minimized ad-hoc cost functions using least squares. Our exhaustive initialization tests on EuRoC dataset show that our proposal largely outperforms the best methods in the literature, being able to initialize in less than 4 seconds in almost any point of the trajectory, with a scale error of 5.3% on average. This initialization has been integrated into ORB-SLAM Visual-Inertial boosting its robustness and efficiency while maintaining its excellent accuracy. Carlos Campos 0001, J. M. M. Montiel, Juan D. Tardós |
ICRA | 2 |
| 2020 | Direct Sparse MappingabstractPhotometric bundle adjustment (PBA) accurately estimates geometry from video. However, current PBA systems have a temporary map that cannot manage scene reobservations. We present, direct sparse mapping, a full monocular visual simultaneous localization and mapping (SLAM) based on PBA. Its persistent map handles reobservations, yielding the most accurate results up to date on EuRoC for a direct method. Jon Zubizarreta, Iker Aguinaga, J. M. M. Montiel |
IEEE Trans. Robotics | 3 |
| 2019 | Fast and Robust Initialization for Visual-Inertial SLAMabstractVisual-inertial SLAM (VI-SLAM) requires a good initial estimation of the initial velocity, orientation with respect to gravity and gyroscope and accelerometer biases. In this paper we build on the initialization method proposed by Martinelli [1] and extended by Kaiser et al. [2], modifying it to be more general and efficient. We improve accuracy with several rounds of visual-inertial bundle adjustment, and robustify the method with novel observability and consensus tests, that discard erroneous solutions. Our results on the EuRoC dataset show that, while the original method produces scale errors up to 156%, our method is able to consistently initialize in less than two seconds with scale errors around 5%, which can be further reduced to less than 1% performing visual-inertial bundle adjustment after ten seconds. Carlos Campos 0001, J. M. M. Montiel, Juan D. Tardós |
ICRA | 2 |
| 2019 | ORBSLAM-Atlas: a robust and accurate multi-map systemabstractWe propose ORBSLAM-Atlas, a system able to handle an unlimited number of disconnected sub-maps, that includes a robust map merging algorithm able to detect submaps with common regions and seamlessly fuse them. The outstanding robustness and accuracy of ORBSLAM are due to its ability to detect wide-baseline matches between keyframes, and to exploit them by means of non-linear optimization, however it only can handle a single map. ORBSLAM-Atlas brings the wide-baseline matching detection and exploitation to the multiple map arena. The result is a SLAM system significantly more general and robust, able to perform multisession mapping. If tracking is lost during exploration, instead of freezing the map, a new sub-map is launched, and it can be fused with the previous map when common parts are visited. Our criteria to declare the camera lost contrast with previous approaches that simply count the number of tracked points, we propose to discard also inaccurately estimated camera poses due to bad geometrical conditioning. As a result, the map is split into more accurate sub-maps, that are eventually merged in a more accurate global map, thanks to the multi-mapping capabilities.We provide extensive experimental validation in the EuRoC datasets, where ORBSLAM-Atlas obtains accurate monocular and stereo results in the difficult sequences where ORBSLAM failed. We also build global maps after multiple sessions in the same room, obtaining the best results to date, between 2 and 3 times more accurate than competing multi-map approaches. We also show the robustness and capability of our system to deal with dynamic scenes, quantitatively in the EuRoC datasets and qualitatively in a densely populated corridor where camera occlusions and tracking losses are frequent. Richard Elvira, Juan D. Tardós, J. M. M. Montiel |
IROS | 3 |
| 2019 | Live Tracking and Dense Reconstruction for Handheld Monocular EndoscopyabstractContemporary endoscopic simultaneous localization and mapping (SLAM) methods accurately compute endoscope poses; however, they only provide a sparse 3-D reconstruction that poorly describes the surgical scene. We propose a novel dense SLAM method whose qualities are: 1) monocular, requiring only RGB images of a handheld monocular endoscope; 2) fast, providing endoscope positional tracking and 3-D scene reconstruction, running in parallel threads; 3) dense, yielding an accurate dense reconstruction; 4) robust, to the severe illumination changes, poor texture and small deformations that are typical in endoscopy; and 5) self-contained, without needing any fiducials nor external tracking devices and, therefore, it can be smoothly integrated into the surgical workflow. It works as follows. First, accurate cluster frame poses are estimated using the sparse SLAM feature matches. The system segments clusters of video frames according to parallax criteria. Next, dense matches between cluster frames are computed in parallel by a variational approach that combines zero mean normalized cross correlation and a gradient Huber norm regularizer. This combination copes with challenging lighting and textures at an affordable time budget on a modern GPU. It can outperform pure stereo reconstructions, because the frames cluster can provide larger parallax from the endoscope's motion. We provide an extensive experimental validation on real sequences of the porcine abdominal cavity, both in-vivo and ex-vivo. We also show a qualitative evaluation on human liver. In addition, we show a comparison with the other dense SLAM methods showing the performance gain in terms of accuracy, density, and computation time. Nader Mahmoud, Toby Collins, Alexandre Hostettler, Luc Soler, Christophe Doignon, J. M. M. Montiel |
IEEE Trans. Medical Imaging | 6 |
| 2016 | Mode-shape interpretation: Re-thinking modal space for recovering deformable shapesabstractThis paper describes an on-line approach for estimating non-rigid shape and camera pose from monocular video sequences. We assume an initial estimate of the shape at rest to be given and represented by a triangulated mesh, which is encoded by a matrix of the distances between every pair of vertexes. By applying spectral analysis on this matrix, we are then able to compute a low-dimensional shape basis, that in contrast to standard approaches, has a very direct physical interpretation and requires a much smaller number of modes to span a large variety of deformations, either for inextensible or extensible configurations. Based on this low-rank model, we then sequentially retrieve both camera motion and non-rigid shape in each image, optimizing the model parameters with bundle adjustment over a sliding window of image frames. Since the number of these parameters is small, specially when considering physical priors, our approach may potentially achieve real-time performance. Experimental results on real videos for different scenarios demonstrate remarkable robustness to artifacts such as missing and noisy observations. Antonio Agudo, J. M. M. Montiel, Begoña Calvo, Francesc Moreno-Noguer |
WACV | 2 |
| 2016 | Real-time 3D reconstruction of non-rigid shapes with a single moving camera
Antonio Agudo, Francesc Moreno-Noguer, Begoña Calvo, J. M. M. Montiel |
Comput. Vis. Image Underst. | 4 |
| 2016 | Sequential Non-Rigid Structure from Motion Using Physical PriorsabstractWe propose a new approach to simultaneously recover camera pose and 3D shape of non-rigid and potentially extensible surfaces from a monocular image sequence. For this purpose, we make use of the Extended Kalman Filter based Simultaneous Localization And Mapping (EKF-SLAM) formulation, a Bayesian optimization framework traditionally used in mobile robotics for estimating camera pose and reconstructing rigid scenarios. In order to extend the problem to a deformable domain we represent the object's surface mechanics by means of Navier's equations, which are solved using a Finite Element Method (FEM). With these main ingredients, we can further model the material's stretching, allowing us to go a step further than most of current techniques, typically constrained to surfaces undergoing isometric deformations. We extensively validate our approach in both real and synthetic experiments, and demonstrate its advantages with respect to competing methods. More specifically, we show that besides simultaneously retrieving camera pose and non-rigid shape, our approach is adequate for both isometric and extensible surfaces, does not require neither batch processing all the frames nor tracking points over the whole sequence and runs at several frames per second. Antonio Agudo, Francesc Moreno-Noguer, Begoña Calvo, J. M. M. Montiel |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2015 | Layout aware visual tracking and mappingabstractNowadays real time visual Simultaneous Localization And Mapping (SLAM) algorithms exist and rely on consistent measurements across multiple views. In indoor environments, where majority of robot's activity takes place, severe occlusions can occur, e.g., when turning around a corner or moving from one room to another. In these situations, SLAM algorithms can not establish correspondences across views, which leads to failures in camera localization or map construction. This work takes advantage of the recent scene box layout descriptor to make the above mentioned SLAM systems occlusion aware. This room box reasoning helps the sequential tracker to reason about possible occlusions and therefore look for matches in only potentially visible features instead of the entire map. This increases the life of the tracker, as it does not consider itself lost under the occlusion state. Additionally, focusing on the potentially visible portion of the map, i.e., the current room features, it improves the computational efficiency without compromising the accuracy. Finally, this room level reasoning helps in better image selection for bundle adjustment. The image bundle coming from the same room has little occlusion, which leads to better dense reconstruction. We demonstrate the superior performance of layout aware SLAM on several long monocular sequences acquired in difficult indoor situations, specifically in a room-room transition and turning around a corner. Marta Salas, Muhammad Wajahat Hussain, Alejo Concha, Luis Montano, Javier Civera 0001, J. M. M. Montiel |
IROS | 6 |
| 2015 | RoboEarth Semantic Mapping: A Cloud Enabled Knowledge-Based ApproachabstractThe vision of the RoboEarth project is to design a knowledge-based system to provide web and cloud services that can transform a simple robot into an intelligent one. In this work, we describe the RoboEarth semantic mapping system. The semantic map is composed of: 1) an ontology to code the concepts and relations in maps and objects and 2) a SLAM map providing the scene geometry and the object locations with respect to the robot. We propose to ground the terminological knowledge in the robot perceptions by means of the SLAM map of objects. RoboEarth boosts mapping by providing: 1) a subdatabase of object models relevant for the task at hand, obtained by semantic reasoning, which improves recognition by reducing computation and the false positive rate; 2) the sharing of semantic maps between robots; and 3) software as a service to externalize in the cloud the more intensive mapping computations, while meeting the mandatory hard real time constraints of the robot. To demonstrate the RoboEarth cloud mapping system, we investigate two action recipes that embody semantic map building in a simple mobile robot. The first recipe enables semantic map building for a novel environment while exploiting available prior information about the environment. The second recipe searches for a novel object, with the efficiency boosted thanks to the reasoning on a semantically annotated map. Our experimental results demonstrate that, by using RoboEarth cloud services, a simple robot can reliably and efficiently build the semantic maps needed to perform its quotidian tasks. In addition, we show the synergetic relation of the SLAM map of objects that grounds the terminological knowledge coded in the ontology. Luis Riazuelo, Moritz Tenorth, Daniel Di Marco, Marta Salas, Dorian Gálvez-López, Lorenz Mösenlechner, Lars Kunze, Michael Beetz, Juan D. Tardós, Luis Montano, J. M. M. Montiel |
IEEE Trans Autom. Sci. Eng. | 11 |
| 2015 | ORB-SLAM: A Versatile and Accurate Monocular SLAM SystemabstractThis paper presents ORB-SLAM, a feature-based monocular simultaneous localization and mapping (SLAM) system that operates in real time, in small and large indoor and outdoor environments. The system is robust to severe motion clutter, allows wide baseline loop closing and relocalization, and includes full automatic initialization. Building on excellent algorithms of recent years, we designed from scratch a novel system that uses the same features for all SLAM tasks: tracking, mapping, relocalization, and loop closing. A survival of the fittest strategy that selects the points and keyframes of the reconstruction leads to excellent robustness and generates a compact and trackable map that only grows if the scene content changes, allowing lifelong operation. We present an exhaustive evaluation in 27 sequences from the most popular datasets. ORB-SLAM achieves unprecedented performance with respect to other state-of-the-art monocular SLAM approaches. For the benefit of the community, we make the source code public. Raul Mur-Artal, J. M. M. Montiel, Juan D. Tardós |
IEEE Trans. Robotics | 2 |
| 2014 | Online Dense Non-Rigid 3D Shape and Camera Motion Recovery
Antonio Agudo, J. M. M. Montiel, Lourdes Agapito, Begoña Calvo |
BMVC | 2 |
| 2014 | Good Vibrations: A Modal Analysis Approach for Sequential Non-rigid Structure from MotionabstractWe propose an online solution to non-rigid structure from motion that performs camera pose and 3D shape estimation of highly deformable surfaces on a frame-by-frame basis. Our method models non-rigid deformations as a linear combination of some mode shapes obtained using modal analysis from continuum mechanics. The shape is first discretized into linear elastic triangles, modelled by means of finite elements, which are used to pose the force balance equations for an undamped free vibrations model. The shape basis computation comes down to solving an eigenvalue problem, without the requirement of a learning step. The camera pose and time varying weights that define the shape at each frame are then estimated on the fly, in an online fashion, using bundle adjustment over a sliding window of image frames. The result is a low computational cost method that can run sequentially in real-time. We show experimental results on synthetic sequences with ground truth 3D data and real videos for different scenarios ranging from sparse to dense scenes. Our system exhibits a good trade-off between accuracy and computational budget, it can handle missing data and performs favourably compared to competing methods. Antonio Agudo, Lourdes Agapito, Begoña Calvo, J. M. M. Montiel |
CVPR | 4 |
| 2014 | Visual SLAM for Handheld Monocular EndoscopeabstractSimultaneous localization and mapping (SLAM) methods provide real-time estimation of 3-D models from the sole input of a handheld camera, routinely in mobile robotics scenarios. Medical endoscopic sequences mimic a robotic scenario in which a handheld camera (monocular endoscope) moves along an unknown trajectory while observing an unknown cavity. However, the feasibility and accuracy of SLAM methods have not been extensively validated with human in vivo image sequences. In this work, we propose a monocular visual SLAM algorithm tailored to deal with medical image sequences in order to provide an up-to-scale 3-D map of the observed cavity and the endoscope trajectory at frame rate. The algorithm is validated over synthetic data and human in vivo sequences corresponding to 15 laparoscopic hernioplasties where accurate ground-truth distances are available. It can be concluded that the proposed procedure is: 1) noninvasive, because only a standard monocular endoscope and a surgical tool are used; 2) convenient, because only a hand-controlled exploratory motion is needed; 3) fast, because the algorithm provides the 3-D map and the trajectory in real time; 4) accurate, because it has been validated with respect to ground-truth; and 5) robust to inter-patient variability, because it has performed successfully over the validation sequences. Oscar G. Grasa, Ernesto Bernal, Santiago Casado, Ismael Gil, J. M. M. Montiel |
IEEE Trans. Medical Imaging | 5 |
| 2012 | Finite Element based sequential Bayesian Non-Rigid Structure from MotionabstractNavier's equations modelling linear elastic solid deformations are embedded within an Extended Kalman Filter (EKF) to compute a sequential Bayesian estimate for the Non-Rigid Structure from Motion problem. The algorithm processes every single frame of a sequence gathered with a full perspective camera. No prior data association is assumed because matches are computed within the EKF prediction-match-update cycle. Scene is coded as a Finite Element Method (FEM) elastic thin-plate solid, where the discretization nodes are the sparse set of scene points salient in the image. It is assumed a set of Gaussian forces acting on solid nodes to cause scene deformation. The EKF combines in a feedback loop an approximate FEM model and the frame rate measurements from the camera, resulting in an efficient method to embed Navier's equations without resorting to expensive non-linear FEM models. Classical FEM modelling has implied an interactive identification of boundary points to constrain the scene rigid motion, in this work this dissatisfying prior knowledge is no longer needed. The scene and camer rigid motion are combined in a unique pose vector and the estimation is coded relative to the camera. Additionally, the deforming effect of the Gaussian forces on the thin-plate is computed by means of the Moore-Penrose pseudoinverse of the FEM stiffness matrix. The proposed algorithm is validated with three real sequences gathered with hand-held camera observing isometric and non-isometric deformations. It is also shown the consistency of the EKF estimation with respect to ground truth computed from stereo. Antonio Agudo, Begoña Calvo, J. M. M. Montiel |
CVPR | 3 |
| 2012 | Creating and using RoboEarth object modelsabstractThis paper presented an approach to create 3D object models for robotic and vision applications in a fast and inexpensive way compared to established approaches. By using the RoboEarth system for storing the created object models users have world-wide access to the data and can immediately reuse a model as soon as it was created and uploaded. The approach shows general applicability for different kinds of cameras. In this work this was shown by two example implementations for the recognition process of objects. The quality of the recognition can be verified in the video. Combined with the knowledge saved in the RoboEarth database the objects can also be properly classified. Daniel Di Marco, Andreas Koch 0003, Oliver Zweigle, Kai Häussermann, Björn Schießle, Paul Levi, Dorian Gálvez-López, Luis Riazuelo, Javier Civera 0001, J. M. M. Montiel, Moritz Tenorth, Alexander Clifford Perzylo, Markus Waibel, René van de Molengraft |
ICRA | 10 |
| 2012 | Impact of Landmark Parametrization on Monocular EKF-SLAM with Points and Lines
Joan Solà, Teresa Vidal-Calleja, Javier Civera 0001, J. M. M. Montiel |
Int. J. Comput. Vis. | 4 |
| 2012 | Visual SLAM: Why filter?
Hauke Strasdat, J. M. M. Montiel, Andrew J. Davison |
Image Vis. Comput. | 2 |
| 2011 | Double window optimisation for constant time visual SLAMabstractWe present a novel and general optimisation framework for visual SLAM, which scales for both local, highly accurate reconstruction and large-scale motion with long loop closures. We take a two-level approach that combines accurate pose-point constraints in the primary region of interest with a stabilising periphery of pose-pose soft constraints. Our algorithm automatically builds a suitable connected graph of keyposes and constraints, dynamically selects inner and outer window membership and optimises both simultaneously. We demonstrate in extensive simulation experiments that our method approaches the accuracy of offline bundle adjustment while maintaining constant-time operation, even in the hard case of very loopy monocular camera motion. Furthermore, we present a set of real experiments for various types of visual sensor and motion, including large scale SLAM with both monocular and stereo cameras, loopy local browsing with either monocular or RGB-D cameras, and dense RGB-D object model building. Hauke Strasdat, Andrew J. Davison, J. M. M. Montiel, Kurt Konolige |
ICCV | 3 |
| 2011 | EKF monocular SLAM with relocalization for laparoscopic sequencesabstractIn recent years, research on visual SLAM has produced robust algorithms providing, in real time at 30 Hz, both the 3D model of the observed rigid scene and the 3D camera motion using as only input the gathered image sequence. These algorithms have been extensively validated in rigid human-made environments -indoor and outdoor- showing robust performance in dealing with clutter, occlusions or sudden motions. Medical endoscopic sequences naturally pose a monocular SLAM problem: an unknown camera motion in an unknown environment. The corresponding map would be useful in providing 3D information to assist surgeons, to support augmented reality insertions or to be exploited by medical robots. In this paper we propose the combination EKF Monocular SLAM + 1-Point RANSAC + Randomised List Relocalization to process laparoscopic sequences -abdominal cavity images-. The sequences are challenging due to: 1) cluttering produced by tools; 2) sudden motions of the camera; 3) laparoscope frequently goes in and out of abdominal cavity; 4) tissue deformation caused by respiration, heartbeats and/or surgical tools. Real medical image sequences provide experimental validation. Oscar G. Grasa, Javier Civera 0001, J. M. M. Montiel |
ICRA | 3 |
| 2011 | Towards semantic SLAM using a monocular cameraabstractMonocular SLAM systems have been mainly focused on producing geometric maps just composed of points or edges; but without any associated meaning or semantic content. In this paper, we propose a semantic SLAM algorithm that merges in the estimated map traditional meaningless points with known objects. The non-annotated map is built using only the information extracted from a monocular image sequence. The known object models are automatically computed from a sparse set of images gathered by cameras that may be different from the SLAM camera. The models include both visual appearance and tridimensional information. The semantic or annotated part of the map -the objects- are estimated using the information in the image sequence and the precomputed object models. The proposed algorithm runs an EKF monocular SLAM parallel to an object recognition thread. This latest one informs of the presence of an object in the sequence by searching for SURF correspondences and checking afterwards their geometric compatibility. When an object is recognized it is inserted in the SLAM map, being its position measured and hence refined by the SLAM algorithm in subsequent frames. Experimental results show real-time performance for a hand held camera imaging a desktop environment and for a camera mounted in a robot moving in a room-sized scenario. Javier Civera 0001, Dorian Gálvez-López, Luis Riazuelo, Juan D. Tardós, J. M. M. Montiel |
IROS | 5 |
| 2010 | Real-time monocular SLAM: Why filter?abstractWhile the most accurate solution to off-line structure from motion (SFM) problems is undoubtedly to extract as much correspondence information as possible and perform global optimisation, sequential methods suitable for live video streams must approximate this to fit within fixed computational bounds. Two quite different approaches to real-time SFM - also called monocular SLAM (Simultaneous Localisation and Mapping) - have proven successful, but they sparsify the problem in different ways. Filtering methods marginalise out past poses and summarise the information gained over time with a probability distribution. Keyframe methods retain the optimisation approach of global bundle adjustment, but computationally must select only a small number of past frames to process. In this paper we perform the first rigorous analysis of the relative advantages of filtering and sparse optimisation for sequential monocular SLAM. A series of experiments in simulation as well using a real image SLAM system were performed by means of covariance propagation and Monte Carlo methods, and comparisons made using a combined cost/accuracy measure. With some well-discussed reservations, we conclude that while filtering may have a niche in systems with low processing resources, in most modern applications keyframe optimisation gives the most accuracy per unit of computing time. Hauke Strasdat, J. M. M. Montiel, Andrew J. Davison |
ICRA | 2 |
| 2009 | Camera self-calibration for sequential Bayesian structure from motionabstractComputer vision researchers have proved the feasibility of camera self-calibration —the estimation of a camera's internal parameters from an image sequence without any known scene structure. Various self-calibration algorithms have been published. Nevertheless, all of the recent sequential approaches to 3D structure and motion estimation from image sequences which have arisen in robotics and aim at real-time operation (often classed as visual SLAM or visual odometry) have relied on pre-calibrated cameras and have not attempted online calibration. Javier Civera 0001, Diana R. Bueno, Andrew J. Davison, J. M. M. Montiel |
ICRA | 4 |
| 2009 | 1-point RANSAC for EKF-based Structure from MotionabstractRecently, classical pairwise Structure From Motion (SfM) techniques have been combined with non-linear global optimization (Bundle Adjustment, BA) over a sliding window to recursively provide camera pose and feature location estimation from long image sequences. Normally called Visual Odometry, these algorithms are nowadays able to estimate with impressive accuracy trajectories of hundreds of meters; either from an image sequence (usually stereo) as the only input, or combining visual and propioceptive information from inertial sensors or wheel odometry. This paper has a double objective. First, we aim to illustrate for the first time how similar accuracy and trajectory length can be achieved by filtering-based visual SLAM methods. Specifically, a camera-centered Extended Kalman Filter is used here to process a monocular sequence as the only input, with 6DOF motion estimated. Features are kept live in the filter while visible as the camera explores forward, and are deleted from the state once they go out of view. This permits an increase in the number of tracked features per frame from tens to around a hundred. While improving the accuracy of the estimation, it makes computationally infeasible the exhaustive Branch and Bound search performed by standard JCBB for match outlier rejection. As a second contribution that overcomes this problem, we present here a RANSAC-like algorithm that exploits the probabilistic prediction of the filter. This use of prior information makes it possible to reduce the size of the minimal data subset to instantiate a hypothesis to the minimum possible of 1 point, greatly increasing the efficiency of the outlier rejection stage. Experimental results from real image sequences covering trajectories of hundreds of meters are presented and compared against RTK GPS ground truth. Estimation errors are about 1% of the trajectory for trajectories up to 650 metres. Javier Civera 0001, Oscar G. Grasa, Andrew J. Davison, J. M. M. Montiel |
IROS | 4 |
| 2009 | Drift-Free Real-Time Sequential Mosaicing
Javier Civera 0001, Andrew J. Davison, Juan A. Magallon, J. M. M. Montiel |
Int. J. Comput. Vis. | 4 |
| 2008 | Interacting multiple model monocular SLAMabstractRecent work has demonstrated the benefits of adopting a fully probabilistic SLAM approach in sequential motion and structure estimation from an image sequence. Unlike standard Structure from Motion (SFM) methods, this 'monocular SLAM' approach is able to achieve drift-free estimation with high frame-rate real-time operation, particularly benefitting from highly efficient active feature search, map management and mismatch rejection. A consistent thread in this research on real-time monocular SLAM has been to reduce the assumptions required. In this paper we move towards the logical conclusion of this direction by implementing a fully Bayesian Interacting Multiple Models (IMM) framework which can switch automatically between parameter sets in a dimensionless formulation of monocular SLAM. Remarkably, our approach of full sequential probability propagation means that there is no need for penalty terms to achieve the Occam property of favouring simpler models - this arises automatically. We successfully tackle the known stiffness in on-the-fly monocular SLAM start up without known patterns in the scene. The search regions for matches are also reduced in size with respect to single model EKF increasing the rejection of spurious matches. We demonstrate our method with results on a complex real image sequence with varied motion. Javier Civera 0001, Andrew J. Davison, J. M. M. Montiel |
ICRA | 3 |
| 2008 | Inverse Depth Parametrization for Monocular SLAMabstractWe present a new parametrization for point features within monocular simultaneous localization and mapping (SLAM) that permits efficient and accurate representation of uncertainty during undelayed initialization and beyond, all within the standard extended Kalman filter (EKF). The key concept is direct parametrization of the inverse depth of features relative to the camera locations from which they were first viewed, which produces measurement equations with a high degree of linearity. Importantly, our parametrization can cope with features over a huge range of depths, even those that are so far from the camera that they present little parallax during motion---maintaining sufficient representative uncertainty that these points retain the opportunity to "come in'' smoothly from infinity if the camera makes larger movements. Feature initialization is undelayed in the sense that even distant features are immediately used to improve camera motion estimates, acting initially as bearing references but not permanently labeled as such. The inverse depth parametrization remains well behaved for features at all stages of SLAM processing, but has the drawback in computational terms that each point is represented by a 6-D state vector as opposed to the standard three of a EuclideanXYZrepresentation. We show that once the depth estimate of a feature is sufficiently accurate, its representation can safely be converted to the EuclideanXYZform, and propose a linearity index that allows automatic detection and conversion to maintain maximum efficiency---only low parallax features need be maintained in inverse depth form for long periods. We present a real-time implementation at 30 Hz, where the parametrization is validated in a fully automatic 3-D SLAM system featuring a handheld single camera with no additional sensing. Experiments show robust operation in challenging indoor and outdoor environments with a very large ranges of scene depth, varied motion, and also real time 360degloop closing. Javier Civera 0001, Andrew J. Davison, J. M. M. Montiel |
IEEE Trans. Robotics | 3 |
| 2007 | Inverse Depth to Depth Conversion for Monocular SLAMabstractRecently it has been shown that an inverse depth parametrization can improve the performance of real-time monocular EKF SLAM, permitting undelayed initialization of features at all depths. However, the inverse depth parametrization requires the storage of 6 parameters in the state vector for each map point. This implies a noticeable computing overhead when compared with the standard 3 parameter XYZ Euclidean encoding of a 3D point, since the computational complexity of the EKF scales poorly with state vector size. In this work we propose to restrict the inverse depth parametrization only to cases where the standard Euclidean encoding implies a departure from linearity in the measurement equations. Every new map feature is still initialized using the 6 parameter inverse depth method. However, as the estimation evolves, if according to a linearity index the alternative XYZ coding can be considered linear, we show that feature parametrization can be transformed from inverse depth to XYZ for increased computational efficiency with little reduction in accuracy. We present a theoretical development of the necessary linearity indices, along with simulations to analyze the influence of the conversion threshold. Experiments performed with with a 30 frames per second real-time system are reported. An analysis of the increase in the map size that can be successfully managed is included. Javier Civera 0001, Andrew J. Davison, J. M. M. Montiel |
ICRA | 3 |
| 2006 | A Visual Compass based on SLAMabstractAccurate full 3 axis orientation is computed using a low cost calibrated camera. We present a simultaneous sensor location and mapping method that uses a purely rotating camera as sensor and distant points, ideally at infinity, as features. A smooth constant angular velocity pure rotation motion model codifies the camera location. Because of the sequential EKF approach used, and the number of features in the map, about a hundred, the proposed method has been implemented in real time at standard video rates. Experimental results with real images show that the system is able to close loops with 360deg pan and 360deg cyclotorsion rotations. Sequences show good performance under challenging conditions: hand-held camera, varying natural outdoor illumination, low cost camera and lens and people moving in the scene J. M. M. Montiel, Andrew J. Davison |
ICRA | 1 |
| 2006 | Adaptive Scale Robust Segmentation for 2D Laser ScannerabstractThis paper presents a robust algorithm for segmentation and line detection in 2D range scans. The described method exploits the multimodal probability density function of the residual error. It is capable of segmenting the range data in clusters, estimate the straight segments parameters, and estimate the scale of inliers error noise successfully, despite of high level of spurious data. No prior knowledge about the sensor and object properties is given to the algorithm. The mode seeking is based on mean shift algorithm, which has been widely used and tested in 3D laser scan segmentation, machine learning and pattern recognition applications. We show the reliability of the technique with experimental indoor and outdoor manmade environment. Compared with classical methods, a good compromise between false positive, false negative, wrong segment split and wrong segment merge is achieved, with improved accuracy in the estimated parameters. Ruben Martinez-Cantin, José A. Castellanos 0001, Juan D. Tardós, J. M. M. Montiel |
IROS | 4 |
| 2004 | Relocation using Laser and VisionabstractWe present a method for solving the first location problem using 2D laser and vision. Our observation is a two-dimensional laser scan together with its corresponding image. The observation is segmented into textured vertical planes; each vertical plane contains geometrical information about its location given by the laser scan, plus the gray level image obtained by the camera. The rich plane texture allows a safe plane recognition. Once two planes are recognized as correspondent, the computer vision geometry allows to compute the relative camera motion. The proposed algorithm outperforms both laser-only and vision-only algorithms. This is shown in the experimental results where a map composed of 8 observations of a 20/spl times/3 meter corridor is used to successfully locate the robot (without any other prior) in 163 out of 192 initial test robot locations. Diego Ortin, José Neira, J. M. M. Montiel |
ICRA | 3 |
| 2003 | Automated multisensor polyhedral model acquisitionabstractWe describe a method for automatically generating accurate piecewise planar models for indoor scenes using a combination of a 2D laser scanner and a camera on a mobile platform. The method exploits the complementarity of the sensors. Mapping techniques applied to 2D laser scans simultaneously compute a map and the location of the sensor in the unknown environment. This provides an initial estimate for the vision algorithms by compensating the rotation, foreshortening and the scale change between images. The vision algorithms are then able to compute a very accurate registration (via a plane to plane homography), which is used to segment the model into planar facets, and to improve the estimate of the model and sensor position. Results are demonstrated on a man made scene using a 2D laser scanner and a calibrated camera mounted on a trolley. Diego Ortin, J. M. M. Montiel, Andrew Zisserman |
ICRA | 2 |
| 2000 | Structure and motion from straight line segments
J. M. M. Montiel, Juan D. Tardós, Luis Montano |
Pattern Recognit. | 1 |
| 1999 | Goal Directed Reactive Robot Navigation with Relocation Using Laser and VisionabstractThis paper presents a method to perform a goal directed reactive navigation in unknown indoor environments. Two sensors cooperate to accomplish this task: trinocular vision and 3D laser rangefinder. Trinocular vision selects the initial goal location for the navigation task. Laser is used to accomplish a reactive navigation to avoid the obstacles and to periodically relocate the goal with respect to the robot, so the dead-reckoning drift is compensated. An extended Kalman filter is used to solve the data association problem and to perform the goal relocation while the robot navigates. Experimental results involving a real mobile robot are presented, validating the proposed method. José R. Asensio, J. M. M. Montiel, Luis Montano |
ICRA | 2 |
| 1999 | Continuous Mobile Robot Localization: Vision vs. LaserabstractWe present a comparative study of the performance of map-based robot localisation processes based on diverse sensing devices such as monocular and trinocular vision systems and laser rangefinders. We study both the precision (error with respect to the true values) and robustness (sensor measurements correctly paired with map features) of each localisation process. The experiment design we used allows one to compare these processes under exactly the same conditions. We conclude that comparable precision levels can be attained with each of the three sensors. With respect to robustness, monocular and trinocular vision pose more complex matching problems than laser, requiring more elaborate solutions to make the process robust. J. A. Pérez, José A. Castellanos 0001, J. M. M. Montiel, José Neira, Juan D. Tardós |
ICRA | 3 |
| 1999 | Probabilistic structure from camera location using straight segments
J. M. M. Montiel, Luis Montano |
Image Vis. Comput. | 1 |
| 1999 | The SPmap: a probabilistic framework for simultaneous localization and map buildingabstractThis article describes a rigorous and complete framework for the simultaneous localization and map building problem for mobile robots: the symmetries and perturbation map (SPmap), which is based on a general probabilistic representation of uncertain geometric information. We present a complete experiment with a LabMate/sup TM/ mobile robot navigating in a human-made indoor environment and equipped with a rotating 2D laser rangefinder. Experiments validate the appropriateness of our approach and provide a real measurement of the precision of the algorithms. José A. Castellanos 0001, J. M. M. Montiel, José Neira, Juan D. Tardós |
IEEE Trans. Robotics Autom. | 2 |