EDBT 2026 Demo / reviewers in the wild / expert
Javier Civera 0001
dblp:53/826 · also Javier Civera Sancho
· DBLP profile ↗
54ranked-venue papers
8as first author
23since 2021 · last 2025
0000-0003-1368-1151ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 6 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 15 since 2021Systems, architecture and hardware · 23 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MVSAnywhere: Zero-Shot Multi-View StereoabstractComputing accurate depth from multiple views is a fundamental and longstanding challenge in computer vision. However, most existing approaches do not generalize well across different domains and scene types (e.g. indoor vs. outdoor). Training a general-purpose multi-view stereo model is challenging and raises several questions, e.g. how to best make use of transformer-based architectures, how to incorporate additional metadata when there is a variable number of input views, and how to estimate the range of valid depths which can vary considerably across different scenes and is typically not known a priori? To address these issues, we introduce MVSA, a novel and versatile Multi-View Stereo architecture that aims to work Anywhere by generalizing across diverse domains and depth ranges. MVSA combines monocular and multi-view cues with an adaptive cost volume to deal with scale-related issues. We demonstrate state-of-the-art zero-shot depth estimation on the Robust Multi-View Depth Benchmark, surpassing existing multi-view stereo and monocular baselines. Sergio Izquierdo, Mohamed Sayed, Michael Firman, Guillermo Garcia-Hernando, Daniyar Turmukhambetov, Javier Civera 0001, Oisin Mac Aodha, Gabriel J. Brostow, Jamie Watson |
CVPR | 6 |
| 2025 | AnyCalib: On-Manifold Learning for Model-Agnostic Single-View Camera Calibration
Javier Tirado, Javier Civera 0001 |
ICCV | 2 |
| 2025 | Single-Shot Metric Depth from Focused Plenoptic CamerasabstractMetric depth estimation from visual sensors is crucial for robots to perceive, navigate, and interact with their environment. Traditional range imaging setups, such as stereo or structured light cameras, face hassles including calibration, occlusions, and hardware demands, with accuracy limited by the baseline between cameras. Single- and multi-view monocular depth offers a more compact alternative, but is constrained by the unobservability of the metric scale. Light field imaging provides a promising solution for estimating metric depth by using a unique lens configuration through a single device. However, its application to single-view dense metric depth is under-addressed mainly due to the technology's high cost, the lack of public benchmarks, and proprietary geometrical models and software. Our work explores the potential of focused plenoptic cameras for dense metric depth. We propose a novel pipeline that predicts metric depth from a single plenoptic camera shot by first generating a sparse metric point cloud using a neural network, which is then used to scale and align a dense relative depth map regressed by a foundation depth model, resulting in a dense metric depth. To validate it, we curated the Light Field & Stereo Image Dataset11Dataset available at https://zenodo.org/records/14224205. (LFS) of real-world light field images with stereo depth labels, filling a current gap in existing resources. Experimental results show that our pipeline produces accurate metric depth predictions, laying a solid groundwork for future research in this field.22Work partially supported by the DLR Impulse Project SaiNSOR. Blanca Lasheras-Hernandez, Klaus H. Strobl, Sergio Izquierdo, Tim Bodenmüller, Rudolph Triebel, Javier Civera 0001 |
ICRA | 6 |
| 2025 | VSLAM-LAB: A Comprehensive Framework for Visual SLAM Methods and DatasetsabstractVisual Simultaneous Localization and Mapping (VSLAM) research faces significant challenges due to fragmented toolchains, complex system configurations, and inconsistent evaluation methodologies. To address these issues, we present VSLAM-LAB, a unified framework designed to streamline the development, evaluation, and deployment of VSLAM systems. VSLAM-LAB simplifies the entire workflow by enabling seamless compilation and configuration of VSLAM algorithms, automated dataset downloading and preprocessing, and standardized experiment design, execution, and evaluation. All of these features are accessible through a single command-line interface. The framework supports a wide range of VSLAM systems and datasets, offering broad compatibility and extendability while promoting reproducibility through consistent evaluation metrics and analysis tools. By reducing implementation complexity and minimizing configuration overhead, VSLAM-LAB empowers researchers to focus on advancing VSLAM methodologies and accelerates progress toward scalable, real-world solutions. We demonstrate the ease with which user-relevant benchmarks can be created: here, we introduce difficulty-level-based categories, but one could envision environment-specific or condition-specific categories. Alejandro Fontán, Tobias Fischer 0001, Javier Civera 0001, Michael Milford |
IROS | 3 |
| 2025 | MRMT-PR: A Multi-Scale Reverse-View Mamba-Transformer for LiDAR Place RecognitionabstractPlace recognition is a fundamental technology of high relevance for autonomous robot navigation. Existing methods encounter significant challenges arising from scene variations (e.g., illumination changes, dynamic objects), view-point shifts, and difficulties in data fusion and alignment. These factors often lead to a substantial drop in recognition recall, which is typically addressed in the literature by training deep neural networks to learn invariant feature representations. In this paper, we propose MRMT-PR, a novel multi-scale reverse-view Mamba-Transformer architecture for LiDAR-based place recognition that uses a single-frame point cloud as its input. Our MRMT-PR framework consists of a multi-scale reverse-view preprocessing module for LiDAR point clouds, a Mamba-Transformer feature encoder, and a global feature fusion module. This architecture effectively mitigates the impact of perspective and illumination variations, enhances the global representational capacity of LiDAR features, and significantly improves recognition robustness under challenging conditions such as viewpoint changes and long-term localization. Experiments conducted on NCLT dataset with challenging scenarios demonstrate that MRMT-PR outperforms existing LiDAR-based place recognition baselines in terms of overall performance. Jingwen Wang 0009, Hongshan Yu, Yaonan Wang 0001, Javier Civera 0001, Xieyuanli Chen |
IROS | 5 |
| 2024 | Ray-Patch: An Efficient Querying for Light Field TransformersabstractIn this paper we propose the Ray-Patch querying, a novel model to efficiently query transformers to decode implicit representations into target views. Our Ray-Patch decoding reduces the computational footprint and increases inference speed up to one order of magnitude compared to previous models, without losing global attention, and hence maintaining specific task metrics. The key idea of our novel querying is to split the target image into a set of patches, then querying the transformer for each patch to extract a set of feature vectors, which are finally decoded into the target image using convolutional layers. Our experimental results, implementing Ray-Patch in 3 different architectures and evaluating it in 2 different tasks and datasets, demonstrate and quantify the effectiveness of our method, specifically a notable boost in rendering speed for the same task metrics. Tomás Berriel Martins, Javier Civera 0001 |
3DV | 2 |
| 2024 | DAC: Detector-Agnostic Spatial Covariances for Deep Local FeaturesabstractCurrent deep visual local feature detectors do not model the spatial uncertainty of detected features, producing suboptimal results in downstream applications. In this work, we propose two post-hoc covariance estimates that can be plugged into any pretrained deep feature detector: a simple, isotropic estimate that uses the predicted score at a given pixel location, and a full estimate via the local structure tensor of the learned score maps. Both methods are easy to implement and can be applied to any deep feature detector. We show that these covariances are directly related to errors in feature matching, leading to improvements in downstream tasks, including solving the perspective-n-point problem and motion-only bundle adjustment. Code is available at https://github.com/javrtg/DAC. Javier Tirado, Frederik Warburg, Javier Civera 0001 |
3DV | 3 |
| 2024 | Feature Splatting for Better Novel View Synthesis with Low Overlap
Tomás Berriel Martins, Javier Civera 0001 |
BMVC | 2 |
| 2024 | Optimal Transport Aggregation for Visual Place RecognitionabstractThe task of Visual Place Recognition (VPR) aims to match a query image against references from an extensive database of images from different places, relying solely on visual cues. State-of-the-art pipelines focus on the aggre-gation offeatures extractedfrom a deep backbone, in order to form a global descriptor for each image. In this con-text, we introduce SALAD (Sinkhorn Algorithm for Locally Aggregated Descriptors), which reformulates NetVLAD's soft-assignment of local features to clusters as an optimal transport problem. In SALAD, we consider both feature-to-cluster and cluster-to-feature relations and we also in-troduce a ‘dustbin’ cluster, designed to selectively discard features deemed non-informative, enhancing the overall de-scriptor quality. Additionally, we leverage and fine-tune DINOv2 as a backbone, which provides enhanced description power for the local features, and dramatically reduces the required training time. As a result, our single-stage method not only surpasses single-stage baselines in pub-lic VPR datasets, but also surpasses two-stage methods that add a re-ranking with significantly higher cost. Code and models are available at https://github.com/serizbalsalad. Sergio Izquierdo, Javier Civera 0001 |
CVPR | 2 |
| 2024 | From Correspondences to Pose: Non-Minimal Certifiably Optimal Relative Pose Without DisambiguationabstractEstimating the relative camera pose from n ≥ 5 correspondences between two calibrated views is a fundamental task in computer vision. This process typically involves two stages: 1) estimating the essential matrix between the views, and 2) disambiguating among the four candidate relative poses that satisfy the epipolar geometry. In this paper, we demonstrate a novel approach that, for the first time, bypasses the second stage. Specifically, we show that it is possible to directly estimate the correct relative camera pose from correspondences without needing a post-processing step to enforce the cheirality constraint on the correspondences. Building on recent advances in certifiable non-minimal optimization, we frame the relative pose estimation as a Quadratically Constrained Quadratic Program (QCQP). By applying the appropriate constraints, we ensure the estimation of a camera pose that corresponds to a valid 3D geometry and that is globally optimal when certified. We validate our method through exhaustive synthetic and real-world experiments, confirming the efficacy, efficiency and accuracy of the proposed approach. Code is available at https://github.com/javrtg/C2P. Javier Tirado, Javier Civera 0001 |
CVPR | 2 |
| 2024 | Close, But Not There: Boosting Geographic Distance Sensitivity in Visual Place Recognition
Sergio Izquierdo, Javier Civera 0001 |
ECCV (73) | 2 |
| 2024 | Adaptive Outlier Thresholding for Bundle Adjustment in Visual SLAMabstractState-of-the-art V-SLAM pipelines utilize robust cost functions and outlier rejection techniques to remove incorrect correspondences. However, these methods are typically fine-tuned to overfit certain benchmarks and struggle to adapt effectively to changes in the application domain or environmental conditions. This renders them impractical for many robotic applications in which robustness in a wide variety of conditions is essential. In this paper we introduce a novel distribution-based approach for online outlier rejection that reduces the necessity for scene-specific fine-tuning while simultaneously improving the overall SLAM performance. Through experiments across 3 different public datasets, we show that our approach consistently outperforms state-of-the-art methods in various real-world settings. Our code is available at https://github.com/alejandrofontan/ORB_SLAM2_Distribution Alejandro Fontán, Javier Civera 0001, Michael Milford |
ICRA | 2 |
| 2024 | Unifying Local and Global Multimodal Features for Place Recognition in Aliased and Low-Texture EnvironmentsabstractPerceptual aliasing and weak textures pose significant challenges to the task of place recognition, hindering the performance of Simultaneous Localization and Mapping (SLAM) systems. This paper presents a novel model, called UMF (standing for Unifying Local and Global Multimodal Features) that 1) leverages multi-modality by cross-attention blocks between vision and LiDAR features, and 2) includes a re-ranking stage that re-orders based on local feature matching the top-k candidates retrieved using a global representation. Our experiments, particularly on sequences captured on a planetary-analogous environment, show that UMF outperforms significantly previous baselines in those challenging aliased environments. Since our work aims to enhance the reliability of SLAM in all situations, we also explore its performance on the widely used RobotCar dataset, for broader applicability. Code and models are available at https://github.com/DLR-RM/UMF. Alberto García-Hernández, Riccardo Giubilato, Klaus H. Strobl, Javier Civera 0001, Rudolph Triebel |
ICRA | 4 |
| 2023 | Motion-Bias-Free Feature-Based SLAM
Alejandro Fontán, Michael Milford, Javier Civera 0001 |
BMVC | 3 |
| 2023 | SfM-TTR: Using Structure from Motion for Test-Time Refinement of Single-View Depth NetworksabstractEstimating a dense depth map from a single view is geometrically ill-posed, and state-of-the-art methods rely on learning depth's relation with visual appearance using deep neural networks. On the other hand, Structure from Motion (SfM) leverages multi-view constraints to produce very accurate but sparse maps, as matching across images is typically limited by locally discriminative texture. In this work, we combine the strengths of both approaches by proposing a novel test-time refinement (TTR) method, denoted as SfM-TTR, that boosts the performance of single-view depth networks at test time using SfM multi-view cues. Specifically, and differently from the state of the art, we use sparse SfM point clouds as test-time self-supervisory signal, fine-tuning the network encoder to learn a better representation of the test scene. Our results show how the addition of SfM-TTR to several state-of-the-art self-supervised and supervised networks improves significantly their performance, outperforming previous TTR baselines mainly based on photometric multi-view consistency. The code is available at https://github.com/serizba/SfM-TTR. Sergio Izquierdo, Javier Civera 0001 |
CVPR | 2 |
| 2023 | LightDepth: Single-View Depth Self-Supervision from Illumination DeclineabstractSingle-view depth estimation can be remarkably effective if there is enough ground-truth depth data for supervised training. However, there are scenarios, especially in medicine in the case of endoscopies, where such data cannot be obtained. In such cases, multi-view self-supervision and synthetic-to-real transfer serve as alternative approaches, however, with a considerable performance reduction in comparison to supervised case. Instead, we propose a single-view self-supervised method that achieves a performance similar to the supervised case. In some medical devices, such as endoscopes, the camera and light sources are co-located at a small distance from the target surfaces. Thus, we can exploit that, for any given albedo and surface orientation, pixel brightness is inversely proportional to the square of the distance to the surface, providing a strong single-view self-supervisory signal. In our experiments, our self-supervised models deliver accuracies comparable to those of fully supervised ones, while being applicable without depth ground-truth data. Javier Rodriguez Puigvert, Victor M. Batlle, J. M. M. Montiel, Ruben Martinez-Cantin, Pascal Fua, Juan D. Tardós, Javier Civera 0001 |
ICCV | 7 |
| 2023 | Graph-Based Global Robot Localization Informing Situational Graphs with Architectural GraphsabstractIn this paper, we propose a solution for legged robot localization using architectural plans. Our specific contributions towards this goal are several. Firstly, we develop a method for converting the plan of a building into what we denote as an architectural graph (A-Graph). When the robot starts moving in an environment, we assume it has no knowledge about it, and it estimates an online situational graph representation (S-Graph) of its surroundings. We develop a novel graph-to-graph matching method, in order to relate the S-Graph estimated online from the robot sensors and the A-Graph extracted from the building plans. Note the challenge in this, as the S-Graph may show a partial view of the full A-Graph, their nodes are heterogeneous and their reference frames are different. After the matching, both graphs are aligned and merged, resulting in what we denote as an informed Situational Graph (is-Graph), with which we achieve global robot localization and exploitation of prior knowledge from the building plans. Our experiments show that our pipeline shows a higher robustness and a significantly lower pose error than several LiDAR localization baselines. Paper Video: https://youtu.be/3Pv7y8aOsUY Muhammad Shaheer, Jose Andres Millan-Romera, Hriday Bavle, Jose Luis Sanchez-Lopez, Javier Civera 0001, Holger Voos |
IROS | 5 |
| 2023 | The Drunkard's Odometry: Estimating Camera Motion in Deforming ScenesabstractEstimating camera motion in deformable scenes poses a complex and open research challenge. Most existing non-rigid structure from motion techniques assume to observe also static scene parts besides deforming scene parts in order to establish an anchoring reference. However, this assumption does not hold true in certain relevant application cases such as endoscopies. Deformable odometry and SLAM pipelines, which tackle the most challenging scenario of exploratory trajectories, suffer from a lack of robustness and proper quantitative evaluation methodologies. To tackle this issue with a common benchmark, we introduce the Drunkard's Dataset, a challenging collection of synthetic data targeting visual navigation and reconstruction in deformable environments. This dataset is the first large set of exploratory camera trajectories with ground truth inside 3D scenes where every surface exhibits non-rigid deformations over time. Simulations in realistic 3D buildings lets us obtain a vast amount of data and ground truth labels, including camera poses, RGB images and depth, optical flow and normal maps at high resolution and quality. We further present a novel deformable odometry method, dubbed the Drunkard’s Odometry, which decomposes optical flow estimates into rigid-body camera motion and non-rigid scene deformations. In order to validate our data, our work contains an evaluation of several baselines as well as a novel tracking error metric which does not require ground truth data. Dataset and code: https://davidrecasens.github.io/TheDrunkard'sOdometry/ David Recasens, Martin R. Oswald, Marc Pollefeys, Javier Civera 0001 |
NeurIPS | 4 |
| 2022 | HARA: A Hierarchical Approach for Robust Rotation AveragingabstractWe propose a novel hierarchical approach for multiple rotation averaging, dubbed HARA. Our method incrementally initializes the rotation graph based on a hierarchy of triplet support. The key idea is to build a spanning tree by prioritizing the edges with many strong triplet supports and gradually adding those with weaker and fewer supports. This reduces the risk of adding outliers in the spanning tree. As a result, we obtain a robust initial solution that enables us to filter outliers prior to nonlinear optimization. With minimal modification, our approach can also integrate the knowledge of the number of valid 2D-2D correspondences. We perform extensive evaluations on both synthetic and real datasets, demonstrating state-of-the-art results. Seong Hun Lee, Javier Civera 0001 |
CVPR | 2 |
| 2022 | On the Uncertain Single-View Depths in Colonoscopies
Javier Rodriguez Puigvert, David Recasens, Javier Civera 0001, Ruben Martinez-Cantin |
MICCAI (3) | 3 |
| 2021 | Rotation-Only Bundle AdjustmentabstractWe propose a novel method for estimating the global rotations of the cameras independently of their positions and the scene structure. When two calibrated cameras observe five or more of the same points, their relative rotation can be recovered independently of the translation. We extend this idea to multiple views, thereby decoupling the rotation estimation from the translation and structure estimation. Our approach provides several benefits such as complete immunity to inaccurate translations and structure, and the accuracy improvement when used with rotation averaging. We perform extensive evaluations on both synthetic and real datasets, demonstrating consistent and significant gains in accuracy when used with the state-of-the-art rotation averaging method. Seong Hun Lee, Javier Civera 0001 |
CVPR | 2 |
| 2021 | Bayesian Triplet Loss: Uncertainty Quantification in Image RetrievalabstractUncertainty quantification in image retrieval is crucial for downstream decisions, yet it remains a challenging and largely unexplored problem. Current methods for estimating uncertainties are poorly calibrated, computationally expensive, or based on heuristics. We present a new method that views image embeddings as stochastic features rather than deterministic features. Our two main contributions are (1) a likelihood that matches the triplet constraint and that evaluates the probability of an anchor being closer to a positive than a negative; and (2) a prior over the feature space that justifies the conventional l2normalization. To ensure computational efficiency, we derive a variational approximation of the posterior, called the Bayesian triplet loss, that produces state-of-the-art uncertainty estimates and matches the predictive performance of current state-of-the-art methods. Frederik Warburg, Martin Jørgensen, Javier Civera 0001, Søren Hauberg |
ICCV | 3 |
| 2021 | DOT: Dynamic Object Tracking for Visual SLAMabstractIn this paper we present DOT (Dynamic Object Tracking), a front-end that added to existing SLAM systems can significantly improve their robustness and accuracy in highly dynamic environments. DOT combines instance segmentation and multi-view geometry to generate masks for dynamic objects in order to allow SLAM systems based on rigid scene models to avoid such image areas in their optimizations.To determine which objects are actually moving, DOT segments first instances of potentially dynamic objects and then, with the estimated camera motion, tracks such objects by minimizing the photometric reprojection error. This short-term tracking improves the accuracy of the segmentation with respect to other approaches. In the end, only actually dynamicmasks are generated.We have evaluated DOT with ORB-SLAM 2 [1] in three public datasets. Our results show that our approach improves significantly the accuracy and robustness of ORB-SLAM2, especially in highly dynamic scenes. Irene Ballester, Alejandro Fontán, Javier Civera 0001, Klaus H. Strobl, Rudolph Triebel |
ICRA | 3 |
| 2020 | Information-Driven Direct RGB-D OdometryabstractThis paper presents an information-theoretic approach to point selection in direct RGB-D odometry. The aim is to select only the most informative measurements, in order to reduce the optimization problem with a minimal impact in the accuracy. It is usual practice in visual odometry/SLAM to track several hundreds of points, achieving real-time performance in high-end desktop PCs. Reducing their computational footprint will facilitate the implementation of odometry and SLAM in low-end platforms such as small robots and AR/VR glasses. Our experimental results show that our novel information-based selection criterion allows us to reduce the number of tracked points an order of magnitude (down to only 24 of them), achieving an accuracy similar to the state of the art (sometimes outperforming it) while reducing 10 times the computational demand. Alejandro Fontán, Javier Civera 0001, Rudolph Triebel |
CVPR | 2 |
| 2020 | Mapillary Street-Level Sequences: A Dataset for Lifelong Place RecognitionabstractLifelong place recognition is an essential and challenging task in computer vision with vast applications in robust localization and efficient large-scale 3D reconstruction. Progress is currently hindered by a lack of large, diverse, publicly available datasets. We contribute with Mapillary Street-Level Sequences (SLS), a large dataset for urban and suburban place recognition from image sequences. It contains more than 1.6 million images curated from the Mapillary collaborative mapping platform. The dataset is orders of magnitude larger than current data sources, and is designed to reflect the diversities of true lifelong learning. It features images from 30 major cities across six continents, hundreds of distinct cameras, and substantially different viewpoints and capture times, spanning all seasons over a nine year period. All images are geo-located with GPS and compass, and feature high-level attributes such as road type. We propose a set of benchmark tasks designed to push state-of-the-art performance and provide baseline studies. We show that current state-of-the-art methods still have a long way to go, and that the lack of diversity in existing datasets have prevented generalization to new environments. The dataset and benchmarks are available for academic research. Frederik Warburg, Søren Hauberg, Manuel Lopez-Antequera, Pau Gargallo, Yubin Kuang, Javier Civera 0001 |
CVPR | 6 |
| 2020 | From Points to Planes - Adding Planar Constraints to Monocular SLAM Factor GraphsabstractPlanar structures are common in man-made environments. Their addition to monocular SLAM algorithms is of relevance in order to achieve more complete and higher- level scene representations. Also, the additional constraints they introduce might reduce the estimation errors in certain situations. In this paper we present a novel formulation to incorporate plane landmarks and planar constraints to feature- based monocular SLAM. Specifically, we enforce in-plane points to lie exactly in the plane they belong to, propagating such information to the rest of the states. Our formulation, differently from the state of the art, allows us to incorporate general planes, independently of depth information or CNN segmentation being available (although we could also use them). We evaluate our method in several sequences of public databases, showing accurate plane estimations and pose accuracy on par with state- of-the-art point-only monocular SLAM. Charlotte Arndt, Reza Sabzevari, Javier Civera 0001 |
IROS | 3 |
| 2020 | Incremental Learning of Object Models From Natural Human-Robot InteractionsabstractIn order to perform complex tasks in realistic human environments, robots need to be able to learn new concepts in the wild, incrementally, and through their interactions with humans. This article presents an end-to-end pipeline to learn object models incrementally during the human-robot interaction (HRI). The pipeline we propose consists of three parts: 1) recognizing the interaction type; 2) detecting the object that the interaction is targeting; and 3) learning incrementally the models from data recorded by the robot sensors. Our main contributions lie in the target object detection, guided by the recognized interaction, and in the incremental object learning. The novelty of our approach is the focus on natural, heterogeneous, and multimodal HRIs to incrementally learn new object models. Throughout the article, we highlight the main challenges associated with this problem, such as high degree of occlusion and clutter, domain change, low-resolution data, and interaction ambiguity. This article shows the benefits of using multiview approaches and combining visual and language features, and our experimental results outperform standard baselines. Pablo Azagra, Javier Civera 0001, Ana Cristina Murillo |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2019 | Triangulation: Why Optimize?
Seong Hun Lee, Javier Civera 0001 |
BMVC | 2 |
| 2019 | CAM-Convs: Camera-Aware Multi-Scale Convolutions for Single-View DepthabstractSingle-view depth estimation suffers from the problem that a network trained on images from one camera does not generalize to images taken with a different camera model. Thus, changing the camera model requires collecting an entirely new training dataset. In this work, we propose a new type of convolution that can take the camera parameters into account, thus allowing neural networks to learn calibration-aware patterns. Experiments confirm that this improves the generalization capabilities of depth prediction networks considerably, and clearly outperforms the state of the art when the train and test images are acquired with different cameras. José M. Fácil, Benjamin Ummenhofer, Huizhong Zhou, Luis Montesano, Thomas Brox, Javier Civera 0001 |
CVPR | 6 |
| 2019 | Closed-Form Optimal Two-View Triangulation Based on Angular ErrorsabstractIn this paper, we study closed-form optimal solutions to two-view triangulation with known internal calibration and pose. By formulating the triangulation problem as L1and L∞minimization of angular reprojection errors, we derive the exact closed-form solutions that guarantee global optimality under respective cost functions. To the best of our knowledge, we are the first to present such solutions. Since the angular error is rotationally invariant, our solutions can be applied for any type of central cameras, be it perspective, fisheye or omnidirectional. Our methods also require significantly less computation than the existing optimal methods. Experimental results on synthetic and real datasets validate our theoretical derivations. Seong Hun Lee, Javier Civera 0001 |
ICCV | 2 |
| 2018 | Visual-Inertial SLAM Initialization: A General Linear Formulation and a Gravity-Observing Non-Linear OptimizationabstractThe initialization is one of the less reliable pieces of Visual-Inertial SLAM (VI-SLAM) and Odometry (VI-O). The estimation of the initial state (camera poses, IMU states and landmark positions) from the first data readings lacks the accuracy and robustness of other parts of the pipeline, and most algorithms have high failure rates and/or initialization delays up to tens of seconds. Such initialization is critical for AR systems, as the failures and delays of the current approaches can ruin the user experience or mandate impractical guided calibration. In this paper we address the state initialization problem using a monocular-inertial sensor setup, the most common in AR platforms. Our contributions are 1) a general linear formulation to obtain an initialization seed, and 2) a non-linear optimization scheme, including gravity, to refine the seed. Our experimental results, in a public dataset, show that our approach improves the accuracy and robustness of current VI state initialization schemes. Javier Dominguez-Conti, Jianfeng Yin, Yacine Alami, Javier Civera 0001 |
ISMAR | 4 |
| 2017 | A multimodal dataset for object model learning from natural human-robot interactionabstractLearning object models in the wild from natural human interactions is an essential ability for robots to perform general tasks. In this paper we present a robocentric multimodal dataset addressing this key challenge. Our dataset focuses on interactions where the user teaches new objects to the robot in various ways. It contains synchronized recordings of visual (3 cameras) and audio data which provide a challenging evaluation framework for different tasks. Additionally, we present an end-to-end system that learns object models using object patches extracted from the recorded natural interactions. Our proposed pipeline follows these steps: (a) recognizing the interaction type, (b) detecting the object that the interaction is focusing on, and (c) learning the models from the extracted data. Our main contribution lies in the steps towards identifying the target object patches of the images. We demonstrate the advantages of combining language and visual features for the interaction recognition and use multiple views to improve the object modelling. Our experimental results show that our dataset is challenging due to occlusions and domain change with respect to typical object learning frameworks. The performance of common out-of-the-box classifiers trained on our data is low. We demonstrate that our algorithm outperforms such baselines. Pablo Azagra, Florian Golemo, Yoan Mollard, Manuel Lopes 0001, Javier Civera 0001, Ana Cristina Murillo |
IROS | 5 |
| 2017 | RGBDTAM: A cost-effective and accurate RGB-D tracking and mapping systemabstractSimultaneous Localization and Mapping using RGB-D cameras has been a fertile research topic in the latest decade, due to the suitability of such sensors for indoor robotics. In this paper we propose a direct RGB-D SLAM algorithm with state-of-the-art accuracy and robustness at a los cost. Our experiments in the RGB-D TUM dataset [34] effectively show a better accuracy and robustness in CPU real time than direct RGB-D SLAM systems that make use of the GPU. The key ingredients of our approach are mainly two. Firstly, the combination of a semi-dense photometric and dense geometric error for the pose tracking (see Figure 1), which we demonstrate to be the most accurate alternative. And secondly, a model of the multi-view constraints and their errors in the mapping and tracking threads, which adds extra information over other approaches. We release the open-source implementation of our approach1. The reader is referred to a video with our results2for a more illustrative visualization of its performance. Alejo Concha, Javier Civera 0001 |
IROS | 2 |
| 2017 | Guest Editorial Special Issue on Wearable and Ego-Vision Systems for Augmented ExperienceabstractThe papers in this special section focus on the deployment of wearable computing technologies and ego-vision systems for augmented reality applications. Rapid progress in the development of low-level component technologies such as wearable sensors, wearable displays, and wearable computers is making our digital lives grow, connect, and play a relevant role in reality. To name a few examples, body-mounted sensors and displays help athletes in training by presenting real-time performance metrics such as speed, distance, and heart rate.Wearable systems allow medical staff in hospitals to consult specialists located anywhere in the world, in real time, providing optimal patient care. And within the context of assistive technologies, a head-mounted camera can be used to identify and convey the presence of objects, people, or text to a visually impaired user. Giuseppe Serra 0001, Rita Cucchiara, Kris Makoto Kitani, Javier Civera 0001 |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2016 | Visual-inertial direct SLAMabstractThe so-called direct visual SLAM methods have shown a great potential in estimating a semidense or fully dense reconstruction of the scene, in contrast to the sparse reconstructions of the traditional feature-based algorithms. In this paper, we propose for the first time a direct, tightly-coupled formulation for the combination of visual and inertial data. Our algorithm runs in real-time on a standard CPU. The processing is split in three threads. The first thread runs at frame rate and estimates the camera motion by a joint non-linear optimization from visual and inertial data given a semidense map. The second one creates a semidense map of high-gradient areas only for camera tracking purposes. Finally, the third thread estimates a fully dense reconstruction of the scene at a lower frame rate. We have evaluated our algorithm in several real sequences with ground truth trajectory data, showing a state-of-the-art performance. Alejo Concha, Giuseppe Loianno, Vijay Kumar 0001, Javier Civera 0001 |
ICRA | 4 |
| 2016 | Dealing with small data and training blind spots in the Manhattan worldabstractLeveraging Manhattan assumption we generate metrically rectified novel views from a single image, even for non-box scenarios. Our novel views enable the already trained classifiers to handle training data missing views (blind spots) without additional training. We demonstrate this on end-to-end scene text spotting under perspective. Additionally, utilizing our fronto-parallel views, we discover unsuspended invariant mid-level patches given a few widely separated training examples (small data domain). These invariant patches outperform various baselines on small data image retrieval challenge. Muhammad Wajahat Hussain, Javier Civera 0001, Luis Montano, Martial Hebert |
WACV | 2 |
| 2015 | DPPTAM: Dense piecewise planar tracking and mapping from a monocular sequenceabstractThis paper proposes a direct monocular SLAM algorithm that estimates a dense reconstruction of a scene in real-time on a CPU. Highly textured image areas are mapped using standard direct mapping techniques [1], that minimize the photometric error across different views. We make the assumption that homogeneous-color regions belong to approximately planar areas. Our contribution is a new algorithm for the estimation of such planar areas, based on the information of a superpixel segmentation and the semidense map from highly textured areas. We compare our approach against several alternatives using the public TUM dataset [2] and additional live experiments with a hand-held camera. We demonstrate that our proposal for piecewise planar monocular SLAM is faster, more accurate and more robust than the piecewise planar baseline [3]. In addition, our experimental results show how the depth regularization of monocular maps can damage its accuracy, being the piecewise planar assumption a reasonable option in indoor scenarios. Alejo Concha, Javier Civera 0001 |
IROS | 2 |
| 2015 | Stereo parallel tracking and mapping for robot localizationabstractThis paper describes a visual SLAM system based on stereo cameras and focused on real-time localization for mobile robots. To achieve this, it heavily exploits the parallel nature of the SLAM problem, separating the time-constrained pose estimation from less pressing matters such as map building and refinement tasks. On the other hand, the stereo setting allows to reconstruct a metric 3D map for each frame of stereo images, improving the accuracy of the mapping process with respect to monocular SLAM and avoiding the well-known bootstrapping problem. Also, the real scale of the environment is an essential feature for robots which have to interact with their surrounding workspace. A series of experiments, on-line on a robot as well as off-line with public datasets, are performed to validate the accuracy and real-time performance of the developed method. Taihú Pire, Thomas Fischer 0006, Javier Civera 0001, Pablo de Cristóforis, Julio Jacobo-Berlles |
IROS | 3 |
| 2015 | Layout aware visual tracking and mappingabstractNowadays real time visual Simultaneous Localization And Mapping (SLAM) algorithms exist and rely on consistent measurements across multiple views. In indoor environments, where majority of robot's activity takes place, severe occlusions can occur, e.g., when turning around a corner or moving from one room to another. In these situations, SLAM algorithms can not establish correspondences across views, which leads to failures in camera localization or map construction. This work takes advantage of the recent scene box layout descriptor to make the above mentioned SLAM systems occlusion aware. This room box reasoning helps the sequential tracker to reason about possible occlusions and therefore look for matches in only potentially visible features instead of the entire map. This increases the life of the tracker, as it does not consider itself lost under the occlusion state. Additionally, focusing on the potentially visible portion of the map, i.e., the current room features, it improves the computational efficiency without compromising the accuracy. Finally, this room level reasoning helps in better image selection for bundle adjustment. The image bundle coming from the same room has little occlusion, which leads to better dense reconstruction. We demonstrate the superior performance of layout aware SLAM on several long monocular sequences acquired in difficult indoor situations, specifically in a room-room transition and turning around a corner. Marta Salas, Muhammad Wajahat Hussain, Alejo Concha, Luis Montano, Javier Civera 0001, J. M. M. Montiel |
IROS | 5 |
| 2015 | Guest Editorial Special Issue on Cloud Robotics and AutomationabstractThe articles in this special section focus on the use of cloud computing in the robotics industry. The Internet and the availability of vast computational resources, ever-growing data and storage capacity have the potential to define a new paradigm for robotics and automation. An intelligent system connected to the Internet can expand its onboard local data, computation and sensors with huge data repositories from similar and very different domains, massive parallel computation from server farms and sensor/actuator streams from other robots and automata. It is the potential and also the research challenges of the field that become the focus on this special section. The goal is to group together and to show the state-of-the-art of this newly emerged field, identify the relevant advances and topics, point out the current lines of research and potential applications, and discuss the main research challenges and future work directions. Javier Civera 0001, Matei T. Ciocarlie, Alper Aydemir, Kostas E. Bekris, Sanjay E. Sarma |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2014 | Grounding Acoustic Echoes in Single View Geometry EstimationabstractExtracting the 3D geometry plays an important part in scene understanding. Recently, robust visual descriptors are proposed for extracting the indoor scene layout from a passive agent’s perspective, specifically from a single image. Their robustness is mainly due to modelling the physical interaction of the underlying room geometry with the objects and the humans present in the room. In this work we add the physical constraints coming from acoustic echoes, generated by an audio source, to this visual model. Our audio-visual 3D geometry descriptor improves over the state of the art in passive perception models as we show in our experiments. Muhammad Wajahat Hussain, Javier Civera 0001, Luis Montano |
AAAI | 2 |
| 2014 | Using superpixels in monocular SLAMabstractMonocular SLAM and Structure from Motion have been traditionally based on finding point correspondences in highly-textured image areas. Large textureless regions, usually found in indoor and urban environments, are difficult to reconstruct by these systems. In this paper we augment for the first time the traditional point-based monocular SLAM maps with superpixels. Superpixels are middle-level features consisting of image regions of homogeneous texture. We propose a novel scheme for superpixel matching, 3D initialization and optimization that overcomes the difficulties of salient point-based approaches in these areas of homogeneous texture. Our experimental results show the validity of our approach. First, we compare our proposal with a state-of-the-art multiview stereo system; being able to reconstruct the textureless regions that the latest cannot. Secondly, we present experimental results of our algorithm integrated with the point-based PTAM [1]; estimating, now in real-time, the superpixel textureless areas. Finally, we show the accuracy of the presented algorithm with a quantitative analysis of the estimation error. Alejo Concha, Javier Civera 0001 |
ICRA | 2 |
| 2012 | Learning object class detectors from weakly annotated videoabstractObject detectors are typically trained on a large set of still images annotated by bounding-boxes. This paper introduces an approach for learning object detectors from real-world web videos known only to contain objects of a target class. We propose a fully automatic pipeline that localizes objects in a set of videos of the class and learns a detector for it. The approach extracts candidate spatio-temporal tubes based on motion segmentation and then selects one tube per video jointly over all videos. To compare to the state of the art, we test our detector on still images, i.e., Pascal VOC 2007. We observe that frames extracted from web videos can differ significantly in terms of quality to still images taken by a good camera. Thus, we formulate the learning from videos as a domain adaptation task. We show that training from a combination of weakly annotated videos and fully annotated still images using domain adaptation improves the performance of a detector trained from still images alone. Alessandro Prest, Christian Leistner, Javier Civera 0001, Cordelia Schmid, Vittorio Ferrari |
CVPR | 3 |
| 2012 | Creating and using RoboEarth object modelsabstractThis paper presented an approach to create 3D object models for robotic and vision applications in a fast and inexpensive way compared to established approaches. By using the RoboEarth system for storing the created object models users have world-wide access to the data and can immediately reuse a model as soon as it was created and uploaded. The approach shows general applicability for different kinds of cameras. In this work this was shown by two example implementations for the recognition process of objects. The quality of the recognition can be verified in the video. Combined with the knowledge saved in the RoboEarth database the objects can also be properly classified. Daniel Di Marco, Andreas Koch 0003, Oliver Zweigle, Kai Häussermann, Björn Schießle, Paul Levi, Dorian Gálvez-López, Luis Riazuelo, Javier Civera 0001, J. M. M. Montiel, Moritz Tenorth, Alexander Clifford Perzylo, Markus Waibel, René van de Molengraft |
ICRA | 9 |
| 2012 | Impact of Landmark Parametrization on Monocular EKF-SLAM with Points and Lines
Joan Solà, Teresa Vidal-Calleja, Javier Civera 0001, J. M. M. Montiel |
Int. J. Comput. Vis. | 3 |
| 2011 | EKF monocular SLAM with relocalization for laparoscopic sequencesabstractIn recent years, research on visual SLAM has produced robust algorithms providing, in real time at 30 Hz, both the 3D model of the observed rigid scene and the 3D camera motion using as only input the gathered image sequence. These algorithms have been extensively validated in rigid human-made environments -indoor and outdoor- showing robust performance in dealing with clutter, occlusions or sudden motions. Medical endoscopic sequences naturally pose a monocular SLAM problem: an unknown camera motion in an unknown environment. The corresponding map would be useful in providing 3D information to assist surgeons, to support augmented reality insertions or to be exploited by medical robots. In this paper we propose the combination EKF Monocular SLAM + 1-Point RANSAC + Randomised List Relocalization to process laparoscopic sequences -abdominal cavity images-. The sequences are challenging due to: 1) cluttering produced by tools; 2) sudden motions of the camera; 3) laparoscope frequently goes in and out of abdominal cavity; 4) tissue deformation caused by respiration, heartbeats and/or surgical tools. Real medical image sequences provide experimental validation. Oscar G. Grasa, Javier Civera 0001, J. M. M. Montiel |
ICRA | 2 |
| 2011 | Dense multi-planar scene estimation from a sparse set of imagesabstractEgo-motion estimation and 3D scene reconstruction from image data has been a long term aim both in the Robotics and Computer Vision communities. Nevertheless, while both visual SLAM and Structure from Motion already provide an accurate ego-motion estimation, visual scene estimation does not offer yet such a satisfactory result; being in most cases limited to a sparse set of salient points. In this paper we propose an algorithm to densify a sparse point-based reconstruction into a dense multi-plane based one, from the only input of a set of sparse images. Alberto Argiles, Javier Civera 0001, Luis Montesano |
IROS | 2 |
| 2011 | Towards semantic SLAM using a monocular cameraabstractMonocular SLAM systems have been mainly focused on producing geometric maps just composed of points or edges; but without any associated meaning or semantic content. In this paper, we propose a semantic SLAM algorithm that merges in the estimated map traditional meaningless points with known objects. The non-annotated map is built using only the information extracted from a monocular image sequence. The known object models are automatically computed from a sparse set of images gathered by cameras that may be different from the SLAM camera. The models include both visual appearance and tridimensional information. The semantic or annotated part of the map -the objects- are estimated using the information in the image sequence and the precomputed object models. The proposed algorithm runs an EKF monocular SLAM parallel to an object recognition thread. This latest one informs of the presence of an object in the sequence by searching for SURF correspondences and checking afterwards their geometric compatibility. When an object is recognized it is inserted in the SLAM map, being its position measured and hence refined by the SLAM algorithm in subsequent frames. Experimental results show real-time performance for a hand held camera imaging a desktop environment and for a camera mounted in a robot moving in a room-sized scenario. Javier Civera 0001, Dorian Gálvez-López, Luis Riazuelo, Juan D. Tardós, J. M. M. Montiel |
IROS | 1 |
| 2009 | Camera self-calibration for sequential Bayesian structure from motionabstractComputer vision researchers have proved the feasibility of camera self-calibration —the estimation of a camera's internal parameters from an image sequence without any known scene structure. Various self-calibration algorithms have been published. Nevertheless, all of the recent sequential approaches to 3D structure and motion estimation from image sequences which have arisen in robotics and aim at real-time operation (often classed as visual SLAM or visual odometry) have relied on pre-calibrated cameras and have not attempted online calibration. Javier Civera 0001, Diana R. Bueno, Andrew J. Davison, J. M. M. Montiel |
ICRA | 1 |
| 2009 | 1-point RANSAC for EKF-based Structure from MotionabstractRecently, classical pairwise Structure From Motion (SfM) techniques have been combined with non-linear global optimization (Bundle Adjustment, BA) over a sliding window to recursively provide camera pose and feature location estimation from long image sequences. Normally called Visual Odometry, these algorithms are nowadays able to estimate with impressive accuracy trajectories of hundreds of meters; either from an image sequence (usually stereo) as the only input, or combining visual and propioceptive information from inertial sensors or wheel odometry. This paper has a double objective. First, we aim to illustrate for the first time how similar accuracy and trajectory length can be achieved by filtering-based visual SLAM methods. Specifically, a camera-centered Extended Kalman Filter is used here to process a monocular sequence as the only input, with 6DOF motion estimated. Features are kept live in the filter while visible as the camera explores forward, and are deleted from the state once they go out of view. This permits an increase in the number of tracked features per frame from tens to around a hundred. While improving the accuracy of the estimation, it makes computationally infeasible the exhaustive Branch and Bound search performed by standard JCBB for match outlier rejection. As a second contribution that overcomes this problem, we present here a RANSAC-like algorithm that exploits the probabilistic prediction of the filter. This use of prior information makes it possible to reduce the size of the minimal data subset to instantiate a hypothesis to the minimum possible of 1 point, greatly increasing the efficiency of the outlier rejection stage. Experimental results from real image sequences covering trajectories of hundreds of meters are presented and compared against RTK GPS ground truth. Estimation errors are about 1% of the trajectory for trajectories up to 650 metres. Javier Civera 0001, Oscar G. Grasa, Andrew J. Davison, J. M. M. Montiel |
IROS | 1 |
| 2009 | Drift-Free Real-Time Sequential Mosaicing
Javier Civera 0001, Andrew J. Davison, Juan A. Magallon, J. M. M. Montiel |
Int. J. Comput. Vis. | 1 |
| 2008 | Interacting multiple model monocular SLAMabstractRecent work has demonstrated the benefits of adopting a fully probabilistic SLAM approach in sequential motion and structure estimation from an image sequence. Unlike standard Structure from Motion (SFM) methods, this 'monocular SLAM' approach is able to achieve drift-free estimation with high frame-rate real-time operation, particularly benefitting from highly efficient active feature search, map management and mismatch rejection. A consistent thread in this research on real-time monocular SLAM has been to reduce the assumptions required. In this paper we move towards the logical conclusion of this direction by implementing a fully Bayesian Interacting Multiple Models (IMM) framework which can switch automatically between parameter sets in a dimensionless formulation of monocular SLAM. Remarkably, our approach of full sequential probability propagation means that there is no need for penalty terms to achieve the Occam property of favouring simpler models - this arises automatically. We successfully tackle the known stiffness in on-the-fly monocular SLAM start up without known patterns in the scene. The search regions for matches are also reduced in size with respect to single model EKF increasing the rejection of spurious matches. We demonstrate our method with results on a complex real image sequence with varied motion. Javier Civera 0001, Andrew J. Davison, J. M. M. Montiel |
ICRA | 1 |
| 2008 | Inverse Depth Parametrization for Monocular SLAMabstractWe present a new parametrization for point features within monocular simultaneous localization and mapping (SLAM) that permits efficient and accurate representation of uncertainty during undelayed initialization and beyond, all within the standard extended Kalman filter (EKF). The key concept is direct parametrization of the inverse depth of features relative to the camera locations from which they were first viewed, which produces measurement equations with a high degree of linearity. Importantly, our parametrization can cope with features over a huge range of depths, even those that are so far from the camera that they present little parallax during motion---maintaining sufficient representative uncertainty that these points retain the opportunity to "come in'' smoothly from infinity if the camera makes larger movements. Feature initialization is undelayed in the sense that even distant features are immediately used to improve camera motion estimates, acting initially as bearing references but not permanently labeled as such. The inverse depth parametrization remains well behaved for features at all stages of SLAM processing, but has the drawback in computational terms that each point is represented by a 6-D state vector as opposed to the standard three of a EuclideanXYZrepresentation. We show that once the depth estimate of a feature is sufficiently accurate, its representation can safely be converted to the EuclideanXYZform, and propose a linearity index that allows automatic detection and conversion to maintain maximum efficiency---only low parallax features need be maintained in inverse depth form for long periods. We present a real-time implementation at 30 Hz, where the parametrization is validated in a fully automatic 3-D SLAM system featuring a handheld single camera with no additional sensing. Experiments show robust operation in challenging indoor and outdoor environments with a very large ranges of scene depth, varied motion, and also real time 360degloop closing. Javier Civera 0001, Andrew J. Davison, J. M. M. Montiel |
IEEE Trans. Robotics | 1 |
| 2007 | Inverse Depth to Depth Conversion for Monocular SLAMabstractRecently it has been shown that an inverse depth parametrization can improve the performance of real-time monocular EKF SLAM, permitting undelayed initialization of features at all depths. However, the inverse depth parametrization requires the storage of 6 parameters in the state vector for each map point. This implies a noticeable computing overhead when compared with the standard 3 parameter XYZ Euclidean encoding of a 3D point, since the computational complexity of the EKF scales poorly with state vector size. In this work we propose to restrict the inverse depth parametrization only to cases where the standard Euclidean encoding implies a departure from linearity in the measurement equations. Every new map feature is still initialized using the 6 parameter inverse depth method. However, as the estimation evolves, if according to a linearity index the alternative XYZ coding can be considered linear, we show that feature parametrization can be transformed from inverse depth to XYZ for increased computational efficiency with little reduction in accuracy. We present a theoretical development of the necessary linearity indices, along with simulations to analyze the influence of the conversion threshold. Experiments performed with with a 30 frames per second real-time system are reported. An analysis of the increase in the map size that can be successfully managed is included. Javier Civera 0001, Andrew J. Davison, J. M. M. Montiel |
ICRA | 1 |