Cesar Dario Cadena Lerma

dblp:43/7748 · also Cesar Cadena 0001, César Cadena 0001 · DBLP profile ↗
← Back
55ranked-venue papers
8as first author
25since 2021 · last 2026
0000-0002-2972-6011ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 50 · 6 first-author · 22 since 2021Systems, architecture and hardware · 37 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Real-time frame- and event-based object detection with spiking neural networks on edge neuromorphic hardware: Design, deployment and benchmark
Udayanga G. W. K. N. Gamage, Cesar Dario Cadena Lerma, Matteo Fumagalli 0001, Silvia Tolu
Neurocomputing3
2025 ForestLPR: LiDAR Place Recognition in Forests Attentioning Multiple BEV Density Images
abstract
Place recognition is essential to maintain global consistency in large-scale localization systems. While research in urban environments has progressed significantly using LiDARs or cameras, applications in natural forest-like environments remain largely under-explored. Furthermore, forests present particular challenges due to high self-similarity and substantial variations in vegetation growth over time. In this work, we propose a robust LiDAR-based place recognition method for natural forests, ForestLPR. We hypothesize that a set of cross-sectional images of the forest’s geometry at different heights contains the information needed to recognize revisiting a place. The cross-sectional images are represented by bird’s-eye view (BEV) density images of horizontal slices of the point cloud at different heights. Our approach utilizes a visual transformer as the shared backbone to produce sets of local descriptors and introduces a multi-BEV interaction module to attend to information at different heights adaptively. It is followed by an aggregation layer that produces a rotation-invariant place descriptor. We evaluated the efficacy of our method extensively on real-world data from public benchmarks as well as robotic datasets and compared it against the state-of-the-art (SOTA) methods. The results indicate that ForestLPR has consistently good performance on all evaluations and achieves an average increase of 7.38% and 9.11% on Recall@1 over the closest competitor on intra-sequence loop closure detection and inter-sequence re-localization, respectively, validating our hypothesis1.
Yanqing Shen, Turcan Tuna, Marco Hutter 0001, Cesar Dario Cadena Lerma, Nanning Zheng 0001
CVPR4
2025 TartanGround: A Large-Scale Dataset for Ground Robot Perception and Navigation
abstract
We present TartanGround, a large-scale, multi-modal dataset to advance the perception and autonomy of ground robots operating in diverse environments. This dataset, collected in various photorealistic simulation environments includes multiple RGB stereo cameras for 360-degree coverage, along with depth, optical flow, stereo disparity, LiDAR point clouds, ground truth poses, semantic segmented images, and occupancy maps with semantic labels. Data is collected using an integrated automatic pipeline, which generates trajectories mimicking the motion patterns of various ground robot platforms, including wheeled and legged robots. We collect 878 trajectories across 63 environments, resulting in 1.44 million samples. Evaluations on occupancy prediction and SLAM tasks reveal that state-of-the-art methods trained on existing datasets struggle to generalize across diverse scenes. TartanGround can serve as a testbed for training and evaluation of a broad range of learning-based tasks, including occupancy prediction, SLAM, neural scene representation, perception-based navigation, and more, enabling advancements in robotic perception and autonomy towards achieving robust models generalizable to more diverse scenarios. The dataset and codebase are available on the webpage: https://tartanair.org/tartanground
Manthan Patel, Fan Yang 0092, Yuheng Qiu, Cesar Dario Cadena Lerma, Sebastian A. Scherer, Marco Hutter 0001
IROS4
2025 Continuous-Time State Estimation Methods in Robotics: A Survey
abstract
Accurate, efficient, and robust state estimation is more important than ever in robotics as the variety of platforms and complexity of tasks continue to grow. Historically, discrete-time filters and smoothers have been the dominant approach, in which the estimated variables are states at discrete sample times. The paradigm of continuous-time state estimation proposes an alternative strategy by estimating variables that express the state as a continuous function of time, which can be evaluated at any query time. Not only can this benefit downstream tasks such as planning and control, but it also significantly increases estimator performance and flexibility, as well as reduces sensor preprocessing and interfacing complexity. Despite this, continuous-time methods remain underutilized, potentially because they are less well-known within robotics. To remedy this, this work presents a unifying formulation of these methods and the most exhaustive literature review to date, systematically categorizing prior work by methodology, application, state variables, historical context, and theoretical contribution to the field. By surveying splines and Gaussian process together and contextualizing works from other research domains, this work identifies and analyzes open problems in continuous-time state estimation and suggests new research directions.
William Talbot, Julian Nubert, Turcan Tuna, Cesar Dario Cadena Lerma, Frederike Dümbgen, Jesus Tordesillas, Tim D. Barfoot, Marco Hutter 0001
IEEE Trans. Robotics4
2024 COIN-LIO: Complementary Intensity-Augmented LiDAR Inertial Odometry
abstract
We present COIN-LIO, a LiDAR Inertial Odometry pipeline that tightly couples information from LiDAR intensity with geometry-based point cloud registration. The focus of our work is to improve the robustness of LiDAR-inertial odometry in geometrically degenerate scenarios, like tunnels or flat fields. We project LiDAR intensity returns into an image, and present a novel image processing pipeline that produces filtered images with improved brightness consistency within the image as well as across different scenes. We effectively leverage intensity as an additional modality, using our new feature selection scheme that detects uninformative directions in the point cloud registration and explicitly selects patches with complementary image information. Photometric error minimization in the image patches is then fused with inertial measurements and point-to-plane registration in an iterated Extended Kalman Filter. The proposed approach improves accuracy and robustness on a public dataset. We additionally publish a new dataset, that captures five real-world environments in challenging, geometrically degenerate scenes. By using the additional photometric information, our approach shows drastically improved robustness against geometric degeneracy in environments where all compared baseline approaches fail.
Patrick Pfreundschuh, Helen Oleynikova, Cesar Dario Cadena Lerma, Roland Siegwart, Olov Andersson
ICRA3
2024 Resilient Legged Local Navigation: Learning to Traverse with Compromised Perception End-to-End
abstract
Autonomous robots must navigate reliably in unknown environments even under compromised exteroceptive perception, or perception failures. Such failures often occur when harsh environments lead to degraded sensing, or when the perception algorithm misinterprets the scene due to limited generalization. In this paper, we model perception failures as invisible obstacles and pits, and train a reinforcement learning (RL) based local navigation policy to guide our legged robot. Unlike previous works relying on heuristics and anomaly detection to update navigational information, we train our navigation policy to reconstruct the environment information in the latent space from corrupted perception and react to perception failures end-to-end. To this end, we incorporate both proprioception and exteroception into our policy inputs, thereby enabling the policy to sense collisions on different body parts and pits, prompting corresponding reactions. We validate our approach in simulation and on the real quadruped robot ANYmal running in real-time (<10ms CPU inference). In a quantitative comparison with existing heuristic-based locally reactive planners, our policy increases the success rate over 30% when facing perception failures. Project Page: https://bit.ly/45NBTuh.
Jonas Frey, Nikita Rudin, Matías Mattamala, Cesar Dario Cadena Lerma, Marco Hutter 0001
ICRA6
2024 Temporal- and Viewpoint-Invariant Registration for Under-Canopy Footage using Deep-Learning-based Bird's-Eye View Prediction
abstract
Conducting visual assessments under the canopy using mobile robots is an emerging task in smart farming and forestry. However, it is challenging to register images across different data-collection days, especially across seasons, due to the self-occluding geometry and temporal dynamics in forests and orchards. This paper proposes a new approach for registering under-canopy image sequences in general and in these situations. Our methodology leverages standard GPS data and deep-learning-based perspective to bird’s-eye view conversion to provide an initial estimation of the positions of the trees in images and their association across datasets. Furthermore, it introduces an innovative strategy for extracting tree trunks and clean ground surfaces from noisy and sparse 3D reconstructions created from the image sequences, utilizing these features to achieve precise alignment. Our robust alignment method effectively mitigates position and scale drift, which may arise from GPS inaccuracies and Sparse Structure from Motion (SfM) limitations. We evaluate our approach on three challenging real-world datasets, demonstrating that our method outperforms ICP-based methods on average by 50%, and surpasses FGR and TEASER++ by over 90% in alignment accuracy. These results highlight our method’s cost efficiency and robustness, even in the presence of severe outliers and sparsity. https://github.com/VIS4ROB-lab/bev_undercanopy_registration
Jiawei Zhou 0004, Ruben Mascaro, Cesar Dario Cadena Lerma, Margarita Chli, Lucas Teixeira
IROS3
2024 VIRUS-NeRF - Vision, InfraRed and UltraSonic based Neural Radiance Fields
abstract
Autonomous mobile robots are an increasingly integral part of modern factory and warehouse operations. Obstacle detection, avoidance and path planning are critical safety-relevant tasks, which are often solved using expensive LiDAR sensors and depth cameras. We propose to use cost-effective low-resolution ranging sensors, such as ultrasonic and infrared time-of-flight sensors by developing VIRUS-NeRF - Vision, InfraRed, and UltraSonic based Neural Radiance Fields.Building upon Instant Neural Graphics Primitives with a Multiresolution Hash Encoding (Instant-NGP), VIRUS-NeRF incorporates depth measurements from ultrasonic and infrared sensors and utilizes them to update the occupancy grid used for ray marching. Experimental evaluation in 2D demonstrates that VIRUS-NeRF achieves comparable mapping performance to LiDAR point clouds regarding coverage. Notably, in small environments, its accuracy aligns with that of LiDAR measurements, while in larger ones, it is bounded by the utilized ultrasonic sensors. An in-depth ablation study reveals that adding ultrasonic and infrared sensors is highly effective when dealing with sparse data and low view variation. Further, the proposed occupancy grid of VIRUS-NeRF improves the mapping capabilities and increases the training speed by 46% compared to Instant-NGP. Overall, VIRUS-NeRF presents a promising approach for cost-effective local mapping in mobile robotics, with potential applications in safety and navigation tasks. The code can be found at https://github.com/ethz-asl/virus_nerf.
Nicolaj Schmid, Cornelius von Einem, Cesar Dario Cadena Lerma, Roland Siegwart, Lorenz Hruby, Florian Tschopp
IROS3
2023 Robust and Efficient Edge-guided Pose Estimation with Resolution-conditioned NeRF
Liesbeth Claessens, Fabian Manhardt, Ricardo Martin-Brualla, Roland Siegwart, Cesar Dario Cadena Lerma, Federico Tombari
BMVC5
2023 Unsupervised Continual Semantic Adaptation Through Neural Rendering
abstract
An increasing amount of applications rely on data-driven models that are deployed for perception tasks across a sequence of scenes. Due to the mismatch between training and deployment data, adapting the model on the new scenes is often crucial to obtain good performance. In this work, we study continual multi-scene adaptation for the task of semantic segmentation, assuming that no ground-truth labels are available during deployment and that performance on the previous scenes should be maintained. We propose training a Semantic-NeRF network for each scene by fusing the predictions of a segmentation model and then using the view-consistent rendered semantic labels as pseudo-labels to adapt the model. Through joint training with the segmentation model, the Semantic-NeRF model effectively enables 2D-3D knowledge transfer. Furthermore, due to its compact size, it can be stored in a long-term memory and subsequently used to render data from arbitrary viewpoints to reduce forgetting. We evaluate our approach on Scan-Net, where we outperform both a voxel-based baseline and a state-of-the-art unsupervised domain adaptation method.
Zhizheng Liu, Francesco Milano 0001, Jonas Frey, Roland Siegwart, Hermann Blum, Cesar Dario Cadena Lerma
CVPR6
2023 SphNet: A Spherical Network for Semantic Pointcloud Segmentation
abstract
Semantic segmentation for robotic systems can enable a wide range of applications, from self-driving cars and augmented reality systems to domestic robots. We argue that a spherical representation is a natural one for egocentric pointclouds. Thus, in this work, we present a novel framework exploiting such a representation of LiDAR pointclouds for the task of semantic segmentation. Our approach is based on a spherical convolutional neural network that can seamlessly handle observations from various sensor systems (e.g., different LiDAR systems) and provides an accurate segmentation of the environment. We operate in two distinct stages: First, we encode the projected input pointclouds to spherical features. Second, we decode and back-project the spherical features to achieve an accurate semantic segmentation of the pointcloud. We evaluate our method with respect to state-of-the-art projection-based semantic segmentation approaches using well-known public datasets. We demonstrate that the spherical representation enables us to provide more accurate segmentation and to have a better generalization to sensors with different field-of-view and number of beams than what was seen during training.
Lukas Bernreiter, Lionel Ott, Roland Siegwart, Cesar Dario Cadena Lerma
ICRA4
2023 3D VSG: Long-term Semantic Scene Change Prediction through 3D Variable Scene Graphs
abstract
Numerous applications require robots to operate in environments shared with other agents, such as humans or other robots. However, such shared scenes are typically subject to different kinds of long-term semantic scene changes. The ability to model and predict such changes is thus crucial for robot autonomy. In this work, we formalize the task of semantic scene variability estimation and identify three main varieties of semantic scene change: changes in the position of an object, its semantic state, or the composition of a scene as a whole. To represent this variability, we propose the Variable Scene Graph (VSG), which augments existing 3D Scene Graph (SG) representations with the variability attribute, representing the likelihood of discrete long-term change events. We present a novel method, DeltaVSG, to estimate the variability of VSGs in a supervised fashion. We evaluate our method on the 3RScan long-term dataset, showing notable improvements in this novel task over existing approaches. Our method DeltaVsgachieves an accuracy of 77.1% and a recall of 72.3%, often mimicking human intuition about how indoor scenes change over time. We further show the utility of VSG prediction in the task of active robotic change detection, speeding up task completion by 66.0% compared to a scene-change-unaware planner. We make our code available as open-source.
Samuel Looper, Javier Rodriguez Puigvert, Roland Siegwart, Cesar Dario Cadena Lerma, Lukas Schmid 0001
ICRA4
2023 Obstacle avoidance using Raycasting and Riemannian Motion Policies at kHz rates for MAVs
abstract
This paper presents a novel method for using Riemannian Motion Policies on volumetric maps, shown in the example of obstacle avoidance for Micro Aerial Vehicles (MAVs), Today, most robotic obstacle avoidance algorithms rely on sampling or optimization-based planners with volumetric maps. However, they are computationally expensive and often have inflexible monolithic architectures. Riemannian Motion Policies are a modular, parallelizable, and efficient navigation alternative but are challenging to use with the widely used voxel-based environment representations. We propose using GPU raycasting and tens of thousands of concurrent policies to provide direct obstacle avoidance using Riemannian Motion Policies in voxelized maps without needing map smoothing or pre-processing. Additionally, we present how the same method can directly plan on LiDAR scans without any intermediate map. We show how this reactive approach compares favorably to traditional planning methods and can evaluate up to 200 million rays per second. We demonstrate the planner successfully on a real MAV for static and dynamic obstacles. The presented planner is made available as an open-source package11https://github.com/ethz-asl/reactive_avoidance.
Michael Pantic, Isar Meijer, Rik Girod, Nikhilesh Alatur, Olov Andersson, Cesar Dario Cadena Lerma, Roland Siegwart, Lionel Ott
ICRA6
2023 Local and Global Information in Obstacle Detection on Railway Tracks
abstract
Reliable obstacle detection on railways could help prevent collisions that result in injuries and potentially damage or derail the train. Unfortunately, generic object detectors do not have enough classes to account for all possible scenarios, and datasets featuring objects on railways are challenging to obtain. We propose utilizing a shallow network to learn railway segmentation from normal railway images. The limited receptive field of the network prevents overconfident predictions and allows the network to focus on the locally very distinct and repetitive patterns of the railway environment. Additionally, we explore the controlled inclusion of global information by learning to hallucinate obstacle-free images. We evaluate our method on a custom dataset featuring railway images with artificially augmented obstacles. Our proposed method outperforms other learning-based baseline methods.
Matthias Brucker, Andrei Cramariuc, Cornelius von Einem, Roland Siegwart, Cesar Dario Cadena Lerma
IROS5
2022 Collaborative Robot Mapping using Spectral Graph Analysis
abstract
In this paper, we deal with the problem of creating globally consistent pose graphs in a centralized multi-robot SLAM framework. For each robot to act autonomously, individual onboard pose estimates and maps are maintained, which are then communicated to a central server to build an optimized global map. However, inconsistencies between onboard and server estimates can occur due to onboard odometry drift or failure. Furthermore, robots do not benefit from the collaborative map if the server provides no feedback in a computationally tractable and bandwidth-efficient manner. Motivated by this challenge, this paper proposes a novel collaborative mapping framework to enable accurate global mapping among robots and server. In particular, structural differences between robot and server graphs are exploited at different spatial scales using graph spectral analysis to generate necessary constraints for the individual robot pose graphs. The proposed approach is thoroughly analyzed and validated using several real-world multi-robot field deployments where we show improvements of the onboard system up to 90%.
Lukas Bernreiter, Shehryar Khattak, Lionel Ott, Roland Siegwart, Marco Hutter 0001, Cesar Dario Cadena Lerma
ICRA6
2022 Panoptic Multi-TSDFs: a Flexible Representation for Online Multi-resolution Volumetric Mapping and Long-term Dynamic Scene Consistency
abstract
For robotic interaction in environments shared with other agents, access to volumetric and semantic maps of the scene is crucial. However, such environments are inevitably subject to long-term changes, which the map needs to account for. We thus propose panoptic multi-TSDFs as a novel representation for multi-resolution volumetric mapping in changing environments. By leveraging high-level information for 3D reconstruction, our proposed system allocates high resolution only where needed. Through reasoning on the object level, semantic consistency over time is achieved. This enables our method to maintain up-to-date reconstructions with high accuracy while improving coverage by incorporating previous data. We show in thorough experimental evaluation that our map can be efficiently constructed, maintained, and queried during online operation, and that the presented approach can operate robustly on real depth sensors using non-optimized panoptic segmentation as input.
Lukas Schmid 0001, Jeffrey A. Delmerico, Johannes L. Schönberger, Juan I. Nieto 0001, Marc Pollefeys, Roland Siegwart, Cesar Dario Cadena Lerma
ICRA7
2022 See Yourself in Others: Attending Multiple Tasks for Own Failure Detection
abstract
Autonomous robots deal with unexpected scenarios in real environments. Given input images, various visual perception tasks can be performed, e.g., semantic segmentation, depth estimation and normal estimation. These different tasks provide rich information for the whole robotic perception system. All tasks have their own characteristics while sharing some latent correlations. However, some of the task predictions may suffer from the unreliability dealing with complex scenes and anomalies. We propose an attention-based failure detection approach by exploiting the correlations among multiple tasks. The proposed framework infers task failures by evaluating the individual prediction, across multiple visual perception tasks for different regions in an image. The formulation of the evaluations is based on an attention network supervised by multi-task uncertainty estimation and their corresponding prediction errors. Our proposed framework11Code link https://github.com/ethz-asl/uncertainty_with_multiple_tasks. generates more accurate estimations of the prediction error for the different task's predictions.
Jiaxu Xing, Hermann Blum, Roland Siegwart, Cesar Dario Cadena Lerma
ICRA5
2022 Don't Share My Face: Privacy Preserving Inpainting for Visual Localization
abstract
Visual localization is an important task for many robotic and augmented reality applications. As localizing within large scale maps can be memory and computationally de-manding, cloud-based localization services are appealing for developers. However, such services raise important privacy concerns for both passive and active users. In particular, some sensitive information might be revealed by an attacker who intercepts data during the data-sharing process. Therefore, the sensitive data in the image should be concealed before it is shared. As a motivating case, we demonstrated the exposure generated by the common feature descriptor SIFT when attempting to recover private content. In this paper, we propose a pipeline to effectively conceal privacy-sensitive image regions from possible attacks to the transmission or localization services, by making use of learning-based image inpainting techniques while preserving, and even boosting, the localization performance. We tested our pipeline with two off-the-shelf localization services based on deep neural networks on the publicly available Oxford Robotcar dataset, showing that the localization performance on our generated private concealed images is on par with the non-private baseline.1
Saad Himmi, Oguzhan Ilter, François Pailleau, Roland Siegwart, Berta Bescós, Cesar Dario Cadena Lerma
IROS6
2022 SCIM: Simultaneous Clustering, Inference, and Mapping for Open-World Semantic Scene Understanding
Hermann Blum, Marcus Gerhard Müller, Abel Gawel, Roland Siegwart, Cesar Dario Cadena Lerma
ISRR5
2021 NeuralBlox: Real-Time Neural Representation Fusion for Robust Volumetric Mapping
abstract
We present a novel 3D mapping method leveraging the recent progress in neural implicit representation for 3D reconstruction. Most existing state-of-the-art neural implicit representation methods are limited to object-level reconstructions and can not incrementally perform updates given new data. In this work, we propose a fusion strategy and training pipeline to incrementally build and update neural implicit representations that enable the reconstruction of large scenes from sequential partial observations. By representing an arbitrarily sized scene as a grid of latent codes and performing updates directly in latent space, we show that incrementally built occupancy maps can be obtained in real-time even on a CPU. Compared to traditional approaches such as Truncated Signed Distance Fields (TSDFs), our map representation is significantly more robust in yielding a better scene completeness given noisy inputs. We demonstrate the performance of our approach in thorough experimental validation on real-world datasets with varying degrees of added pose noise.
Stefan Lionar, Lukas Schmid 0001, Cesar Dario Cadena Lerma, Roland Siegwart, Andrei Cramariuc
3DV3
2021 Pixel-Wise Anomaly Detection in Complex Driving Scenes
abstract
The inability of state-of-the-art semantic segmentation methods to detect anomaly instances hinders them from being deployed in safety-critical and complex applications, such as autonomous driving. Recent approaches have focused on either leveraging segmentation uncertainty to identify anomalous areas or re-synthesizing the image from the semantic label map to find dissimilarities with the input image. In this work, we demonstrate that these two methodologies contain complementary information and can be combined to produce robust predictions for anomaly segmentation. We present a pixel-wise anomaly detection framework that uses uncertainty maps to improve over existing re-synthesis methods in finding dissimilarities between the input and generated images. Our approach works as a general framework around already trained segmentation networks, which ensures anomaly detection without compromising segmentation accuracy, while significantly out performing all similar methods. Top-2 performance across a range of different anomaly datasets shows the robustness of our approach to handling different anomaly instances.
Giancarlo Di Biase, Hermann Blum, Roland Siegwart, Cesar Dario Cadena Lerma
CVPR4
2021 Spherical Multi-Modal Place Recognition for Heterogeneous Sensor Systems
abstract
In this paper, we propose a robust end-to-end multi-modal pipeline for place recognition where the sensor systems can differ from the map building to the query. Our approach operates directly on images and LiDAR scans without requiring any local feature extraction modules. By projecting the sensor data onto the unit sphere, we learn a multi-modal descriptor of partially overlapping scenes using a spherical convolutional neural network. The employed spherical projection model enables the support of arbitrary LiDAR and camera systems readily without losing information. Loop closure candidates are found using a nearest-neighbor lookup in the embedding space. We tackle the problem of correctly identifying the closest place by correlating the candidates’ power spectra, obtaining a confidence value per prospect. Our estimate for the correct place corresponds then to the candidate with the highest confidence. We evaluate our proposal w.r.t. state-of-the-art approaches in place recognition using real-world data acquired using different sensors. Our approach can achieve a recall that is up to 10% and 5% higher than for a LiDAR- and vision-based system, respectively, when the sensor setup differs between model training and deployment. Additionally, our place selection can correctly identify up to 95% matches from the candidate set.
Lukas Bernreiter, Lionel Ott, Juan I. Nieto 0001, Roland Siegwart, Cesar Dario Cadena Lerma
ICRA5
2021 CalQNet - Detection of Calibration Quality for Life-Long Stereo Camera Setups
abstract
Many mobile robotic platforms rely on an accurate knowledge of the extrinsic calibration parameters, especially systems performing visual stereo matching. Although a number of accurate stereo camera calibration methods have been developed, which provide good initial “factory” calibrations, the determined parameters can lose their validity over time as the sensors are exposed to environmental conditions and external effects. Thus, on autonomous platforms on-board diagnostic methods for an early detection of the need to repeat calibration procedures have the potential to prevent critical failures of crucial systems, such as state estimation or obstacle detection. In this work, we present a novel data-driven method to estimate the quality of extrinsic calibration and detect discrepancies between the original calibration and the current system state for stereo camera systems. The framework consists of a novel dataset generation pipeline to train CalQNet, a deep convolutional neural network. CalQNet can estimate the extrinsic calibration quality using a new metric that approximates the degree of miscalibration in stereo setups. We show the framework's ability to predict the divergence of a state-of-the-art stereo-visual odometry system following a degraded calibration in two real-world experiments.
Jiapeng Zhong, Zheyu Ye, Andrei Cramariuc, Florian Tschopp, Jen Jen Chung, Roland Siegwart, Cesar Dario Cadena Lerma
IV7
2021 The Fishyscapes Benchmark: Measuring Blind Spots in Semantic Segmentation
abstract
Abstract Deep learning has enabled impressive progress in the accuracy of semantic segmentation. Yet, the ability to estimate uncertainty and detect failure is key for safety-critical applications like autonomous driving. Existing uncertainty estimates have mostly been evaluated on simple tasks, and it is unclear whether these methods generalize to more complex scenarios. We present Fishyscapes, the first public benchmark for anomaly detection in a real-world task of semantic segmentation for urban driving. It evaluates pixel-wise uncertainty estimates towards the detection of anomalous objects. We adapt state-of-the-art methods to recent semantic segmentation models and compare uncertainty estimation approaches based on softmax confidence, Bayesian learning, density estimation, image resynthesis, as well as supervised anomaly detection methods. Our results show that anomaly detection is far from solved even for ordinary situations, while our benchmark allows measuring advancements beyond the state-of-the-art. Results, data and submission information can be found at https://fishyscapes.com/ .
Hermann Blum, Paul-Edouard Sarlin, Juan I. Nieto 0001, Roland Siegwart, Cesar Dario Cadena Lerma
Int. J. Comput. Vis.5
2021 Empty Cities: A Dynamic-Object-Invariant Space for Visual SLAM
abstract
In this article, we present a data-driven approach to obtain the static image of a scene, eliminating dynamic objects that might have been present at the time of traversing the scene with a camera. The general objective is to improve vision-based localization and mapping tasks in dynamic environments, where the presence (or absence) of different dynamic objects in different moments makes these tasks less robust. We introduce an end-to-end deep learning framework to turn images of an urban environment that include dynamic content, such as vehicles or pedestrians, into realistic static frames suitable for localization and mapping. This objective faces two main challenges: detecting the dynamic objects, and inpainting the static occluded background. The first challenge is addressed by the use of a convolutional network that learns a multiclass semantic segmentation of the image. The second challenge is approached with a generative adversarial model that, taking as input the original dynamic image and the computed dynamic/static binary mask, is capable of generating the final static image. This framework makes use of two new losses, one based on image steganalysis techniques, useful to improve the inpainting quality, and another one based on ORB features, designed to enhance feature matching between real and hallucinated image regions. To validate our approach, we perform an extensive evaluation on different tasks that are affected by dynamic entities, i.e.,visual odometry, place recognition, and multiview stereo, with the hallucinated images. Code has been made available on https://github.com/bertabescos/EmptyCities_SLAM.
Berta Bescós, Cesar Dario Cadena Lerma, José Neira
IEEE Trans. Robotics2
2020 Learning Camera Miscalibration Detection
abstract
Self-diagnosis and self-repair are some of the key challenges in deploying robotic platforms for long-term real-world applications. One of the issues that can occur to a robot is miscalibration of its sensors due to aging, environmental transients, or external disturbances. Precise calibration lies at the core of a variety of applications, due to the need to accurately perceive the world. However, while a lot of work has focused on calibrating the sensors, not much has been done towards identifying when a sensor needs to be recalibrated. This paper focuses on a data-driven approach to learn the detection of miscalibration in vision sensors, specifically RGB cameras. Our contributions include a proposed miscalibration metric for RGB cameras and a novel semi-synthetic dataset generation pipeline based on this metric. Additionally, by training a deep convolutional neural network, we demonstrate the effectiveness of our pipeline to identify whether a recalibration of the camera’s intrinsic parameters is required or not. The code is available at http://github.com/ethz-asl/camera_miscalib_detection.
Andrei Cramariuc, Aleksandar Petrov, Rohit Suri, Mayank Mittal, Roland Siegwart, Cesar Dario Cadena Lerma
ICRA6
2020 MOZARD: Multi-Modal Localization for Autonomous Vehicles in Urban Outdoor Environments
abstract
Visually poor scenarios are one of the main sources of failure in visual localization systems in outdoor environments. To address this challenge, we present MOZARD, a multi-modal localization system for urban outdoor environments using vision and LiDAR. By fusing key point based visual multi-session information with semantic data, an improved localization recall can be achieved across vastly different appearance conditions. In particular we focus on the use of curbstone information because of their broad distribution and reliability within urban environments. We present thorough experimental evaluations on several driving kilometers in challenging urban outdoor environments, analyze the recall and accuracy of our localization system and demonstrate in a case study possible failure cases of each subsystem. We demonstrate that MOZARD is able to bridge scenarios where our previous key point based visual approach, VIZARD, fails, hence yielding an increased recall performance, while a similar localization accuracy of 0.2m is achieved.
Lukas Schaupp, Patrick Pfreundschuh, Mathias Bürki, Cesar Dario Cadena Lerma, Roland Siegwart, Juan I. Nieto 0001
IROS4
2020 Learning Common and Transferable Feature Representations for Multi-Modal Data
abstract
LiDAR sensors are crucial in automotive perception for accurate object detection. However, LiDAR data is hard to interpret for humans and consequently time-consuming to label. Whereas camera data is easy interpretable and thus, comparably simpler to label. Within this work we present a transductive transfer learning approach to transfer the knowledge for the object detection task from images to point cloud data. We propose a multi-modal adversarial Auto Encoder architecture which disentangles uni-modal features into two groups: common (transferable) features, and complementary (modality-specific) features. This disentanglement is based on the hypothesis that a set of common features exist. An important point of our framework is that the disentanglement is learned in an unsupervised manner. Furthermore, the results show that only a small amount of multi-modal data is needed to learn the disentanglement, and thus to transfer the knowledge between modalities. As a result we our experiments show that training with 75% less data of the KITTI objects, the classification accuracy achieved is of 71.75%, only 3.12% less than when using the full data set. The implications of these findings can have great impact in perception pipelines based on LIDAR data.
Julia Nitsch, Juan I. Nieto 0001, Roland Siegwart, Max Schmidt, Cesar Dario Cadena Lerma
IV5
2019 From Coarse to Fine: Robust Hierarchical Localization at Large Scale
abstract
Robust and accurate visual localization is a fundamental capability for numerous applications, such as autonomous driving, mobile robotics, or augmented reality. It remains, however, a challenging task, particularly for large-scale environments and in presence of significant appearance changes. State-of-the-art methods not only struggle with such scenarios, but are often too resource intensive for certain real-time applications. In this paper we propose HF-Net, a hierarchical localization approach based on a monolithic CNN that simultaneously predicts local features and global descriptors for accurate 6-DoF localization. We exploit the coarse-to-fine localization paradigm: we first perform a global retrieval to obtain location hypotheses and only later match local features within those candidate places. This hierarchical approach incurs significant runtime savings and makes our system suitable for real-time operation. By leveraging learned descriptors, our method achieves remarkable localization robustness across large variations of appearance and sets a new state-of-the-art on two challenging benchmarks for large-scale localization.
Paul-Edouard Sarlin, Cesar Dario Cadena Lerma, Roland Siegwart, Marcin Dymczyk
CVPR2
2019 Empty Cities: Image Inpainting for a Dynamic-Object-Invariant Space
abstract
In this paper we present an end-to-end deep learning framework to turn images that show dynamic content, such as vehicles or pedestrians, into realistic static frames. This objective encounters two main challenges: detecting all the dynamic objects, and inpainting the static occluded background with plausible imagery. The former challenge is addressed by the use of a convolutional network that learns a multiclass semantic segmentation of the image. The second problem is approached with a conditional generative adversarial model that, taking as input the original dynamic image and its dynamic/static binary mask, is capable of generating the final static image. These generated images can be used for applications such as augmented reality or vision-based robot localization purposes. To validate our approach, we show both qualitative and quantitative comparisons against other state-of-the-art inpainting methods by removing the dynamic objects and hallucinating the static structure behind them. Furthermore, to demonstrate the potential of our results, we carry out pilot experiments that show the benefits of our proposal for visual place recognition.
Berta Bescós, José Neira, Roland Siegwart, Cesar Dario Cadena Lerma
ICRA4
2019 An Approach for Semantic Segmentation of Tree-like Vegetation
abstract
This paper presents a pipeline for semantic segmentation of trees into their components. Given a single RGB-D image of a tree, we employ a deep network to predict labels to classify each pixel of the tree into trunk, branches, twigs and leaves. Multiple convolutional neural network architectures to combine the complementary modalities of depth and colour data are investigated. An asynchronous training approach where two networks trained separately on RGB and depth encoded as a 3-channel HHA image are combined using a late fusion architecture with different learning rates performs the best. Training and evaluation are performed on a synthetic dataset of 6 species of broadleaf trees. We further demonstrate the network's generalization capabilities, across various tree species on the synthetic dataset, achieving an accuracy of upto 92.5%. Furthermore, we present a qualitative evaluation of our approach on real-world data.
Sundara Tejaswi Digumarti, Lukas Schmid 0001, Giuseppe Maria Rizzi, Juan I. Nieto 0001, Roland Siegwart, Paul A. Beardsley, Cesar Dario Cadena Lerma
ICRA7
2019 Object Classification Based on Unsupervised Learned Multi-Modal Features For Overcoming Sensor Failures
abstract
For autonomous driving applications it is critical to know which type of road users and road side infrastructure are present to plan driving manoeuvres accordingly. Therefore autonomous cars are equipped with different sensor modalities to robustly perceive its environment. However, for classification modules based on machine learning techniques it is challenging to overcome unseen sensor noise. This work presents an object classification module operating on unsupervised learned multi-modal features with the ability to overcome gradual or total sensor failure. A two stage approach composed of an unsupervised feature training and a uni-modal and multimodal classifiers training is presented. We propose a simple but effective decision module switching between uni-modal and multi-modal classifiers based on the closeness in the feature space to the training data. Evaluations on the ModelNet 40 data set show that the proposed approach has a 14% accuracy gain compared to a late fusion approach operating on a noisy point cloud data and a 6% accuracy gain when operating on noisy image data.
Julia Nitsch, Juan I. Nieto 0001, Roland Siegwart, Max Schmidt, Cesar Dario Cadena Lerma
ICRA5
2019 Flexible Trinocular: Non-rigid Multi-Camera-IMU Dense Reconstruction for UAV Navigation and Mapping
abstract
In this paper, we propose a visual-inertial framework able to efficiently estimate the camera poses of a non-rigid trinocular baseline for long-range depth estimation on-board a fast moving aerial platform. The estimation of the time-varying baseline is based on relative inertial measurements, a photometric relative pose optimizer, and a probabilistic wing model fused in an efficient Extended Kalman Filter (EKF) formulation. The estimated depth measurements can be integrated into a geo-referenced global map to render a reconstruction of the environment useful for local replanning algorithms. Based on extensive real-world experiments we describe the challenges and solutions for obtaining the probabilistic wing model, reliable relative inertial measurements, and vision-based relative pose updates and demonstrate the computational efficiency and robustness of the overall system under challenging conditions.
Timo Hinzmann, Cesar Dario Cadena Lerma, Juan I. Nieto 0001, Roland Siegwart
IROS2
2019 Free-Space Features: Global Localization in 2D Laser SLAM Using Distance Function Maps
abstract
In many applications, maintaining a consistent map of the environment is key to enabling robotic platforms to perform higher-level decision making. Detection of already visited locations is one of the primary ways in which map consistency is maintained, especially in situations where external positioning systems are unavailable or unreliable. Mapping in 2D is an important field in robotics, largely due to the fact that man-made environments such as warehouses and homes, where robots are expected to play an increasing role, can often be approximated as planar. Place recognition in this context remains challenging: 2D lidar scans contain scant information with which to characterize, and therefore recognize, a location. This paper introduces a novel approach aimed at addressing this problem. At its core, the system relies on the use of the distance function for representation of geometry. This representation allows extraction of features which describe the geometry of both surfaces and free-space in the environment. We propose a feature for this purpose. Through evaluations on public datasets, we demonstrate the utility of free-space in the description of places, and show an increase in localization performance over a state-of-the-art descriptor extracted from surface geometry.
Alexander Millane, Helen Oleynikova, Juan I. Nieto 0001, Roland Siegwart, Cesar Dario Cadena Lerma
IROS5
2019 OREOS: Oriented Recognition of 3D Point Clouds in Outdoor Scenarios
abstract
We introduce a novel method for oriented place recognition with 3D LiDAR scans. A Convolutional Neural Network is trained to extract compact descriptors from single 3D LiDAR scans. These can be used both to retrieve near-by place candidates from a map, and to estimate the yaw discrepancy needed for bootstrapping local registration methods. We employ a triplet loss function for training and use a hard-negative mining strategy to further increase the performance of our descriptor extractor. In an extensive evaluation on the NCLT and KITTI datasets, we demonstrate that our method outperforms related state-of-the-art approaches based on both data-driven and handcrafted data representation in challenging long-term outdoor conditions.
Lukas Schaupp, Mathias Bürki, Renaud Dubé, Roland Siegwart, Cesar Dario Cadena Lerma
IROS5
2019 VIZARD: Reliable Visual Localization for Autonomous Vehicles in Urban Outdoor Environments
abstract
Changes in appearance is one of the main sources of failure in visual localization systems in outdoor environments. To address this challenge, we present VIZARD, a visual localization system for urban outdoor environments. By combining a local localization algorithm with the use of multi-session maps, a high localization recall can be achieved across vastly different appearance conditions. The fusion of the visual localization constraints with wheel-odometry in a state estimation framework further guarantees smooth and accurate pose estimates. In an extensive experimental evaluation on several hundreds of driving kilometers in challenging urban outdoor environments, we analyze the recall and accuracy of our localization system, investigate its key parameters and boundary conditions, and compare different types of feature descriptors. Our results show that VIZARD is able to achieve nearly 100% recall with a localization accuracy below 0.5m under varying outdoor appearance conditions, including at night-time.
Mathias Bürki, Lukas Schaupp, Marcin Dymczyk, Renaud Dubé, Cesar Dario Cadena Lerma, Roland Siegwart, Juan I. Nieto 0001
IV5
2018 A Data-driven Model for Interaction-Aware Pedestrian Motion Prediction in Object Cluttered Environments
abstract
This paper reports on a data-driven, interaction-aware motion prediction approach for pedestrians in environments cluttered with static obstacles. When navigating in such workspaces shared with humans, robots need accurate motion predictions of the surrounding pedestrians. Human navigation behavior is mostly influenced by their surrounding pedestrians and by the static obstacles in their vicinity. In this paper we introduce a new model based on Long-Short Term Memory (LSTM) neural networks, which is able to learn human motion behavior from demonstrated data. To the best of our knowledge, this is the first approach using LSTMs, that incorporates both static obstacles and surrounding pedestrians for trajectory forecasting. As part of the model, we introduce a new way of encoding surrounding pedestrians based on a 1d-grid in polar angle space. We evaluate the benefit of interaction-aware motion prediction and the added value of incorporating static obstacles on both simulation and real-world datasets by comparing with state-of-the-art approaches. The results show, that our new approach outperforms the other approaches while being very computationally efficient and that taking into account static obstacles for motion predictions significantly improves the prediction accuracy, especially in cluttered environments.
Mark Pfeiffer, Giuseppe Paolo, Hannes Sommer, Juan I. Nieto 0001, Roland Siegwart, Cesar Dario Cadena Lerma
ICRA6
2018 Modular Sensor Fusion for Semantic Segmentation
abstract
Sensor fusion is a fundamental process in robotic systems as it extends the perceptual range and increases robustness in real-world operations. Current multi-sensor deep learning based semantic segmentation approaches do not provide robustness to under-performing classes in one modality, or require a specific architecture with access to the full aligned multi-sensor training data. In this work, we analyze statistical fusion approaches for semantic segmentation that overcome these drawbacks while keeping a competitive performance. The studied approaches are modular by construction, allowing to have different training sets per modality and only a much smaller subset is needed to calibrate the statistical models. We evaluate a range of statistical fusion approaches and report their performance against state-of-the-art baselines on both realworld and simulated data. In our experiments, the approach improves performance in IoU over the best single modality segmentation results by up to 5%. We make all implementations and configurations publicly available.
Hermann Blum, Abel Gawel, Roland Siegwart, Cesar Dario Cadena Lerma
IROS4
2018 C-blox: A Scalable and Consistent TSDF-based Dense Mapping Approach
abstract
In many applications, maintaining a consistent dense map of the environment is key to enabling robotic platforms to perform higher level decision making. Several works have addressed the challenge of creating precise dense 3D maps from visual sensors providing depth information. However, during operation over longer missions, reconstructions can easily become inconsistent due to accumulated camera tracking error and delayed loop closure. Without explicitly addressing the problem of map consistency, recovery from such distortions tends to be difficult. We present a novel system for dense 3D mapping which addresses the challenge of building consistent maps while dealing with scalability. Central to our approach is the representation of the environment as a collection of overlapping Truncated Signed Distance Field (TSDF) subvolumes. These subvolumes are localized through feature-based camera tracking and bundle adjustment. Our main contribution is a pipeline for identifying stable regions in the map, and to fuse the contributing subvolumes. This approach allows us to reduce map growth while still maintaining consistency. We demonstrate the proposed system on a publicly available dataset and simulation engine, and demonstrate the efficacy of the proposed approach for building consistent and scalable maps. Finally we demonstrate our approach running in real-time onboard a lightweight Micro Aerial Vehicle (MAV).
Alexander Millane, Zachary Taylor, Helen Oleynikova, Juan I. Nieto 0001, Roland Siegwart, Cesar Dario Cadena Lerma
IROS6
2018 Map Management for Efficient Long-Term Visual Localization in Outdoor Environments
abstract
We present a complete map management process for a visual localization system designed for multi-vehicle long-term operations in resource constrained outdoor environments. Outdoor visual localization generates large amounts of data that need to be incorporated into a lifelong visual map in order to allow localization at all times and under all appearance conditions. Processing these large quantities of data is non-trivial, as it is subject to limited computational and storage capabilities both on the vehicle and on the mapping backend. We address this problem with a two-fold map update paradigm capable of, either, adding new visual cues to the map, or updating co-observation statistics. The former, in combination with offline map summarization techniques, allows enhancing the appearance coverage of the lifelong map while keeping the map size limited. On the other hand, the latter is able to significantly boost the appearance-based landmark selection for efficient online localization without incurring any additional computational or storage burden. Our evaluation in challenging outdoor conditions shows that our proposed map management process allows building and maintaining maps for precise visual localization over long time spans in a tractable and scalable fashion.
Mathias Bürki, Marcin Dymczyk, Igor Gilitschenski, Cesar Dario Cadena Lerma, Roland Siegwart, Juan I. Nieto 0001
Intelligent Vehicles Symposium4
2017 SegMatch: Segment based place recognition in 3D point clouds
abstract
Place recognition in 3D data is a challenging task that has been commonly approached by adapting image-based solutions. Methods based on local features suffer from ambiguity and from robustness to environment changes while methods based on global features are viewpoint dependent. We propose SegMatch, a reliable place recognition algorithm based on the matching of 3D segments. Segments provide a good compromise between local and global descriptions, incorporating their strengths while reducing their individual drawbacks. SegMatch does not rely on assumptions of `perfect segmentation', or on the existence of `objects' in the environment, which allows for reliable execution on large scale, unstructured environments. We quantitatively demonstrate that SegMatch can achieve accurate localization at a frequency of 1Hz on the largest sequence of the KITTI odometry dataset. We furthermore show how this algorithm can reliably detect and close loops in real-time, during online operation. In addition, the source code for the SegMatch algorithm is made publicly available.
Renaud Dubé, Daniel Dugas, Elena Stumm, Juan I. Nieto 0001, Roland Siegwart, Cesar Dario Cadena Lerma
ICRA6
2017 TSDF-based change detection for consistent long-term dense reconstruction and dynamic object discovery
abstract
Robots that are operating for extended periods of time need to be able to deal with changes in their environment and represent them adequately in their maps. In this paper, we present a novel 3D reconstruction algorithm based on an extended Truncated Signed Distance Function (TSDF) that enables to continuously refine the static map while simultaneously obtaining 3D reconstructions of dynamic objects in the scene. This is a challenging problem because map updates happen incrementally and are often incomplete. Previous work typically performs change detection on point clouds, surfels or maps, which are not able to distinguish between unexplored and empty space. In contrast, our TSDF-based representation naturally contains this information and thus allows us to more robustly solve the scene differencing problem. We demonstrate the algorithms performance as part of a system for unsupervised object discovery and class recognition. We evaluated our algorithm on challenging datasets that we recorded over several days with RGB-D enabled tablets. To stimulate further research in this area, all of our datasets are publicly available3.
Marius Fehr, Fadri Furrer, Ivan Dryanovski, Jürgen Sturm, Igor Gilitschenski, Roland Siegwart, Cesar Dario Cadena Lerma
ICRA7
2017 From perception to decision: A data-driven approach to end-to-end motion planning for autonomous ground robots
abstract
Learning from demonstration for motion planning is an ongoing research topic. In this paper we present a model that is able to learn the complex mapping from raw 2D-laser range findings and a target position to the required steering commands for the robot. To our best knowledge, this work presents the first approach that learns a target-oriented end-to-end navigation model for a robotic platform. The supervised model training is based on expert demonstrations generated in simulation with an existing motion planner. We demonstrate that the learned navigation model is directly transferable to previously unseen virtual and, more interestingly, real-world environments. It can safely navigate the robot through obstacle-cluttered environments to reach the provided targets. We present an extensive qualitative and quantitative evaluation of the neural network-based motion planner, and compare it to a grid-based global approach, both in simulation and in real-world experiments.
Mark Pfeiffer, Michael Schaeuble, Juan I. Nieto 0001, Roland Siegwart, Cesar Dario Cadena Lerma
ICRA5
2017 An online multi-robot SLAM system for 3D LiDARs
abstract
Using multiple cooperative robots is advantageous for time critical Search and Rescue (SaR) missions as they permit rapid exploration of the environment and provide higher redundancy than using a single robot. A considerable number of applications such as autonomous driving and disaster response could benefit from merging mapping data from several agents. Online multi-robot localization and mapping has mainly been addressed for robots equipped with cameras or 2D LiDARs. However, in unstructured and ill-lighted real-life scenarios, a mapping system can potentially benefit from a rich 3D geometric solution. In this work, we present an online localization and mapping system for multiple robots equipped with 3D LiDARs. This system is based on incremental sparse pose-graph optimization using sequential and place recognition constraints, the latter being identified using a 3D segment matching approach. The result is a unified representation of the world and relative robot trajectories. The complete system runs in real-time and is evaluated with two experiments in different environments: one urban and one disaster scenario. The system is available open source and easy-to-run demonstrations are publicly available.
Renaud Dubé, Abel Gawel, Hannes Sommer, Juan I. Nieto 0001, Roland Siegwart, Cesar Dario Cadena Lerma
IROS6
2016 Direct semi-dense SLAM for rolling shutter cameras
abstract
In this paper, we present a monocular Direct and Semi-dense SLAM (Simultaneous Localization And Mapping) system for rolling shutter cameras. In a rolling shutter camera, the pose is different for each row of each image, and this yields poor pose estimates and poor structure estimates when using a state-of-the-art semi-dense direct method designed for global shutter cameras. To address this issue in tracking, we model the smooth and continuous camera trajectory using a B-spline curve of degree k??1 for poses in the Lie algebra, se(3).We solve for the camera poses at each row-time by a direct optimisation of photometric error as a function of the control points of the spline. Likewise for mapping, we develop generalised epipolar geometry for the rolling shutter case and solve for point depths using photometric error. Although each of these issues has been previously tackled, to the best of our knowledge ours is the first full solution to monocular, direct (feature-less) SLAM. We benchmark our method for pose accuracy and map accuracy against the state-of-the-art semi-dense SLAM system, LSD-SLAM, demonstrating the improved efficacy of our approach when using rolling shutter cameras via synthetic sequences with known ground-truth and real sequences.
Jae-Hak Kim, Cesar Dario Cadena Lerma, Ian D. Reid 0001
ICRA2
2016 Measuring the performance of single image depth estimation methods
abstract
We consider the question of benchmarking the performance of methods used for estimating the depth of a scene from a single image. We describe various measures that have been used in the past, discuss their limitations and demonstrate that each is deficient in one or more ways. We propose a new measure of performance for depth estimation that overcomes these deficiencies, and has a number of desirable properties. We show that in various cases of interest the new measure enables visualisation of the performance of a method that is otherwise obfuscated by existing metrics. Our proposed method is capable of illuminating the relative performance of different algorithms on different kinds of data, such as the difference in efficacy of a method when estimating the depth of the ground plane versus estimating the depth of other generic scene structure. We showcase the method by comparing a number of existing single-view methods against each other and against more traditional depth estimation methods such as binocular stereo.
Cesar Dario Cadena Lerma, Yasir Latif, Ian D. Reid 0001
IROS1
2016 Past, Present, and Future of Simultaneous Localization and Mapping: Toward the Robust-Perception Age
abstract
Simultaneous localization and mapping (SLAM) consists in the concurrent construction of a model of the environment (the map), and the estimation of the state of the robot moving within it. The SLAM community has made astonishing progress over the last 30 years, enabling large-scale real-world applications and witnessing a steady transition of this technology to industry. We survey the current state of SLAM and consider future directions. We start by presenting what is now the de-facto standard formulation for SLAM. We then review related work, covering a broad set of topics including robustness and scalability in long-term mapping, metric and semantic representations for mapping, theoretical performance guarantees, active SLAM and exploration, and other new frontiers. This paper simultaneously serves as a position paper and tutorial to those who are users of SLAM. By looking at the published research with a critical eye, we delineate open challenges and new research issues, that still deserve careful scientific investigation. The paper also contains the authors' take on two questions that often animate discussions during robotics conferences: Do robots need SLAM? and Is SLAM solved?
Cesar Dario Cadena Lerma, Luca Carlone, Henry Carrillo, Yasir Latif, Davide Scaramuzza 0001, José Neira, Ian D. Reid 0001, John J. Leonard
IEEE Trans. Robotics1
2015 A fast, modular scene understanding system using context-aware object detection
abstract
We propose a semantic scene understanding system that is suitable for real robotic operations. The system solves different tasks (semantic segmentation and object detections) in an opportunistic and distributed fashion but still allows communication between modules to improve their respective performances. We propose the use of the semantic space to improve specific out-of-the-box object detectors and an update model to take the evidence from different detection into account in the semantic segmentation process. Our proposal is evaluated with the KITTI dataset, on the object detection benchmark and on five different sequences manually annotated for the semantic segmentation task, demonstrating the efficacy of our approach.
Cesar Dario Cadena Lerma, Anthony R. Dick, Ian D. Reid 0001
ICRA1
2014 Semantic segmentation with heterogeneous sensor coverages
abstract
We propose a new approach to semantic parsing, which can seamlessly integrate evidence from multiple sensors with overlapping but possibly different fields of view (FOV), account for missing data and predict semantic labels over the spatial union of sensors coverages. The existing approaches typically carry out semantic segmentation using only one modality, incorrectly interpolate measurements of other modalities or at best assign semantic labels only to the spatial intersection of coverages of different sensors. In this work we remedy these problems by proposing an effective and efficient strategy for inducing the graph structure of Conditional Random Field used for inference and a novel method for computing the sensor domain dependent potentials. We focus on RGB cameras and 3D data from lasers or depth sensors. The proposed approach achieves superior performance, compared to state of the art and obtains labels for the union of spatial coverages of both sensors, while effectively using appearance or 3D cues when they are available. The efficiency of the approach is amenable to realtime implementation. We quantitatively validate our proposal in two publicly available datasets from indoors and outdoors real environments. The obtained semantic understanding of the acquired sensory information can enable higher level tasks for autonomous mobile robots and facilitate semantic mapping of the environments.
Cesar Dario Cadena Lerma, Jana Kosecka
ICRA1
2014 Robust graph SLAM back-ends: A comparative analysis
abstract
In this work, we provide an in-depth analysis of several recent robust Simultaneous Localization And Mapping (SLAM) back-end techniques that aim to recover the correct graph estimate in the presence of outliers in loop closure constraints. We present a benchmark dataset for evaluation of such methods by augmenting the KITTI Vision Benchmark with ground truth as well as generated loop closure hypotheses and present a detailed analysis of recently proposed robust SLAM methods using this benchmark. We also look into how these methods achieve the desired robustness and what are the implications for the SLAM problem. We discuss the issues involved in using the output of these robust back-ends for tasks such as path planning and how they can be addressed. The problem of robustness needs to be addressed adequately in order to have a complete and reliable solution to the SLAM problem.
Yasir Latif, Cesar Dario Cadena Lerma, José Neira
IROS2
2013 Recursive Inference for Prediction of Objects in Urban Environments
Cesar Dario Cadena Lerma, Jana Kosecka
ISRR1
2012 Realizing, reversing, recovering: Incremental robust loop closing over time using the iRRR algorithm
abstract
The ability to reconsider information over time allows to detect failures and is crucial for long term robust autonomous robot applications. This applies to loop closure decisions in localization and mapping systems. This paper describes a method to analyze all available information up to date in order to robustly remove past incorrect loop closures from the optimization process. The main novelties of our algorithm are: 1. incrementally reconsidering loop closures and 2. handling multi-session, spatially related or unrelated experiments. We validate our proposal in real multi-session experiments showing better results than those obtained by state of the art methods.
Yasir Latif, Cesar Dario Cadena Lerma, José Neira
IROS2
2012 Robust Place Recognition With Stereo Sequences
abstract
We propose a place recognition algorithm for simultaneous localization and mapping (SLAM) systems using stereo cameras that considers both appearance and geometric information of points of interest in the images. Both near and far scene points provide information for the recognition process. Hypotheses about loop closings are generated using a fast appearance-only technique based on the bag-of-words (BoW) method. We propose several important improvements to BoWs that profit from the fact that, in this problem, images are provided in sequence. Loop closing candidates are evaluated using a novel normalized similarity score that measures similarity in the context of recent images in the sequence. In cases where similarity is not sufficiently clear, loop closing verification is carried out using a method based on conditional random fields (CRFs). We build on CRF matching with two main novelties: We use both image and 3-D geometric information, and we carry out inference on a minimum spanning tree (MST), instead of a densely connected graph. Our results show that MSTs provide an adequate representation of the problem, with the additional advantages that exact inference is possible and that the computational cost of the inference process is limited. We compare our system with the state of the art using visual indoor and outdoor data from three different locations and show that our system can attain at least full precision (no false positives) for a higher recall (fewer false negatives).
Cesar Dario Cadena Lerma, Dorian Gálvez-López, Juan D. Tardós, José Neira
IEEE Trans. Robotics1
2010 Robust place recognition with stereo cameras
abstract
Place recognition is a challenging task in any SLAM system. Algorithms based on visual appearance are becoming popular to detect locations already visited, also known as loop closures, because cameras are easily available and provide rich scene detail. These algorithms typically result in pairs of images considered depicting the same location. To avoid mismatches, most of them rely on epipolar geometry to check spatial consistency. In this paper we present an alternative system that makes use of stereo vision and combines two complementary techniques: bag-of-words to detect loop closing candidate images, and conditional random fields to discard those which are not geometrically consistent. We evaluate this system in public indoor and outdoor datasets from the Rawseeds project, with hundred-metre long trajectories. Our system achieves more robust results than using spatial consistency based on epipolar geometry.
Cesar Dario Cadena Lerma, Dorian Gálvez-López, Fabio Ramos 0001, Juan D. Tardós, José Neira
IROS1
2009 SLAM in O(log n) with the Combined Kalman - Information filter
abstract
In this paper we show that SLAM can be executed in as low as O(log n) per step. Our algorithm, the Combined Filter SLAM, uses a combination of Extended Kalman and Extended Information filters in such a way that the total cost of building a map can be reduced to O(n log n), as compared with O(n3) for standard EKF SLAM, and O(n2) for Divide and Conquer (D&C) SLAM and the Sparse Local Submap Joining Filter (SLSJF). We discuss the computational improvements that have been proposed for Kalman and Information filters, discuss the advantages and limitations of each, and how a judicious combination results in the possibility of reducing the computational cost per step to O(log n).We use simulations and real datasets to show the advantages of the proposed algorithm.
Cesar Dario Cadena Lerma, José Neira
IROS1