Andrew J. Davison

dblp:d/AndrewJDavison · DBLP profile ↗
← Back
115ranked-venue papers
10as first author
28since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 96 · 8 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 61 · 7 first-author · 15 since 2021Systems, architecture and hardware · 35 · 1 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 8 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 1 since 2021
YearPublicationVenuePosition
2025 U-ARE-ME: Uncertainty-Aware Rotation Estimation in Manhattan Environments
abstract
Camera rotation estimation from a single image is a challenging task, often requiring depth data and/or camera intrinsics, which are generally not available for in-the-wild videos. Although external sensors such as inertial measurement units (IMUs) can help, they often suffer from drift and are not applicable in non-inertial reference frames. We present U-ARE-ME, an algorithm that estimates camera rotation along with uncertainty from uncalibrated RGB images. Using a Manhattan World assumption, our method leverages the per-pixel geometric priors encoded in single-image surface normal predictions and performs optimisation over the SOC 3) manifold. Given a sequence of images, we can use the per-frame rotation estimates and their uncertainty to perform multi-frame optimisation, achieving robustness and temporal consistency. Our experiments demonstrate that U-ARE-ME performs comparably to RGB-D methods and is more robust than feature-based vanishing point and SLAM methods.
Aalok Patwardhan, Callum Rhodes, Gwangbin Bae, Andrew J. Davison
3DV4
2025 4DTAM: Non-Rigid Tracking and Mapping via Dynamic Surface Gaussians
abstract
We propose the first 4D tracking and mapping method that jointly performs camera localization and non-rigid surface reconstruction via differentiable rendering. Our approach captures 4D scenes from an online stream of color images with depth measurements or predictions by simultaneously optimizing scene geometry, appearance, dynamics, and camera ego-motion. Although natural environments exhibit complex non-rigid motions, 4D-SLAM remains relatively underexplored due to its inherent challenges; even with 2.5D signals, the problem is ill-posed because of the high dimensionality of the optimization space. To overcome these challenges, we first introduce a SLAM method based on Gaussian surface primitives that leverages depth signals more effectively than 3D Gaussians, thereby achieving accurate surface reconstruction. To further model nonrigid deformations, we employ a warp-field represented by a multi-layer perceptron (MLP) and introduce a novel camera pose estimation technique along with surface regularization terms that facilitate spatio-temporal reconstruction. In addition to these algorithmic challenges, a significant hurdle in 4D-SLAM research is the lack of publicly available datasets with reliable ground truth and evaluation protocols. To address this, we present a novel synthetic dataset of everyday objects with diverse motions, leveraging large-scale object models and animation modeling. In summary, we open up the modern 4D-SLAM research by introducing a novel method and evaluation protocols grounded in modern vision and rendering techniques.
Hidenobu Matsuki, Gwangbin Bae, Andrew J. Davison
CVPR3
2025 MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priors
abstract
We present a real-time monocular dense SLAM system designed bottom-up from MASt3R, a two-view 3D reconstruction and matching prior. Equipped with this strong prior, our system is robust on in-the-wild video sequences despite making no assumption on a fixed or parametric camera model beyond a unique camera centre. We introduce efficient methods for pointmap matching, camera tracking and local fusion, graph construction and loop closure, and second-order global optimisation. With known calibration, a simple modification to the system achieves state-of-the-art performance across various benchmarks. Altogether, we propose a plug-and-play monocular SLAM system capable of producing globally consistent poses and dense geometry while operating at 15 FPS.
Riku Murai, Eric Dexheimer, Andrew J. Davison
CVPR3
2024 Rethinking Inductive Biases for Surface Normal Estimation
abstract
Despite the growing demand for accurate surface nor- mal estimation models, existing methods use general- purpose dense prediction models, adopting the same induc- tive biases as other tasks. In this paper, we discuss the inductive biases needed for surface normal estimation and propose to (1) utilize the per-pixel ray direction and (2) en- code the relationship between neighboring surface normals by learning their relative rotation. The proposed method can generate crisp - yet, piecewise smooth - predictions for challenging in-the-wild images of arbitrary resolution and aspect ratio. Compared to a recent ViT-based state- of-the-art model, our method shows a stronger generalization ability, despite being trained on an orders of magni- tude smaller dataset. The code is available at https://github.com/baegwangbin/DSINE.
Gwangbin Bae, Andrew J. Davison
CVPR2
2024 EscherNet: A Generative Model for Scalable View Synthesis
abstract
We introduce EscherNet, a multi-view conditioned diffusion model for view synthesis. EscherNet learns implicit and generative 3D representations coupled with a specialised camera positional encoding, allowing precise and continuous relative control of the camera transformation between an arbitrary number of reference and target views. EscherNet offers exceptional generality, flexibility, and scalability in view synthesis ─it can generate more than 100 consistent target views simultaneously on a single consumer-grade GPU, despite being trained with a fixed number of 3 reference views to 3 target views. As a result, EscherNet not only addresses zero-shot novel view synthesis, but also naturally unifies single- and multi-image 3D reconstruction, combining these diverse tasks into a single, cohesive framework. Our extensive experiments demonstrate that EscherNet achieves state-of-the-art performance in multiple benchmarks, even when compared to methods specifically tailored for each individual problem. This remarkable versatility opens up new directions for designing scalable neural architectures for 3D vision. Project page: https://kxhit.github.io/EscherNet.
Xin Kong, Shikun Liu, Xiaoyang Lyu, Marwan Taher, Xiaojuan Qi 0001, Andrew J. Davison
CVPR6
2024 Gaussian Splatting SLAM
abstract
We present the first application of 3D Gaussian Splatting in monocular SLAM, the most fundamental but the hardest setup for Visual SLAM. Our method, which runs live at 3fps, utilises Gaussians as the only 3D representation, unifying the required representation for accurate, efficient tracking, mapping, and high-quality rendering. Designed for challenging monocular settings, our approach is seamlessly extendable to RGB-D SLAM when an external depth sensor is available. Several innovations are required to continuously reconstruct 3D scenes with high fidelity from a live camera. First, to move beyond the original 3DGS algorithm, which requires accurate poses from an offline Structure from Motion (SfM) system, we formulate camera tracking for 3DGS using direct optimisation against the 3D Gaussians, and show that this enables fast and robust tracking with a wide basin of convergence. Second, by utilising the explicit nature of the Gaussians, we introduce geometric verification and regularisation to handle the ambiguities occurring in incremental 3D dense reconstruction. Finally, we introduce afull SLAM system which not only achieves state-of-the-art results in novel view synthesis and trajectory estimation but also reconstruction of tiny and even transparent objects.
Hidenobu Matsuki, Riku Murai, Paul H. J. Kelly, Andrew J. Davison
CVPR4
2024 SuperPrimitive: Scene Reconstruction at a Primitive Level
abstract
Joint camera pose and dense geometry estimation from a set of images or a monocular video remains a challenging problem due to its computational complexity and inherent visual ambiguities. Most dense incremental reconstruction systems operate directly on image pixels and solve for their 3D positions using multi-view geometry cues. Such pixellevel approaches suffer from ambiguities or violations of multi-view consistency (e.g. caused by textureless or specular surfaces). We address this issue with a new image representation which we call a SuperPrimitive. SuperPrimitives are obtained by splitting images into semantically correlated local regions and enhancing them with estimated surface normal directions, both of which are predicted by state-of-the-art single image neural networks. This provides a local geometry estimate per SuperPrimitive, while their relative positions are adjusted based on multi-view observations. We demonstrate the versatility of our new representation by addressing three 3D reconstruction tasks: depth completion, few-view structure from motion, and monocular dense visual odometry. Project page: https://makezur.github.io/SuperPrimitive/
Kirill Mazur, Gwangbin Bae, Andrew J. Davison
CVPR3
2024 COMO: Compact Mapping and Odometry
Eric Dexheimer, Andrew J. Davison
ECCV (55)2
2024 Learning in Deep Factor Graphs with Gaussian Belief Propagation
abstract
We propose an approach to do learning in Gaussian factor graphs. We treat all relevant quantities (inputs, outputs, parameters, activations) as random variables in a graphical model, and view training and prediction as inference problems with different observed nodes. Our experiments show that these problems can be efficiently solved with belief propagation (BP), whose updates are inherently local, presenting exciting opportunities for distributed and asynchronous training. Our approach can be scaled to deep networks and provides a natural means to do continual learning: use the BP-estimated posterior of the current task as a prior for the next. On a video denoising task we demonstrate the benefit of learnable parameters over a classical factor graph approach and we show encouraging performance of deep factor graphs for continual image classification.
Seth Nabarro, Mark van der Wilk, Andrew J. Davison
ICML3
2024 A Distributed Multi-Robot Framework for Exploration, Information Acquisition and Consensus
abstract
The distributed coordination of robot teams performing complex tasks is challenging to formulate. The different aspects of a complete task such as local planning for obstacle avoidance, global goal coordination and collaborative mapping are often solved separately, when clearly each of these should influence the others for the most efficient behaviour. In this paper we use the example application of distributed information acquisition as a robot team explores a large space to show that we can formulate the whole problem as a single factor graph with multiple connected layers representing each aspect. We use Gaussian Belief Propagation (GBP) as the inference mechanism, which permits parallel, on-demand or asynchronous computation for efficiency when different aspects are more or less important. This is the first time that a distributed GBP multi-robot solver has been proven to enable intelligent collaborative behaviour rather than just guiding robots to individual, selfish goals. We encourage the reader to view our demos at https://aalpatya.github.io/gbpstack.
Aalok Patwardhan, Andrew J. Davison
ICRA2
2024 Fit-NGP: Fitting Object Models to Neural Graphics Primitives
abstract
Accurate 3D object pose estimation is key to enabling many robotic applications that involve challenging object interactions. In this work, we show that the density field created by a state-of-the-art efficient radiance field reconstruction method is suitable for highly accurate and robust pose estimation for objects with known 3D models, even when they are very small and with challenging reflective surfaces. We present a fully automatic object pose estimation system based on a robot arm with a single wrist-mounted camera, which can scan a scene from scratch, detect and estimate the 6-Degrees of Freedom (DoF) poses of multiple objects within a couple of minutes of operation. Small objects such as bolts and nuts are estimated with accuracy on order of 1mm.
Marwan Taher, Ignacio Alzugaray, Andrew J. Davison
ICRA3
2024 A Robot Web for Distributed Many-Device Localization
abstract
We show that a distributed network of robots or other devices which make measurements of each other can collaborate to globally localize via efficient ad hoc peer-to-peer communication. Our Robot Web solution is based on Gaussian belief propagation (GBP) on the fundamental nonlinear factor graph describing the probabilistic structure of all of the observations robots make internally or of each other, and is flexible for any type of robot, motion or sensor. We define a simple and efficient communication protocol which can be implemented by the publishing and reading of web pages or other asynchronous communication technologies. We show in simulations with up to 1000 robots interacting in arbitrary patterns that our solution convergently achieves global accuracy as accurate as a centralized nonlinear factor graph solver while operating with high distributed efficiency of computation and communication. Via the use of robust factors in GBP, our method is tolerant to a high percentage of faulty sensor measurements or dropped communication packets. Furthermore, we showcase that the system operates on real robots with limited onboard computational resources.
Riku Murai, Joseph Ortiz, Sajad Saeedi G., Paul H. J. Kelly, Andrew J. Davison
IEEE Trans. Robotics5
2023 Learning a Depth Covariance Function
abstract
We propose learning a depth covariance function with applications to geometric vision tasks. Given RGB images as input, the covariance function can be flexibly used to define priors over depth functions, predictive distributions given observations, and methods for active point selection. We leverage these techniques for a selection of downstream tasks: depth completion, bundle adjustment, and monocular dense visual odometry.
Eric Dexheimer, Andrew J. Davison
CVPR2
2023 vMAP: Vectorised Object Mapping for Neural Field SLAM
abstract
We present vMAP, an object-level dense SLAM system using neural field representations. Each object is repre-sented by a small MLP, enabling efficient, watertight object modelling without the needfor 3D priors. As an RGB-D camera browses a scene with no prior in-formation, vMAP detects object instances on-the-fly, and dynamically adds them to its map. Specifically, thanks to the power of vectorised training, vMAP can optimise as many as 50 individual objects in a single scene, with an extremely efficient training speed of 5Hz map update. We experimentally demonstrate significantly improved scene-level and object-level reconstruction quality compared to prior neural field SLAM systems. Project page: https://kxhit.github.io/vMAP.
Xin Kong, Shikun Liu, Marwan Taher, Andrew J. Davison
CVPR4
2023 iMODE:Real-Time Incremental Monocular Dense Mapping Using Neural Field
abstract
We present a novel real-time dense and semantic neural field mapping system that uses only monocular images as input. Our scene representation is a dense continuous radiance field represented by a Multi-Layer Perceptron (MLP), trained from scratch in real-time. We build on high-performance sparse visual SLAM and use camera poses and sparse keypoint depths as supervision alongside RGB keyframes. Since no prior training is required, our system flexibly fits to arbitrary scale and structure at runtime, and works even with strong specular reflections. We demonstrate reconstruction over a range of scenes from small indoor to large outdoor spaces. We also show that the method can straightforwardly benefit from additional inputs such as learned depth priors or semantic labels for more precise and advanced mapping.
Hidenobu Matsuki, Edgar Sucar, Tristan Laidlow, Kentaro Wada, Raluca Scona, Andrew J. Davison
ICRA6
2023 Feature-Realistic Neural Fusion for Real-Time, Open Set Scene Understanding
abstract
General scene understanding for robotics requires flexible semantic representation, so that novel objects and structures which may not have been known at training time can be identified, segmented and grouped. We present an algorithm which fuses general learned features from a standard pre-trained network into a highly efficient 3D geometric neural field representation during real-time SLAM. The fused 3D feature maps inherit the coherence of the neural field's geometry representation. This means that tiny amounts of human labelling interacting at runtime enable objects or even parts of objects to be robustly and accurately segmented in an open set manner. Project page: https://makezur.github.io/FeatureRealisticFusion/
Kirill Mazur, Edgar Sucar, Andrew J. Davison
ICRA3
2022 Simultaneous Localisation and Mapping With Quadric Surfaces
abstract
There are many possibilities for how to represent the map in simultaneous localisation and mapping (SLAM). While sparse, keypoint-based SLAM systems have achieved impressive levels of accuracy and robustness, their maps may not be suitable for many robotic tasks. Dense SLAM systems are capable of producing dense reconstructions, but can be computationally expensive and, like sparse systems, lack higher-level information about the structure of a scene. Human-made environments contain a lot of structure, and we seek to take advantage of this by enabling the use of quadric surfaces as features in SLAM systems. We introduce a minimal representation for quadric surfaces and show how this can be included in a least-squares formulation. We also show how our representation can be easily extended to include additional constraints on quadrics such as those found in quadrics of revolution. Finally, we introduce a proof-of-concept SLAM system using our representation, and provide some experimental results using an RGB-D dataset.
Tristan Laidlow, Andrew J. Davison
3DV2
2022 Coarse-to-Fine Q-attention: Efficient Learning for Visual Robotic Manipulation via Discretisation
abstract
We present a coarse-to-fine discretisation method that enables the use of discrete reinforcement learning approaches in place of unstable and data-inefficient actorcritic methods in continuous robotics domains. This approach builds on the recently released ARM algorithm, which replaces the continuous next-best pose agent with a discrete one, with coarse-to-fine Q-attention. Given a voxelised scene, coarse-to-fine Q-attention learns what part of the scene to ‘zoom’ into. When this ‘zooming’ behaviour is applied iteratively, it results in a near-lossless discretisation of the translation space, and allows the use of a discrete action, deep Q-learning method. We show that our new coarse-to-fine algorithm achieves state-of-the-art performance on several difficult sparsely rewarded RLBench vision-based robotics tasks, and can train real-world policies, tabula rasa, in a matter of minutes, with as little as 3 demonstrations.
Stephen James, Kentaro Wada, Tristan Laidlow, Andrew J. Davison
CVPR4
2022 Bootstrapping Semantic Segmentation with Regional Contrast
Shikun Liu, Shuaifeng Zhi, Edward Johns, Andrew J. Davison
ICLR4
2022 Incremental Abstraction in Distributed Probabilistic SLAM Graphs
abstract
Scene graphs represent the key components of a scene in a compact and semantically rich way, but are difficult to build during incremental SLAM operation because of the challenges of robustly identifying abstract scene elements and optimising continually changing, complex graphs. We present a distributed, graph-based SLAM framework for incrementally building scene graphs based on two novel components. First, we propose an incremental abstraction framework in which a neural network proposes abstract scene elements that are incorporated into the factor graph of a feature-based monocular SLAM system. Scene elements are confirmed or rejected through optimisation and incrementally replace the points yielding a more dense, semantic and compact representation. Second, enabled by our novel routing procedure, we use Gaussian Belief Propagation (GBP) for distributed inference on a graph processor. The time per iteration of GBP is structure-agnostic and we demonstrate the speed advantages over direct methods for inference of heterogeneous factor graphs. We run our system on real indoor datasets using planar abstractions and recover the major planes with significant compression.
Joseph Ortiz, Talfan Evans, Edgar Sucar, Andrew J. Davison
ICRA4
2022 ReorientBot: Learning Object Reorientation for Specific-Posed Placement
abstract
Robots need the capability of placing objects in arbitrary, specific poses to rearrange the world and achieve various valuable tasks. Object reorientation plays a crucial role in this as objects may not initially be oriented such that the robot can grasp and then immediately place them in a specific goal pose. In this work, we present a vision-based manipulation system, ReorientBot, which consists of 1) visual scene understanding with pose estimation and volumetric reconstruction using an onboard RGB-D camera; 2) learned waypoint selection for successful and efficient motion generation for reorientation; 3) traditional motion planning to generate a collision-free trajectory from the selected waypoints. We evaluate our method using the YCB objects in both simulation and the real world, achieving 93% overall success, 81% improvement in success rate, and 22% improvement in execution time compared to a heuristic approach. We demonstrate extended multi-object rearrangement showing the general capability of the system.
Kentaro Wada, Stephen James, Andrew J. Davison
ICRA3
2022 SafePicking: Learning Safe Object Extraction via Object-Level Mapping
abstract
Robots need object-level scene understanding to manipulate objects while reasoning about contact, support, and occlusion among objects. Given a pile of objects, object recognition and reconstruction can identify the boundary of object instances, giving important cues as to how the objects form and support the pile. In this work, we present a system, SafePicking, that integrates object-level mapping and learning-based motion planning to generate a motion that safely extracts occluded target objects from a pile. Planning is done by learning a deep Q-network that receives observations of predicted poses and a depth-based heightmap to output a motion trajectory, trained to maximize a safety metric reward. Our results show that the observation fusion of poses and depth-sensing gives both better performance and robustness to the model. We evaluate our methods using the YCB objects in both simulation and the real world, achieving safe object extraction from piles.
Kentaro Wada, Stephen James, Andrew J. Davison
ICRA3
2022 Learning to Complete Object Shapes for Object-level Mapping in Dynamic Scenes
abstract
In this paper, we propose a novel object-level mapping system that can simultaneously segment, track, and reconstruct objects in dynamic scenes. It can further predict and complete their full geometries by conditioning on reconstructions from depth inputs and a category-level shape prior with the aim that completed object geometry leads to better object reconstruction and tracking accuracy. For each incoming RGB-D frame, we perform instance segmentation to detect objects and build data associations between the detection and the existing object maps. A new object map will be created for each unmatched detection. For each matched object, we jointly optimise its pose and latent geometry representations using geometric residual and differential rendering residual towards its shape prior and completed geometry. Our approach shows better tracking and reconstruction performance compared to methods using traditional volumetric mapping or learned shape prior approaches. We evaluate its effectiveness by quantitatively and qualitatively testing it in both synthetic and real-world sequences.
Binbin Xu 0001, Andrew J. Davison, Stefan Leutenegger
IROS2
2022 Event-Based Vision: A Survey
abstract
Event cameras are bio-inspired sensors that differ from conventional frame cameras: Instead of capturing images at a fixed rate, they asynchronously measure per-pixel brightness changes, and output a stream of events that encode the time, location and sign of the brightness changes. Event cameras offer attractive properties compared to traditional cameras: high temporal resolution (in the order of μs), very high dynamic range (140 dB versus 60 dB), low power consumption, and high pixel bandwidth (on the order of kHz) resulting in reduced motion blur. Hence, event cameras have a large potential for robotics and computer vision in challenging scenarios for traditional cameras, such as low-latency, high speed, and high dynamic range. However, novel methods are required to process the unconventional output of these sensors in order to unlock their potential. This paper provides a comprehensive overview of the emerging field of event-based vision, with a focus on the applications and the algorithms developed to unlock the outstanding properties of event cameras. We present event cameras from their working principle, the actual sensors that are available and the tasks that they have been used for, from low-level vision (feature detection and tracking, optic flow, etc.) to high-level vision (reconstruction, segmentation, recognition). We also discuss the techniques developed to process events, including learning-based techniques, as well as specialized processors for these novel sensors, such as spiking neural networks. Additionally, we highlight the challenges that remain to be tackled and the opportunities that lie ahead in the search for a more efficient, bio-inspired way for machines to perceive and interact with the world.
Guillermo Gallego 0002, Tobi Delbruck, Garrick Orchard, Chiara Bartolozzi, Brian Taba, Andrea Censi, Stefan Leutenegger, Andrew J. Davison, Jörg Conradt, Kostas Daniilidis, Davide Scaramuzza 0001
IEEE Trans. Pattern Anal. Mach. Intell.8
2021 SIMstack: A Generative Shape and Instance Model for Unordered Object Stacks
abstract
By estimating 3D shape and instances from a single view, we can capture information about an environment quickly, without the need for comprehensive scanning and multi-view fusion. Solving this task for composite scenes (such as object stacks) is challenging: occluded areas are not only ambiguous in shape but also in instance segmentation; multiple decompositions could be valid. We observe that physics constrains decomposition as well as shape in occluded regions and hypothesise that a latent space learned from scenes built under physics simulation can serve as a prior to better predict shape and instances in occluded regions. To this end we propose SIMstack, a depth-conditioned Variational Auto-Encoder (VAE), trained on a dataset of objects stacked under physics simulation. We formulate instance segmentation as a centre voting task which allows for class-agnostic detection and doesn’t require setting the maximum number of objects in the scene. At test time, our model can generate 3D shape and instance segmentation from a single depth view, probabilistically sampling proposals for the occluded region from the learned latent space. Our method has practical applications in providing robots some of the ability humans have to make rapid intuitive inferences of partially observed scenes. We demonstrate an application for precise (non-disruptive) object grasping of unknown objects from a single depth view.
Zoe Landgraf, Raluca Scona, Tristan Laidlow, Stephen James, Stefan Leutenegger, Andrew J. Davison
ICCV6
2021 iMAP: Implicit Mapping and Positioning in Real-Time
abstract
We show for the first time that a multilayer perceptron (MLP) can serve as the only scene representation in a real-time SLAM system for a handheld RGB-D camera. Our network is trained in live operation without prior data, building a dense, scene-specific implicit 3D model of occupancy and colour which is also immediately used for tracking.Achieving real-time SLAM via continual training of a neural network against a live image stream requires significant innovation. Our iMAP algorithm uses a keyframe structure and multi-processing computation flow, with dynamic information-guided pixel sampling for speed, with tracking at 10 Hz and global map updating at 2 Hz. The advantages of an implicit MLP over standard dense SLAM techniques include efficient geometry representation with automatic detail control and smooth, plausible filling-in of unobserved regions such as the back surfaces of objects.
Edgar Sucar, Shikun Liu, Joseph Ortiz, Andrew J. Davison
ICCV4
2021 In-Place Scene Labelling and Understanding with Implicit Scene Representation
abstract
Semantic labelling is highly correlated with geometry and radiance reconstruction, as scene entities with similar shape and appearance are more likely to come from similar classes. Recent implicit neural reconstruction techniques are appealing as they do not require prior training data, but the same fully self-supervised approach is not possible for semantics because labels are human-defined properties.We extend neural radiance fields (NeRF) to jointly encode semantics with appearance and geometry, so that complete and accurate 2D semantic labels can be achieved using a small amount of in-place annotations specific to the scene. The intrinsic multi-view consistency and smoothness of NeRF benefit semantics by enabling sparse labels to efficiently propagate. We show the benefit of this approach when labels are either sparse or very noisy in room-scale scenes. We demonstrate its advantageous properties in various interesting applications such as an efficient scene labelling tool, novel semantic view synthesis, label denoising, super-resolution, label interpolation and multi-view semantic label fusion in visual semantic mapping systems.
Shuaifeng Zhi, Tristan Laidlow, Stefan Leutenegger, Andrew J. Davison
ICCV4
2021 End-to-End Egospheric Spatial Memory
Daniel Lenton, Stephen James, Ronald Clark, Andrew J. Davison
ICLR4
2020 NodeSLAM: Neural Object Descriptors for Multi-View Shape Reconstruction
abstract
The choice of scene representation is crucial in both the shape inference algorithms it requires and the smart applications it enables. We present efficient and optimisable multi-class learned object descriptors together with a novel probabilistic and differential rendering engine, for principled full object shape inference from one or more RGB-D images. Our framework allows for accurate and robust 3D object reconstruction which enables multiple applications including robot grasping and placing, augmented reality, and the first object-level SLAM system capable of optimising object poses and shapes jointly with camera trajectory.
Edgar Sucar, Kentaro Wada, Andrew J. Davison
3DV3
2020 Bundle Adjustment on a Graph Processor
Joseph Ortiz, Mark Pupilli, Stefan Leutenegger, Andrew J. Davison
CVPR4
2020 MoreFusion: Multi-object Reasoning for 6D Pose Estimation from Volumetric Fusion
abstract
Robots and other smart devices need efficient object-based scene representations from their on-board vision systems to reason about contact, physics and occlusion. Recognized precise object models will play an important role alongside non-parametric reconstructions of unrecognized structures. We present a system which can estimate the accurate poses of multiple known objects in contact and occlusion from real-time, embodied multi-view vision. Our approach makes 3D object pose proposals from single RGB-D views, accumulates pose estimates and non-parametric occupancy information from multiple views as the camera moves, and performs joint optimization to estimate consistent, non-intersecting poses for multiple objects in contact. We verify the accuracy and robustness of our approach experimentally on 2 object datasets: YCB-Video, and our own challenging Cluttered YCB-Video. We demonstrate a real-time robotics application where a robot arm precisely and orderly disassembles complicated piles of objects, using only on-board RGB-D vision.
Kentaro Wada, Edgar Sucar, Stephen James, Daniel Lenton, Andrew J. Davison
CVPR5
2020 Comparing View-Based and Map-Based Semantic Labelling in Real-Time SLAM
Zoe Landgraf, Fabian Falck, Michael Bloesch, Stefan Leutenegger, Andrew J. Davison
ICRA5
2019 End-To-End Multi-Task Learning With Attention
abstract
We propose a novel multi-task learning architecture, which allows learning of task-specific feature-level attention. Our design, the Multi-Task Attention Network (MTAN), consists of a single shared network containing a global feature pool, together with a soft-attention module for each task. These modules allow for learning of task-specific features from the global features, whilst simultaneously allowing for features to be shared across different tasks. The architecture can be trained end-to-end and can be built upon any feed-forward neural network, is simple to implement, and is parameter efficient. We evaluate our approach on a variety of datasets, across both image-to-image predictions and image classification tasks. We show that our architecture is state-of-the-art in multi-task learning compared to existing methods, and is also less sensitive to various weighting schemes in the multi-task loss function. Code is available at https://github.com/lorenmt/mtan.
Shikun Liu, Edward Johns, Andrew J. Davison
CVPR3
2019 SceneCode: Monocular Dense Semantic Reconstruction Using Learned Encoded Scene Representations
abstract
Systems which incrementally create 3D semantic maps from image sequences must store and update representations of both geometry and semantic entities. However, while there has been much work on the correct formulation for geometrical estimation, state-of-the-art systems usually rely on simple semantic representations which store and update independent label estimates for each surface element (depth pixels, surfels, or voxels). Spatial correlation is discarded, and fused label maps are incoherent and noisy. We introduce a new compact and optimisable semantic representation by training a variational auto-encoder that is conditioned on a colour image. Using this learned latent space, we can tackle semantic label fusion by jointly optimising the low-dimenional codes associated with each of a set of overlapping images, producing consistent fused label maps which preserve spatial correlation. We also show how this approach can be used within a monocular keyframe based semantic mapping system where a similar code approach is used for geometry. The probabilistic formulation allows a flexible formulation where we can jointly estimate motion, geometry and semantics in a unified optimisation.
Shuaifeng Zhi, Michael Bloesch, Stefan Leutenegger, Andrew J. Davison
CVPR4
2019 Learning Meshes for Dense Visual SLAM
abstract
Estimating motion and surrounding geometry of a moving camera remains a challenging inference problem. From an information theoretic point of view, estimates should get better as more information is included, such as is done in dense SLAM, but this is strongly dependent on the validity of the underlying models. In the present paper, we use triangular meshes as both compact and dense geometry representation. To allow for simple and fast usage, we propose a view-based formulation for which we predict the in-plane vertex coordinates directly from images and then employ the remaining vertex depth components as free variables. Flexible and continuous integration of information is achieved through the use of a residual based inference technique. This so-called factor graph encodes all information as mapping from free variables to residuals, the squared sum of which is minimised during inference. We propose the use of different types of learnable residuals, which are trained end-to-end to increase their suitability as information bearing models and to enable accurate and reliable estimation. Detailed evaluation of all components is provided on both synthetic and real data which confirms the practicability of the presented approach.
Michael Bloesch, Tristan Laidlow, Ronald Clark, Stefan Leutenegger, Andrew J. Davison
ICCV5
2019 SLAMBench 3.0: Systematic Automated Reproducible Evaluation of SLAM Systems for Robot Vision Challenges and Scene Understanding
abstract
As the SLAM research area matures and the number of SLAM systems available increases, the need for frameworks that can objectively evaluate them against prior work grows. This new version of SLAMBench moves beyond traditional visual SLAM, and provides new support for scene understanding and non-rigid environments (dynamic SLAM). More concretely for dynamic SLAM, SLAMBench 3.0 includes the first publicly available implementation of DynamicFusion, along with an evaluation infrastructure. In addition, we include two SLAM systems (one dense, one sparse) augmented with convolutional neural networks for scene understanding, together with datasets and appropriate metrics. Through a series of use-cases, we demonstrate the newly incorporated algorithms, visulation aids and metrics (6 new metrics, 4 new datasets and 5 new algorithms).
Mihai Bujanca, Paul Gafton, Sajad Saeedi G., Andy Nisbet, Bruno Bodin, Michael F. P. O'Boyle, Andrew J. Davison, Paul H. J. Kelly, Graham D. Riley, Barry Lennox, Mikel Luján, Steve Furber
ICRA7
2019 Characterizing Visual Localization and Mapping Datasets
abstract
Benchmarking mapping and motion estimation algorithms is established practice in robotics and computer vision. As the diversity of datasets increases, in terms of the trajectories, models, and scenes, it becomes a challenge to select datasets for a given benchmarking purpose. Inspired by the Wasserstein distance, this paper addresses this concern by developing novel metrics to evaluate trajectories and the environments without relying on any SLAM or motion estimation algorithm. The metrics, which so far have been missing in the research community, can be applied to the plethora of datasets that exist. Additionally, to improve the robotics SLAM benchmarking, the paper presents a new dataset for visual localization and mapping algorithms. A broad range of real-world trajectories is used in very high-quality scenes and a rendering framework to create a set of synthetic datasets with ground-truth trajectory and dense map which are representative of key SLAM applications such as virtual reality (VR), micro aerial vehicle (MAV) flight, and ground robotics.
Sajad Saeedi G., Eduardo D. C. Carvalho, Wenbin Li 0002, Dimos Tzoumanikas, Stefan Leutenegger, Paul H. J. Kelly, Andrew J. Davison
ICRA7
2019 MID-Fusion: Octree-based Object-Level Multi-Instance Dynamic SLAM
abstract
We propose a new multi-instance dynamic RGB-D SLAM system using an object-level octree-based volumetric representation. It can provide robust camera tracking in dynamic environments and at the same time, continuously estimate geometric, semantic, and motion properties for arbitrary objects in the scene. For each incoming frame, we perform instance segmentation to detect objects and refine mask boundaries using geometric and motion information. Meanwhile, we estimate the pose of each existing moving object using an object-oriented tracking method and robustly track the camera pose against the static scene. Based on the estimated camera pose and object poses, we associate segmented masks with existing models and incrementally fuse corresponding colour, depth, semantic, and foreground object probabilities into each object model. In contrast to existing approaches, our system is the first system to generate an object-level dynamic volumetric map from a single RGB-D camera, which can be used directly for robotic tasks. Our method can run at 2-3 Hz on a CPU, excluding the instance segmentation part. We demonstrate its effectiveness by quantitatively and qualitatively testing it on both synthetic and real-world sequences.
Binbin Xu 0001, Wenbin Li 0002, Dimos Tzoumanikas, Michael Bloesch, Andrew J. Davison, Stefan Leutenegger
ICRA5
2019 Self-Supervised Generalisation with Meta Auxiliary Learning
abstract
Learning with auxiliary tasks can improve the ability of a primary task to generalise. However, this comes at the cost of manually labelling auxiliary data. We propose a new method which automatically learns appropriate labels for an auxiliary task, such that any supervised learning task can be improved without requiring access to any further data. The approach is to train two neural networks: a label-generation network to predict the auxiliary labels, and a multi-task network to train the primary task alongside the auxiliary task. The loss for the label-generation network incorporates the loss of the multi-task network, and so this interaction between the two networks can be seen as a form of meta learning with a double gradient. We show that our proposed method, Meta AuXiliary Learning (MAXL), outperforms single-task learning on 7 image datasets, without requiring any additional data. We also show that MAXL outperforms several other baselines for generating auxiliary labels, and is even competitive when compared with human-defined auxiliary labels. The self-supervised nature of our method leads to a promising new direction towards automated generalisation. Source code can be found at \url{https://github.com/lorenmt/maxl}.
Shikun Liu, Andrew J. Davison, Edward Johns
NeurIPS2
2018 Fusion++: Volumetric Object-Level SLAM
abstract
We propose an online object-level SLAM system which builds a persistent and accurate 3D graph map of arbitrary reconstructed objects. As an RGB-D camera browses a cluttered indoor scene, Mask-RCNN instance segmentations are used to initialise compact per-object Truncated Signed Distance Function (TSDF) reconstructions with object size-dependent resolutions and a novel 3D foreground mask. Reconstructed objects are stored in an optimisable 6DoF pose graph which is our only persistent map representation. Objects are incrementally refined via depth fusion, and are used for tracking, relocalisation and loop closure detection. Loop closures cause adjustments in the relative pose estimates of object instances, but no intra-object warping. Each object also carries semantic information which is refined over time and an existence probability to account for spurious instance predictions. We demonstrate our approach on a hand-held RGB-D sequence from a cluttered office scene with a large number and variety of object instances, highlighting how the system closes loops and makes good use of existing objects on repeated loops. We quantitatively evaluate the trajectory error of our system against a baseline approach on the RGB-D SLAM benchmark, and qualitatively compare reconstruction quality of discovered objects on the YCB video dataset. Performance evaluation shows our approach is highly memory efficient and runs online at 4-8Hz (excluding relocalisation) despite not being optimised at the software level.
John McCormac, Ronald Clark, Michael Bloesch, Andrew J. Davison, Stefan Leutenegger
3DV4
2018 CodeSLAM - Learning a Compact, Optimisable Representation for Dense Visual SLAM
abstract
The representation of geometry in real-time 3D perception systems continues to be a critical research issue. Dense maps capture complete surface shape and can be augmented with semantic labels, but their high dimensionality makes them computationally costly to store and process, and unsuitable for rigorous probabilistic inference. Sparse feature-based representations avoid these problems, but capture only partial scene information and are mainly useful for localisation only. We present a new compact but dense representation of scene geometry which is conditioned on the intensity data from a single image and generated from a code consisting of a small number of parameters. We are inspired by work both on learned depth from images, and auto-encoders. Our approach is suitable for use in a keyframe-based monocular dense SLAM system: While each keyframe with a code can produce a depth map, the code can be optimised efficiently jointly with pose variables and together with the codes of overlapping keyframes to attain global consistency. Conditioning the depth map on the image allows the code to only represent aspects of the local geometry which cannot directly be predicted from the image. We explain how to learn our code representation, and demonstrate its advantageous properties in monocular SLAM.
Michael Bloesch, Jan Czarnowski, Ronald Clark, Stefan Leutenegger, Andrew J. Davison
CVPR5
2018 Learning to Solve Nonlinear Least Squares for Monocular Stereo
Ronald Clark, Michael Bloesch, Jan Czarnowski, Stefan Leutenegger, Andrew J. Davison
ECCV (8)5
2018 SLAMBench2: Multi-Objective Head-to-Head Benchmarking for Visual SLAM
abstract
SLAM is becoming a key component of robotics and augmented reality (AR) systems. While a large number of SLAM algorithms have been presented, there has been little effort to unify the interface of such algorithms, or to perform a holistic comparison of their capabilities. This is a problem since different SLAM applications can have different functional and non-functional requirements. For example, a mobile phone-based AR application has a tight energy budget, while a UAV navigation system usually requires high accuracy. SLAMBench2 is a benchmarking framework to evaluate existing and future SLAM systems, both open and close source, over an extensible list of datasets, while using a comparable and clearly specified list of performance metrics. A wide variety of existing SLAM algorithms and datasets is supported, e.g. ElasticFusion, InfiniTAM, ORB-SLAM2, OKVIS, and integrating new ones is straightforward and clearly specified by the framework. SLAMBench2 is a publicly-available software framework which represents a starting point for quantitative, comparable and val-idatable experimental research to investigate trade-offs across SLAM systems.
Bruno Bodin, Harry Wagstaff, Sajad Saeedi G., Luigi Nardi, Emanuele Vespa, John Mawer, Andy Nisbet, Mikel Luján, Steve Furber, Andrew J. Davison, Paul H. J. Kelly, Michael F. P. O'Boyle
ICRA10
2018 ViCTree: an automated framework for taxonomic classification from protein sequences
abstract
Motivation: The increasing rate of submission of genetic sequences into public databases is providing a growing resource for classifying the organisms that these sequences represent. To aid viral classification, we have developed ViCTree, which automatically integrates the relevant sets of sequences in NCBI GenBank and transforms them into an interactive maximum likelihood phylogenetic tree that can be updated automatically. ViCTree incorporates ViCTreeView, which is a JavaScript-based visualization tool that enables the tree to be explored interactively in the context of pairwise distance data. Results: To demonstrate utility, ViCTree was applied to subfamily Densovirinae of family Parvoviridae. This led to the identification of six new species of insect virus. Availability and implementation: ViCTree is open-source and can be run on any Linux- or Unix-based computer or cluster. A tutorial, the documentation and the source code are available under a GPL3 license, and can be accessed at http://bioinformatics.cvr.ac.uk/victree_web/. Supplementary information: Supplementary data are available at Bioinformatics online.
Sejal Modha, Anil S. Thanki, Susan F. Cotmore, Andrew J. Davison, Joseph Hughes
Bioinform.4
2018 Navigating the Landscape for Real-Time Localization and Mapping for Robotics and Virtual and Augmented Reality
abstract
Visual understanding of 3-D environments in real time, at low power, is a huge computational challenge. Often referred to as simultaneous localization and mapping (SLAM), it is central to applications spanning domestic and industrial robotics, autonomous vehicles, and virtual and augmented reality. This paper describes the results of a major research effort to assemble the algorithms, architectures, tools, and systems software needed to enable delivery of SLAM, by supporting applications specialists in selecting and configuring the appropriate algorithm and the appropriate hardware, and compilation pathway, to meet their performance, accuracy, and energy consumption goals. The major contributions we present are: 1) tools and methodology for systematic quantitative evaluation of SLAM algorithms; 2) automated, machine-learning-guided exploration of the algorithmic and implementation design space with respect to multiple objectives; 3) end-to-end simulation tools to enable optimization of heterogeneous, accelerated architectures for the specific algorithmic requirements of the various SLAM algorithmic approaches; and 4) tools for delivering, where appropriate, accelerated, adaptive SLAM solutions in a managed, JIT-compiled, adaptive runtime context.
Sajad Saeedi G., Bruno Bodin, Harry Wagstaff, Andy Nisbet, Luigi Nardi, John Mawer, Nicolas Melot, Oscar Palomar, Emanuele Vespa, Tom Spink, Cosmin Gorgovan, Andrew M. Webb 0002, James Clarkson, Erik Tomusk, Thomas Debrunner, Kuba Kaszyk, Pablo González de Aledo Marugán, Andrey Rodchenko, Graham D. Riley, Christos Kotselidis, Björn Franke, Michael F. P. O'Boyle, Andrew J. Davison, Paul H. J. Kelly, Mikel Luján, Steve Furber
Proc. IEEE23
2017 SceneNet RGB-D: Can 5M Synthetic Images Beat Generic ImageNet Pre-training on Indoor Segmentation?
abstract
We introduce SceneNet RGB-D, a dataset providing pixel-perfect ground truth for scene understanding problems such as semantic segmentation, instance segmentation, and object detection. It also provides perfect camera poses and depth data, allowing investigation into geometric computer vision problems such as optical flow, camera pose estimation, and 3D scene labelling tasks. Random sampling permits virtually unlimited scene configurations, and here we provide 5M rendered RGB-D images from 16K randomly generated 3D trajectories in synthetic layouts, with random but physically simulated object configurations. We compare the semantic segmentation performance of network weights produced from pretraining on RGB images from our dataset against generic VGG-16 ImageNet weights. After fine-tuning on the SUN RGB-D and NYUv2 real-world datasets we find in both cases that the synthetically pre-trained network outperforms the VGG-16 weights. When synthetic pre-training includes a depth channel (something ImageNet cannot natively provide) the performance is greater still. This suggests that large-scale high-quality synthetic RGB datasets with task-specific labels can be more useful for pretraining than real-world generic pre-training such as ImageNet. We host the dataset at http://robotvault. bitbucket.io/scenenet-rgbd.html.
John McCormac, Ankur Handa, Stefan Leutenegger, Andrew J. Davison
ICCV4
2017 Room layout estimation from rapid omnidirectional exploration
abstract
A new generation of practical, low-cost indoor robots is now using wide-angle cameras to aid navigation, but usually this is limited to position estimation via sparse feature-based SLAM. Such robots usually have little global sense of the dimensions, demarcation or identities of the rooms they are in, information which would be very useful to enable behaviour with much more high level intelligence. In this paper we show that we can augment an omni-directional SLAM pipeline with straightforward dense stereo estimation and simple and robust room model fitting to obtain rapid and reliable estimation of the global shape of typical rooms from short robot motions. We have tested our method extensively in real homes, offices and on synthetic data. We also give examples of how our method can extend to making composite maps of larger rooms, and detecting room transitions.
Robert Lukierski, Stefan Leutenegger, Andrew J. Davison
ICRA3
2017 SemanticFusion: Dense 3D semantic mapping with convolutional neural networks
abstract
Ever more robust, accurate and detailed mapping using visual sensing has proven to be an enabling factor for mobile robots across a wide variety of applications. For the next level of robot intelligence and intuitive user interaction, maps need to extend beyond geometry and appearance - they need to contain semantics. We address this challenge by combining Convolutional Neural Networks (CNNs) and a state-of-the-art dense Simultaneous Localization and Mapping (SLAM) system, ElasticFusion, which provides long-term dense correspondences between frames of indoor RGB-D video even during loopy scanning trajectories. These correspondences allow the CNN's semantic predictions from multiple view points to be probabilistically fused into a map. This not only produces a useful semantic 3D map, but we also show on the NYUv2 dataset that fusing multiple predictions leads to an improvement even in the 2D semantic labelling over baseline single frame predictions. We also show that for a smaller reconstruction dataset with larger variation in prediction viewpoint, the improvement over single frame segmentation increases. Our system is efficient enough to allow real-time interactive use at frame-rates of ≈25Hz.
John McCormac, Ankur Handa, Andrew J. Davison, Stefan Leutenegger
ICRA3
2017 Monocular visual odometry: Sparse joint optimisation or dense alternation?
abstract
Real-time monocular SLAM is increasingly mature and entering commercial products. However, there is a divide between two techniques providing similar performance. Despite the rise of ‘dense’ and ‘semi-dense’ methods which use large proportions of the pixels in a video stream to estimate motion and structure via alternating estimation, they have not eradicated feature-based methods which use a significantly smaller amount of image information from keypoints and retain a more rigorous joint estimation framework. Dense methods provide more complete scene information, but in this paper we focus on how the amount of information and different optimisation methods affect the accuracy of local motion estimation (monocular visual odometry). This topic becomes particularly relevant after the recent results from a direct sparse system. We propose a new method for fairly comparing the accuracy of SLAM frontends in a common setting. We suggest computational cost models for an overall comparison which indicates that there is relative parity between the approaches at the settings allowed by current serial processors when evaluated under equal conditions.
Lukas Platinsky, Andrew J. Davison, Stefan Leutenegger
ICRA2
2017 Application-oriented design space exploration for SLAM algorithms
abstract
In visual SLAM, there are many software and hardware parameters, such as algorithmic thresholds and GPU frequency, that need to be tuned; however, this tuning should also take into account the structure and motion of the camera. In this paper, we determine the complexity of the structure and motion with a few parameters calculated using information theory. Depending on this complexity and the desired performance metrics, suitable parameters are explored and determined. Additionally, based on the proposed structure and motion parameters, several applications are presented, including a novel active SLAM approach which guides the camera in such a way that the SLAM algorithm achieves the desired performance metrics. Real-world and simulated experimental results demonstrate the effectiveness of the proposed design space and its applications.
Sajad Saeedi G., Luigi Nardi, Edward Johns, Bruno Bodin, Paul H. J. Kelly, Andrew J. Davison
ICRA6
2017 Near-lighting Photometric Stereo for unknown scene distance and medium attenuation
Chourmouzios Tsiotsios, Andrew J. Davison, Tae-Kyun Kim 0001
Image Vis. Comput.2
2016 Monocular, Real-Time Surface Reconstruction Using Dynamic Level of Detail
abstract
We present a scalable, real-time capable method for robust surface reconstruction that explicitly handles multiple scales. As a monocular camera browses a scene, our algorithm processes images as they arrive and incrementally builds a detailed surface model.While most of the existing reconstruction approaches rely on volumetric or point-cloud representations of the environment, we perform depth-map and colour fusion directly into a multi-resolution triangular mesh that can be adaptively tessellated using the concept of Dynamic Level of Detail. Our method relies on least-squares optimisation, which enables a probabilistically sound and principled formulation of the fusion algorithm.We demonstrate that our method is capable of obtaining high quality, close-up reconstruction, as well as capturing overall scene geometry, while being memory and computationally efficient.
Jacek Zienkiewicz, Akis Tsiotsios, Andrew J. Davison, Stefan Leutenegger
3DV3
2016 Simultaneous Optical Flow and Intensity Estimation from an Event Camera
abstract
Event cameras are bio-inspired vision sensors which mimic retinas to measure per-pixel intensity change rather than outputting an actual intensity image. This proposed paradigm shift away from traditional frame cameras offers significant potential advantages: namely avoiding high data rates, dynamic range limitations and motion blur. Unfortunately, however, established computer vision algorithms may not at all be applied directly to event cameras. Methods proposed so far to reconstruct images, estimate optical flow, track a camera and reconstruct a scene come with severe restrictions on the environment or on the motion of the camera, e.g. allowing only rotation. Here, we propose, to the best of our knowledge, the first algorithm to simultaneously recover the motion field and brightness image, while the camera undergoes a generic motion through any scene. Our approach employs minimisation of a cost function that contains the asynchronous event data as well as spatial and temporal regularisation within a sliding window time interval. Our implementation relies on GPU optimisation and runs in near real-time. In a series of examples, we demonstrate the successful operation of our framework, including in situations where conventional cameras suffer from dynamic range limitations and motion blur.
Patrick Bardow, Andrew J. Davison, Stefan Leutenegger
CVPR2
2016 Pairwise Decomposition of Image Sequences for Active Multi-view Recognition
abstract
A multi-view image sequence provides a much richer capacity for object recognition than from a single image. However, most existing solutions to multi-view recognition typically adopt hand-crafted, model-based geometric methods, which do not readily embrace recent trends in deep learning. We propose to bring Convolutional Neural Networks to generic multi-view recognition, by decomposing an image sequence into a set of image pairs, classifying each pair independently, and then learning an object classifier by weighting the contribution of each pair. This allows for recognition over arbitrary camera trajectories, without requiring explicit training over the potentially infinite number of camera paths and lengths. Building these pairwise relationships then naturally extends to the next-best-view problem in an active recognition framework. To achieve this, we train a second Convolutional Neural Network to map directly from an observed image to next viewpoint. Finally, we incorporate this into a trajectory optimisation task, whereby the best recognition confidence is sought for a given trajectory length. We present state-of-the-art results in both guided and unguided multi-view recognition on the ModelNet dataset, and show how our method can be used with depth images, greyscale images, or both.
Edward Johns, Stefan Leutenegger, Andrew J. Davison
CVPR3
2016 Real-Time 3D Reconstruction and 6-DoF Tracking with an Event Camera
Hanme Kim, Stefan Leutenegger, Andrew J. Davison
ECCV (6)3
2016 Comparative design space exploration of dense and semi-dense SLAM
abstract
SLAM has matured significantly over the past few years, and is beginning to appear in serious commercial products. While new SLAM systems are being proposed at every conference, evaluation is often restricted to qualitative visualizations or accuracy estimation against a ground truth. This is due to the lack of benchmarking methodologies which can holistically and quantitatively evaluate these systems. Further investigation at the level of individual kernels and parameter spaces of SLAM pipelines is non-existent, which is absolutely essential for systems research and integration. We extend the recently introduced SLAMBench framework to allow comparing two state-of-the-art SLAM pipelines, namely KinectFusion and LSD-SLAM, along the metrics of accuracy, energy consumption, and processing frame rate on two different hardware platforms, namely a desktop and an embedded device. We also analyze the pipelines at the level of individual kernels and explore their algorithmic and hardware design spaces for the first time, yielding valuable insights.
M. Zeeshan Zia, Luigi Nardi, Andrew Jack, Emanuele Vespa, Bruno Bodin, Paul H. J. Kelly, Andrew J. Davison
ICRA7
2016 Deep learning a grasp function for grasping under gripper pose uncertainty
abstract
This paper presents a new method for parallel-jaw grasping of isolated objects from depth images, under large gripper pose uncertainty. Whilst most approaches aim to predict the single best grasp pose from an image, our method first predicts a score for every possible grasp pose, which we denote the grasp function. With this, it is possible to achieve grasping robust to the gripper's pose uncertainty, by smoothing the grasp function with the pose uncertainty function. Therefore, if the single best pose is adjacent to a region of poor grasp quality, that pose will no longer be chosen, and instead a pose will be chosen which is surrounded by a region of high grasp quality. To learn this function, we train a Convolutional Neural Network which takes as input a single depth image of an object, and outputs a score for each grasp pose across the image. Training data for this is generated by use of physics simulation and depth image simulation with 3D object meshes, to enable acquisition of sufficient data without requiring exhaustive real-world experiments. We evaluate with both synthetic and real experiments, and show that the learned grasp score is more robust to gripper pose uncertainty than when this uncertainty is not accounted for.
Edward Johns, Stefan Leutenegger, Andrew J. Davison
IROS3
2016 Real-time height map fusion using differentiable rendering
abstract
We present a robust real-time method which performs dense reconstruction of high quality height maps from monocular video. By representing the height map as a triangular mesh, and using efficient differentiable rendering approach, our method enables rigorous incremental probabilistic fusion of standard locally estimated depth and colour into an immediately usable dense model. We present results for the application of free space and obstacle mapping by a low-cost robot, showing that detailed maps suitable for autonomous navigation can be obtained using only a single forward-looking camera.
Jacek Zienkiewicz, Andrew J. Davison, Stefan Leutenegger
IROS2
2016 Model effectiveness prediction and system adaptation for photometric stereo in murky water
Chourmouzios Tsiotsios, Tae-Kyun Kim 0001, Andrew J. Davison, Srinivasa G. Narasimhan
Comput. Vis. Image Underst.3
2015 Introducing SLAMBench, a performance and accuracy benchmarking methodology for SLAM
abstract
Real-time dense computer vision and SLAM offer great potential for a new level of scene modelling, tracking and real environmental interaction for many types of robot, but their high computational requirements mean that use on mass market embedded platforms is challenging. Meanwhile, trends in low-cost, low-power processing are towards massive parallelism and heterogeneity, making it difficult for robotics and vision researchers to implement their algorithms in a performance-portable way. In this paper we introduce SLAMBench, a publicly-available software framework which represents a starting point for quantitative, comparable and validatable experimental research to investigate trade-offs in performance, accuracy and energy consumption of a dense RGB-D SLAM system. SLAMBench provides a KinectFusion implementation in C++, OpenMP, OpenCL and CUDA, and harnesses the ICL-NUIM dataset of synthetic RGB-D sequences with trajectory and scene ground truth for reliable accuracy comparison of different implementation and algorithms. We present an analysis and breakdown of the constituent algorithmic elements of KinectFusion, and experimentally investigate their execution time on a variety of multicore and GPU-accelerated platforms. For a popular embedded platform, we also present an analysis of energy efficiency for different configuration alternatives.
Luigi Nardi, Bruno Bodin, M. Zeeshan Zia, John Mawer, Andy Nisbet, Paul H. J. Kelly, Andrew J. Davison, Mikel Luján, Michael F. P. O'Boyle, Graham D. Riley, Nigel P. Topham, Steve Furber
ICRA7
2014 Simultaneous Mosaicing and Tracking with an Event Camera
Hanme Kim, Ankur Handa, Ryad Benosman, Sio-Hoi Ieng, Andrew J. Davison
BMVC5
2014 Backscatter Compensated Photometric Stereo with 3 Sources
abstract
Photometric stereo offers the possibility of object shape reconstruction via reasoning about the amount of light reflected from oriented surfaces. However, in murky media such as sea water, the illuminating light interacts with the medium and some of it is backscattered towards the camera. Due to this additive light component, the standard Photometric Stereo equations lead to poor quality shape estimation. Previous authors have attempted to reformulate the approach but have either neglected backscatter entirely or disregarded its non-uniformity on the sensor when camera and lights are close to each other. We show that by compensating effectively for the backscatter component, a linear formulation of Photometric Stereo is allowed which recovers an accurate normal map using only 3 lights. Our backscatter compensation method for point-sources can be used for estimating the uneven backscatter directly from single images without any prior knowledge about the characteristics of the medium or the scene. We compare our method with previous approaches through extensive experimental results, where a variety of objects are imaged in a big water tank whose turbidity is systematically increased, and show reconstruction quality which degrades little relative to clean water results even with a very significant scattering level.
Chourmouzios Tsiotsios, Maria E. Angelopoulou, Tae-Kyun Kim 0001, Andrew J. Davison
CVPR4
2014 A benchmark for RGB-D visual odometry, 3D reconstruction and SLAM
abstract
We introduce the Imperial College London and National University of Ireland Maynooth (ICL-NUIM) dataset for the evaluation of visual odometry, 3D reconstruction and SLAM algorithms that typically use RGB-D data. We present a collection of handheld RGB-D camera sequences within synthetically generated environments. RGB-D sequences with perfect ground truth poses are provided as well as a ground truth surface model that enables a method of quantitatively evaluating the final map or surface reconstruction accuracy. Care has been taken to simulate typically observed real-world artefacts in the synthetic imagery by modelling sensor noise in both RGB and depth data. While this dataset is useful for the evaluation of visual odometry and SLAM trajectory estimation, our main focus is on providing a method to benchmark the surface reconstruction accuracy which to date has been missing in the RGB-D community despite the plethora of ground truth RGB-D datasets available.
Ankur Handa, Thomas Whelan, John McDonald 0001, Andrew J. Davison
ICRA4
2014 Dense planar SLAM
abstract
Using higher-level entities during mapping has the potential to improve camera localisation performance and give substantial perception capabilities to real-time 3D SLAM systems. We present an efficient new real-time approach which densely maps an environment using bounded planes and surfels extracted from depth images (like those produced by RGB-D sensors or dense multi-view stereo reconstruction). Our method offers the every-pixel descriptive power of the latest dense SLAM approaches, but takes advantage directly of the planarity of many parts of real-world scenes via a data-driven process to directly regularize planar regions and represent their accurate extent efficiently using an occupancy approach with on-line compression. Large areas can be mapped efficiently and with useful semantic planar structure which enables intuitive and useful AR applications such as using any wall or other planar surface in a scene to display a user's content.
Renato F. Salas-Moreno, Ben Glocker, Paul H. J. Kelly, Andrew J. Davison
ISMAR4
2014 Dense planar SLAM
abstract
Using higher-level entities during mapping has the potential to improve camera localisation performance and give substantial perception capabilities to real-time 3D SLAM systems. We present an efficient new real-time approach which densely maps an environment using bounded planes and surfels extracted from depth images (like those produced by RGB-D sensors or dense multi-view stereo reconstruction). Our method offers the every-pixel descriptive power of the latest dense SLAM approaches, but takes advantage directly of the planarity of many parts of real-world scenes via a data-driven process to directly regularize planar regions and represent their accurate extent efficiently using an occupancy approach with on-line compression. Large areas can be mapped efficiently and with useful semantic planar structure which enables intuitive and useful AR applications such as using any wall or other planar surface in a scene to display a user's content.
Renato F. Salas-Moreno, Ben Glocker, Paul H. J. Kelly, Andrew J. Davison
ISMAR4
2013 Dense, Auto-Calibrating Visual Odometry from a Downward-Looking Camera
abstract
We present a technique whereby a single camera can be used as a high precision visual odometry sensor in a range of practical settings using simple, computationally efficient techniques.Taking advantage of the local planarity of common floor surfaces, we use real-time dense alignment of a 30Hz video stream as the camera looks down from a fast-moving robot, making use of the whole texture available rather than sparse feature points.Our key novelty, and crucial to the practicality of this approach, is rapid and automatic calibration for 6DoF camera extrinsics relative to the robot frame.Our experiments show robust performance over a range of low-textured real surfaces.
Jacek Zienkiewicz, Robert Lukierski, Andrew J. Davison
BMVC3
2013 SLAM++: Simultaneous Localisation and Mapping at the Level of Objects
abstract
We present the major advantages of a new 'object oriented' 3D SLAM paradigm, which takes full advantage in the loop of prior knowledge that many scenes consist of repeated, domain-specific objects and structures. As a hand-held depth camera browses a cluttered scene, real-time 3D object recognition and tracking provides 6DoF camera-object constraints which feed into an explicit graph of objects, continually refined by efficient pose-graph optimisation. This offers the descriptive and predictive power of SLAM systems which perform dense surface reconstruction, but with a huge representation compression. The object graph enables predictions for accurate ICP-based camera to model tracking at each live frame, and efficient active search for new objects in currently undescribed image regions. We demonstrate real-time incremental SLAM in large, cluttered environments, including loop closure, relocalisation and the detection of moved objects, and of course the generation of an object level scene description with the potential to enable interaction.
Renato F. Salas-Moreno, Richard A. Newcombe, Hauke Strasdat, Paul H. J. Kelly, Andrew J. Davison
CVPR5
2013 Real-Time Dense Stereo Reconstruction Using Convex Optimisation with a Cost-Volume for Image-Guided Robotic Surgery
Ping-Lin Chang, Danail Stoyanov, Andrew J. Davison, Philip J. Edwards
MICCAI (1)3
2013 Gauge-SURF descriptors
Pablo Fernández Alcantarilla, Luis Miguel Bergasa, Andrew J. Davison
Image Vis. Comput.3
2012 KAZE Features
Pablo Fernández Alcantarilla, Adrien Bartoli, Andrew J. Davison
ECCV (6)3
2012 Real-Time Camera Tracking: When is High Frame-Rate Best?
Ankur Handa, Richard A. Newcombe, Adrien Angeli, Andrew J. Davison
ECCV (7)4
2012 Real-time surface light-field capture for augmentation of planar specular surfaces
abstract
A single hand-held camera provides an easily accessible but potentially extremely powerful setup for augmented reality. Capabilities which previously required expensive and complicated infrastructure have gradually become possible from a live monocular video feed, such as accurate camera tracking and, most recently, dense 3D scene reconstruction. A new frontier is to work towards recovering the reflectance properties of general surfaces and the lighting configuration in a scene without the need for probes, omni-directional cameras or specialised light-field cameras. Specular lighting phenomena cause effects in a video stream which can lead current tracking and reconstruction algorithms to fail. However, the potential exists to measure and use these effects to estimate deeper physical details about an environment, enabling advanced scene understanding and more convincing AR. In this paper we present an algorithm for real-time surface light-field capture from a single hand-held camera, which is able to capture dense illumination information for general specular surfaces. Our system incorporates a guidance mechanism to help the user interactively during capture. We then split the light-field into its diffuse and specular components, and show that the specular component can be used for estimation of an environment map. This enables the convincing placement of an augmentation on a specular surface such as a shiny book, with realistic synthesized shadow, reflection and occlusion of specularities as the viewpoint changes. Our method currently works for planar scenes, but the surface light-field representation makes it ideal for future combination with dense 3D reconstruction methods.
Jan Jachnik, Richard A. Newcombe, Andrew J. Davison
ISMAR3
2012 Visual SLAM: Why filter?
Hauke Strasdat, J. M. M. Montiel, Andrew J. Davison
Image Vis. Comput.3
2011 DTAM: Dense tracking and mapping in real-time
abstract
DTAM is a system for real-time camera tracking and reconstruction which relies not on feature extraction but dense, every pixel methods. As a single hand-held RGB camera flies over a static scene, we estimate detailed textured depth maps at selected keyframes to produce a surface patchwork with millions of vertices. We use the hundreds of images available in a video stream to improve the quality of a simple photometric data term, and minimise a global spatially regularised energy functional in a novel non-convex optimisation framework. Interleaved, we track the camera's 6DOF motion precisely by frame-rate whole image alignment against the entire dense model. Our algorithms are highly parallelisable throughout and DTAM achieves real-time performance using current commodity GPU hardware. We demonstrate that a dense model permits superior tracking performance under rapid motion compared to a state of the art method using features; and also show the additional usefulness of the dense model for real-time scene interaction in a physics-enhanced augmented reality application.
Richard A. Newcombe, Steven Lovegrove, Andrew J. Davison
ICCV3
2011 Double window optimisation for constant time visual SLAM
abstract
We present a novel and general optimisation framework for visual SLAM, which scales for both local, highly accurate reconstruction and large-scale motion with long loop closures. We take a two-level approach that combines accurate pose-point constraints in the primary region of interest with a stabilising periphery of pose-pose soft constraints. Our algorithm automatically builds a suitable connected graph of keyposes and constraints, dynamically selects inner and outer window membership and optimises both simultaneously. We demonstrate in extensive simulation experiments that our method approaches the accuracy of offline bundle adjustment while maintaining constant-time operation, even in the hard case of very loopy monocular camera motion. Furthermore, we present a set of real experiments for various types of visual sensor and motion, including large scale SLAM with both monocular and stereo cameras, loopy local browsing with either monocular or RGB-D cameras, and dense RGB-D object model building.
Hauke Strasdat, Andrew J. Davison, J. M. M. Montiel, Kurt Konolige
ICCV2
2011 SLAM-based automatic extrinsic calibration of a multi-camera rig
abstract
Cameras are often a good choice as the primary outward-looking sensor for mobile robots, and a wide field of view is usually desirable for responsive and accurate navigation, SLAMand relocalisation. While this can potentially be provided by a single omnidirectional camera, it can also be flexibly achieved by multiple cameras with standard optics mounted around the robot. However, such setups are difficult to calibrate. Here we present a general method for fully automatic extrinsic auto-calibration of a fixed multi camera rig, with no requirement for calibration patterns or other infrastructure, which works even in the case where the cameras have completely non-overlapping views. The robot is placed in a natural environment and makes a set of programmed movements including a full horizontal rotation and captures a synchronized image sequence from each camera. These sequences are processed individually with a monocular visual SLAM algorithm. The resulting maps are matched and fused robustly based on corresponding invariant features, and then all estimates are optimised full joint bundle adjustment, where we constrain the relative poses of the cameras to be fixed. We present results showing accurate performance of the method for various two and four camera configurations.
Gerardo Carrera, Adrien Angeli, Andrew J. Davison
ICRA3
2011 KinectFusion: Real-time dense surface mapping and tracking
abstract
We present a system for accurate real-time mapping of complex and arbitrary indoor scenes in variable lighting conditions, using only a moving low-cost depth camera and commodity graphics hardware. We fuse all of the depth data streamed from a Kinect sensor into a single global implicit surface model of the observed scene in real-time. The current sensor pose is simultaneously obtained by tracking the live depth frame relative to the global model using a coarse-to-fine iterative closest point (ICP) algorithm, which uses all of the observed depth data available. We demonstrate the advantages of tracking against the growing full surface model compared with frame-to-frame tracking, obtaining tracking and mapping results in constant time within room sized scenes with limited drift and high accuracy. We also show both qualitative and quantitative results relating to various aspects of our tracking and mapping system. Modelling of natural scenes, in real-time with only commodity sensor and GPU hardware, promises an exciting step forward in augmented reality (AR), in particular, it allows dense surfaces to be reconstructed in real-time, with a level of detail and robustness beyond any solution yet presented using passive computer vision.
Richard A. Newcombe, Shahram Izadi, Otmar Hilliges, David Molyneaux, David Kim 0002, Andrew J. Davison, Pushmeet Kohli, Jamie Shotton, Steve Hodges 0001, Andrew W. Fitzgibbon
ISMAR6
2011 Accurate visual odometry from a rear parking camera
abstract
As an increasing number of automatic safety and navigation features are added to modern vehicles, the crucial job of providing real-time localisation is predominantly performed by a single sensor, GPS, despite its well-known failings, particularly in urban environments. Various attempts have been made to supplement GPS to improve localisation performance, but these usually require additional specialised and expensive sensors. Offering increased value to vehicle OEMs, we show that it is possible to use just the video stream from a rear parking camera to produce smooth and locally accurate visual odometry in real-time. We use an efficient whole image alignment approach based on ESM, taking account of both the difficulties and advantages of the fact that a parking camera views only the road surface directly behind a vehicle. Visual odometry is complementary to GPS in offering localisation information at 30 Hz which is smooth and highly accurate locally whilst GPS is course but offers absolute measurements. We demonstrate our system in a large scale experiment covering real urban driving. We also present real-time fusion of our visual estimation with automotive GPS to generate a commodity-cost localisation solution which is smooth, accurate and drift free in global coordinates.
Steven Lovegrove, Andrew J. Davison, Javier Ibañez-Guzmán
Intelligent Vehicles Symposium2
2011 KinectFusion: real-time 3D reconstruction and interaction using a moving depth camera
abstract
KinectFusion enables a user holding and moving a standard Kinect camera to rapidly create detailed 3D reconstructions of an indoor scene. Only the depth data from Kinect is used to track the 3D pose of the sensor and reconstruct, geometrically precise, 3D models of the physical scene in real-time. The capabilities of KinectFusion, as well as the novel GPU-based pipeline are described in full. Uses of the core system for low-cost handheld scanning, and geometry-aware augmented reality and physics-based interactions are shown. Novel extensions to the core GPU pipeline demonstrate object segmentation and user interaction directly in front of the sensor, without degrading camera tracking or reconstruction. These extensions are used to enable real-time multi-touch interactions anywhere, allowing any planar or non-planar reconstructed physical surface to be appropriated for touch.
Shahram Izadi, David Kim 0002, Otmar Hilliges, David Molyneaux, Richard A. Newcombe, Pushmeet Kohli, Jamie Shotton, Steve Hodges 0001, Dustin Freeman, Andrew J. Davison, Andrew W. Fitzgibbon
UIST10
2010 Live Feature Clustering in Video Using Appearance and 3D Geometry
abstract
We present a method for live grouping of feature points into persistent 3D clusters as a single camera browses a static scene, with no additional assumptions, training or infrastructure required. The clusters produced depend both on similar appearance and on 3D proximity information derived from real-time structure from motion, and clustering proceeds via interleaved local and global processes which permit scalable real-time operation in scenes with thousands of feature points. Notably, we use a relative 3D distance between the features which makes it possible to adjust the level of detail of the clusters according to their distance from the camera, such that the nearby scene is broken into more finely detailed clusters than the far background. We demonstrate the quality of our approach with video results showing live clustering of several indoor scenes with varying viewpoints and camera motions. The clusters produced are often consistently associated with single objects in the scene, and we forsee applications of our method both in providing cues for scene segmentation and labelling, and in building efficient 3D descriptors suitable for place recognition.
Adrien Angeli, Andrew J. Davison
BMVC2
2010 Scalable active matching
abstract
In matching tasks in computer vision, and particularly in real-time tracking from video, there are generally strong priors available on absolute and relative correspondence locations thanks to motion and scene models. While these priors are often partially used post-hoc to resolve matching consensus in algorithms like RANSAC, it was recently shown that fully integrating them in an `Active Matching' (AM) approach permits efficient guided image processing with rigorous decisions guided by Information Theory. AM's weakness was that the overhead induced by intermediate Bayesian updates required meant poor scaling to cases where many correspondences were sought. In this paper we show that relaxation of the rigid probabilistic model of AM, where every feature measurement directly affects the prediction of every other, permits dramatically more scalable operation without affecting accuracy. We take a general graph-theoretic view of the structure of prior information in matching to sparsify and approximate the interconnections. We demonstrate the performance of two variations, CLAM and SubAM, in the context of sequential camera tracking. These algorithms are highly competitive with other techniques at matching hundreds of features per frame while retaining great intuitive appeal and the full probabilistic capability to digest prior information.
Ankur Handa, Margarita Chli, Hauke Strasdat, Andrew J. Davison
CVPR4
2010 Live dense reconstruction with a single moving camera
abstract
We present a method which enables rapid and dense reconstruction of scenes browsed by a single live camera. We take point-based real-time structure from motion (SFM) as our starting point, generating accurate 3D camera pose estimates and a sparse point cloud. Our main novel contribution is to use an approximate but smooth base mesh generated from the SFM to predict the view at a bundle of poses around automatically selected reference frames spanning the scene, and then warp the base mesh into highly accurate depth maps based on view-predictive optical flow and a constrained scene flow update. The quality of the resulting depth maps means that a convincing global scene model can be obtained simply by placing them side by side and removing overlapping regions. We show that a cluttered indoor environment can be reconstructed from a live hand-held camera in a few seconds, with all processing performed by current desktop hardware. Real-time monocular dense reconstruction opens up many application areas, and we demonstrate both real-time novel view synthesis and advanced augmented reality where augmentations interact physically with the 3D scene and are correctly clipped by occlusions.
Richard A. Newcombe, Andrew J. Davison
CVPR2
2010 Real-Time Spherical Mosaicing Using Whole Image Alignment
Steven Lovegrove, Andrew J. Davison
ECCV (3)2
2010 Real-time monocular SLAM: Why filter?
abstract
While the most accurate solution to off-line structure from motion (SFM) problems is undoubtedly to extract as much correspondence information as possible and perform global optimisation, sequential methods suitable for live video streams must approximate this to fit within fixed computational bounds. Two quite different approaches to real-time SFM - also called monocular SLAM (Simultaneous Localisation and Mapping) - have proven successful, but they sparsify the problem in different ways. Filtering methods marginalise out past poses and summarise the information gained over time with a probability distribution. Keyframe methods retain the optimisation approach of global bundle adjustment, but computationally must select only a small number of past frames to process. In this paper we perform the first rigorous analysis of the relative advantages of filtering and sparse optimisation for sequential monocular SLAM. A series of experiments in simulation as well using a real image SLAM system were performed by means of covariance propagation and Monte Carlo methods, and comparisons made using a combined cost/accuracy measure. With some well-discussed reservations, we conclude that while filtering may have a niche in systems with low processing resources, in most modern applications keyframe optimisation gives the most accuracy per unit of computing time.
Hauke Strasdat, J. M. M. Montiel, Andrew J. Davison
ICRA3
2009 Automatically and efficiently inferring the hierarchical structure of visual maps
abstract
In Simultaneous Localisation and Mapping (SLAM), it is well known that probabilistic filtering approaches which aim to estimate the robot and map state sequentially suffer from poor computational scaling to large map sizes. Various authors have demonstrated that this problem can be mitigated by approximations which treat estimates of features in different parts of a map as conditionally independent, allowing them to be processed separately. When it comes to the choice of how to divide a large map into such ‘submaps’, straightforward heuristics may be sufficient in maps built using sensors such as laser range-finders with limited range, where a regular grid of submap boundaries performs well. With visual sensing, however, the ideal division of submaps is less clear, since a camera has potentially unlimited range and will often observe spatially distant parts of a scene simultaneously. In this paper we present an efficient and generic method for automatically determining a suitable submap division for SLAM maps, and apply this to visual maps built with a single agile camera. We use the mutual information between predicted measurements of features as an absolute measure of correlation, and cluster highly correlated features into groups. Via tree factorisation, we are able to determine not just a single level submap division but a powerful fully hierarchical correlation and clustering structure. Our analysis and experiments reveal particularly interesting structure in visual maps and give pointers to more efficient approximate visual SLAM algorithms.
Margarita Chli, Andrew J. Davison
ICRA2
2009 Camera self-calibration for sequential Bayesian structure from motion
abstract
Computer vision researchers have proved the feasibility of camera self-calibration —the estimation of a camera's internal parameters from an image sequence without any known scene structure. Various self-calibration algorithms have been published. Nevertheless, all of the recent sequential approaches to 3D structure and motion estimation from image sequences which have arisen in robotics and aim at real-time operation (often classed as visual SLAM or visual odometry) have relied on pre-calibrated cameras and have not attempted online calibration.
Javier Civera 0001, Diana R. Bueno, Andrew J. Davison, J. M. M. Montiel
ICRA3
2009 1-point RANSAC for EKF-based Structure from Motion
abstract
Recently, classical pairwise Structure From Motion (SfM) techniques have been combined with non-linear global optimization (Bundle Adjustment, BA) over a sliding window to recursively provide camera pose and feature location estimation from long image sequences. Normally called Visual Odometry, these algorithms are nowadays able to estimate with impressive accuracy trajectories of hundreds of meters; either from an image sequence (usually stereo) as the only input, or combining visual and propioceptive information from inertial sensors or wheel odometry. This paper has a double objective. First, we aim to illustrate for the first time how similar accuracy and trajectory length can be achieved by filtering-based visual SLAM methods. Specifically, a camera-centered Extended Kalman Filter is used here to process a monocular sequence as the only input, with 6DOF motion estimated. Features are kept live in the filter while visible as the camera explores forward, and are deleted from the state once they go out of view. This permits an increase in the number of tracked features per frame from tens to around a hundred. While improving the accuracy of the estimation, it makes computationally infeasible the exhaustive Branch and Bound search performed by standard JCBB for match outlier rejection. As a second contribution that overcomes this problem, we present here a RANSAC-like algorithm that exploits the probabilistic prediction of the filter. This use of prior information makes it possible to reduce the size of the minimal data subset to instantiate a hypothesis to the minimum possible of 1 point, greatly increasing the efficiency of the outlier rejection stage. Experimental results from real image sequences covering trajectories of hundreds of meters are presented and compared against RTK GPS ground truth. Estimation errors are about 1% of the trajectory for trajectories up to 650 metres.
Javier Civera 0001, Oscar G. Grasa, Andrew J. Davison, J. M. M. Montiel
IROS3
2009 Drift-Free Real-Time Sequential Mosaicing
Javier Civera 0001, Andrew J. Davison, Juan A. Magallon, J. M. M. Montiel
Int. J. Comput. Vis.2
2008 Active Matching
Margarita Chli, Andrew J. Davison
ECCV (1)2
2008 Interacting multiple model monocular SLAM
abstract
Recent work has demonstrated the benefits of adopting a fully probabilistic SLAM approach in sequential motion and structure estimation from an image sequence. Unlike standard Structure from Motion (SFM) methods, this 'monocular SLAM' approach is able to achieve drift-free estimation with high frame-rate real-time operation, particularly benefitting from highly efficient active feature search, map management and mismatch rejection. A consistent thread in this research on real-time monocular SLAM has been to reduce the assumptions required. In this paper we move towards the logical conclusion of this direction by implementing a fully Bayesian Interacting Multiple Models (IMM) framework which can switch automatically between parameter sets in a dimensionless formulation of monocular SLAM. Remarkably, our approach of full sequential probability propagation means that there is no need for penalty terms to achieve the Occam property of favouring simpler models - this arises automatically. We successfully tackle the known stiffness in on-the-fly monocular SLAM start up without known patterns in the scene. The search regions for matches are also reduced in size with respect to single model EKF increasing the rejection of spurious matches. We demonstrate our method with results on a complex real image sequence with varied motion.
Javier Civera 0001, Andrew J. Davison, J. M. M. Montiel
ICRA2
2008 Inverse Depth Parametrization for Monocular SLAM
abstract
We present a new parametrization for point features within monocular simultaneous localization and mapping (SLAM) that permits efficient and accurate representation of uncertainty during undelayed initialization and beyond, all within the standard extended Kalman filter (EKF). The key concept is direct parametrization of the inverse depth of features relative to the camera locations from which they were first viewed, which produces measurement equations with a high degree of linearity. Importantly, our parametrization can cope with features over a huge range of depths, even those that are so far from the camera that they present little parallax during motion---maintaining sufficient representative uncertainty that these points retain the opportunity to "come in'' smoothly from infinity if the camera makes larger movements. Feature initialization is undelayed in the sense that even distant features are immediately used to improve camera motion estimates, acting initially as bearing references but not permanently labeled as such. The inverse depth parametrization remains well behaved for features at all stages of SLAM processing, but has the drawback in computational terms that each point is represented by a 6-D state vector as opposed to the standard three of a EuclideanXYZrepresentation. We show that once the depth estimate of a feature is sufficiently accurate, its representation can safely be converted to the EuclideanXYZform, and propose a linearity index that allows automatic detection and conversion to maintain maximum efficiency---only low parallax features need be maintained in inverse depth form for long periods. We present a real-time implementation at 30 Hz, where the parametrization is validated in a fully automatic 3-D SLAM system featuring a handheld single camera with no additional sensing. Experiments show robust operation in challenging indoor and outdoor environments with a very large ranges of scene depth, varied motion, and also real time 360degloop closing.
Javier Civera 0001, Andrew J. Davison, J. M. M. Montiel
IEEE Trans. Robotics2
2007 Inverse Depth to Depth Conversion for Monocular SLAM
abstract
Recently it has been shown that an inverse depth parametrization can improve the performance of real-time monocular EKF SLAM, permitting undelayed initialization of features at all depths. However, the inverse depth parametrization requires the storage of 6 parameters in the state vector for each map point. This implies a noticeable computing overhead when compared with the standard 3 parameter XYZ Euclidean encoding of a 3D point, since the computational complexity of the EKF scales poorly with state vector size. In this work we propose to restrict the inverse depth parametrization only to cases where the standard Euclidean encoding implies a departure from linearity in the measurement equations. Every new map feature is still initialized using the 6 parameter inverse depth method. However, as the estimation evolves, if according to a linearity index the alternative XYZ coding can be considered linear, we show that feature parametrization can be transformed from inverse depth to XYZ for increased computational efficiency with little reduction in accuracy. We present a theoretical development of the necessary linearity indices, along with simulations to analyze the influence of the conversion threshold. Experiments performed with with a 30 frames per second real-time system are reported. An analysis of the increase in the map size that can be successfully managed is included.
Javier Civera 0001, Andrew J. Davison, J. M. M. Montiel
ICRA2
2007 Integrating Walking and Vision to Increase Humanoid Robot Autonomy
abstract
This video demonstrates our current investigation in developing autonomous behaviors for humanoid robots. Our main goal is to develop functionalities as much generic as possible in order to realize useful behaviors. More particularly this video demonstrates our current status on extending a popular zero momentum problem (ZMP) preview control based pattern generator, and building some links between walking with vision.
Olivier Stasse, Björn Verrelst, Andrew J. Davison, Nicolas Mansard, Bram Vanderborght, Claudia Esteves, François Saïdi, Kazuhito Yokoi
ICRA3
2007 MonoSLAM: Real-Time Single Camera SLAM
abstract
We present a real-time algorithm which can recover the 3D trajectory of a monocular camera, moving rapidly through a previously unknown scene. Our system, which we dub MonoSLAM, is the first successful application of the SLAM methodology from mobile robotics to the "pure vision" domain of a single uncontrolled camera, achieving real time but drift-free performance inaccessible to Structure from Motion approaches. The core of the approach is the online creation of a sparse but persistent map of natural landmarks within a probabilistic framework. Our key novel contributions include an active approach to mapping and measurement, the use of a general motion model for smooth camera movement, and solutions for monocular feature initialization and feature orientation estimation. Together, these add up to an extremely efficient and robust algorithm which runs at 30 Hz with standard PC and camera hardware. This work extends the range of robotic systems in which SLAM can be usefully applied, but also opens up new areas. We present applications of MonoSLAM to real-time 3D localization and mapping for a high-performance full-size humanoid robot and live augmented reality with a hand-held camera.
Andrew J. Davison, Ian D. Reid 0001, Nicholas Molton, Olivier Stasse
IEEE Trans. Pattern Anal. Mach. Intell.1
2006 Real-Time Monocular SLAM with Straight Lines
abstract
The use of line features in real-time visual tracking applications is commonplace when a prior map is available, but building the map while tracking in real-time is much more difficult. We describe how straight lines can be added to a monocular Extended Kalman Filter Simultaneous Mapping and Localisation (EKF SLAM) system in a manner that is both fast and which integrates easily with point features. To achieve real-time operation, we present a fast straight-line detector that hypothesises and tests straight lines connecting detected seed points. We demonstrate that the resulting system provides good camera localisation and mapping in real-time on a standard workstation, using either line features alone, or lines and points combined.
Ian D. Reid 0001, Andrew J. Davison
BMVC3
2006 A Visual Compass based on SLAM
abstract
Accurate full 3 axis orientation is computed using a low cost calibrated camera. We present a simultaneous sensor location and mapping method that uses a purely rotating camera as sensor and distant points, ideally at infinity, as features. A smooth constant angular velocity pure rotation motion model codifies the camera location. Because of the sequential EKF approach used, and the number of features in the map, about a hundred, the proposed method has been implemented in real time at standard video rates. Experimental results with real images show that the system is able to close loops with 360deg pan and 360deg cyclotorsion rotations. Sequences show good performance under challenging conditions: hand-held camera, varying natural outdoor illumination, low cost camera and lens and people moving in the scene
J. M. M. Montiel, Andrew J. Davison
ICRA2
2006 Active Control for Single Camera SLAM
abstract
In this paper we consider a single hand-held camera performing SLAM at video rate with generic 6DOF motion. The aim is to optimise both the localisation of the sensor and building of the feature map by computing the most appropriate control actions or movements. The actions belong to a discrete set (e.g. go forward, go left, go up, turn right, etc), and are chosen so as to maximise the mutual information gain between posterior states and measurements. Maximising the mutual information helps the camera avoid making ill-conditioned measurements appropriate to bearing-only SLAM. Moreover, orientation changes are determined by maximising the trace of the Fisher information matrix. In this way, we allow the camera to continue looking at those landmarks with large uncertainty, but from better-posed directions. Various position and gaze control strategies are first tested in a simulated environment, and then validated in a video-rate implementation. Given that our system is capable of producing motion commands for a real-time 6DOF visual SLAM, it could be used with any type of mobile platform, without the need of other sensors
Teresa Vidal-Calleja, Andrew J. Davison, Juan Andrade-Cetto, David William Murray 0001
ICRA2
2006 Real-time 3D SLAM for Humanoid Robot considering Pattern Generator Information
abstract
Humanoid robotics and SLAM (simultaneous localisation and mapping) are certainly two of the most significant themes of the current worldwide robotics research effort, but the two fields have up until now largely run independent parallel paths, despite the obvious benefit to be gained in joining the two. The next major step forward in humanoid robotics will be increased autonomy, and the ability of a robot to create its own world map on the fly will be a significant enabling technology. Meanwhile, SLAM techniques have found most success with robot platforms and sensor configurations which are outside of the humanoid domain. Humanoid robots move with high linear and angular accelerations in full 3D, and normally only vision is available as an outward-looking sensor. Building on recently published work on monocular SLAM using vision, and on pattern generation, we show that real-time SLAM for a humanoid can indeed be achieved. Using HRP-2, we present results in which a sparse 3D map of visual landmarks is acquired on the fly using a single camera and demonstrated loop closing and drift-free 3D motion estimation within a typical cluttered indoor environment. This is achieved by tightly coupling the pattern generator, the robot odometry and inertial sensing to aid visual mapping within a standard EKF framework. To our knowledge this is the first implementation of real-time 3D SLAM for a humanoid robot able to demonstrate loop closing
Olivier Stasse, Andrew J. Davison, Ramzi Sellaouti, Kazuhito Yokoi
IROS2
2006 Simultaneous Stereoscope Localization and Soft-Tissue Mapping for Minimal Invasive Surgery
Peter Mountney, Danail Stoyanov, Andrew J. Davison, Guang-Zhong Yang
MICCAI (1)3
2005 Active Search for Real-Time Vision
abstract
In most cases when information is to be extracted from an image, there are priors available on the state of the world and therefore on the detailed measurements which are obtained. While such priors are commonly combined with the actual measurements via Bayes' rule to calculate posterior probability distributions on model parameters, their additional value in guiding efficient image processing has almost always been overlooked. Priors tell us where to look for information in an image, how much computational effort we can expect to expend to extract it, and of how much utility to the task in hand it is likely to be. Such considerations are of importance in all practical real time vision systems, where the processing resources available at each frame in a sequence are strictly limited - and it is exactly in high frame rate real time systems such as trackers where strong priors are most likely to be available. In this paper, we use Shannon information theory to analyse the fundamental value of measurements using mutual information scores in absolute units of bits, specifically looking at the overwhelming case where uncertainty can be characterised by Gaussian probability distributions. We then compare these measurement values with the computational cost of the image processing required to obtain them. This theory puts on a firm footing for the first time principles of 'active search' for efficient guided image processing, in which candidate features of possibly different types can be compared and selected automatically for measurement.
Andrew J. Davison
ICCV1
2004 Interaction between hand and wearable camera in 2d and 3d environments
abstract
This paper is concerned with allowing the user of a wearable, portable, vision system to interact with the visual information using hand movements and gestures. Two example scenarios are explored. The first, in 2D, uses the wearer’s hand to both guide an active wearable camera and to highlight objects of interest using a grasping vector. The second is based in 3D, and builds on earlier work which recovers 3D scene structure at video-rate, allowing real-time purposive redirection of the camera to any scene point. Here, a range of hand gestures are used to highlight and select 3D points within the structure and in this instance used to insert 3D graphical objects into the scene. Structure recovery, gesture recognition, scene annotation and augmentation are achieved in parallel and at video-rate.
Walterio W. Mayol-Cuevas, Andrew J. Davison, Ben Tordoff, Nicholas Molton, David William Murray 0001
BMVC2
2004 Locally Planar Patch Features for Real-Time Structure from Motion
abstract
The performance of sequential structure from motion systems, where scene mapping is sparse to permit real-time operation, depends greatly on the ability to repeatedly measure the same visual features from a wide range of viewpoints. While previous systems have tracked features as 2D templates in image space, we show that long-term tracking is improved by treating salient feature patches as observations of locally planar regions on 3D world surfaces. Within a SLAM framework for motion and structure estimation, a gradient-based image alignment method is used to deduce estimates feature surface normal estimates, enabling pre-warping of templates for matching. As an added benefit these normals provide a richer description of the scene. 1
Nicholas Molton, Andrew J. Davison, Ian D. Reid 0001
BMVC2
2004 Advanced Visual Tracking
Simon J. Julier, Andrew J. Davison, Andrew W. Fitzgibbon
ISMAR2
2003 Real-Time Simultaneous Localisation and Mapping with a Single Camera
abstract
Ego-motion estimation for an agile single camera moving through general, unknown scenes becomes a much more challenging problem when real-time performance is required rather than under the off-line processing conditions under which most successful structure from motion work has been achieved. This task of estimating camera motion from measurements of a continuously expanding set of self-mapped visual features is one of a class of problems known as Simultaneous Localisation and Mapping (SLAM) in the robotics community, and we argue that such real-time mapping research, despite rarely being camera-based, is more relevant here than off-line structure from motion methods due to the more fundamental emphasis placed on propagation of uncertainty. We present a top-down Bayesian framework for single-camera localisation via mapping of a sparse set of natural features using motion modelling and an information-guided active measurement strategy, in particular addressing the difficult issue of real-time feature initialisation via a factored sampling approach. Real-time handling of uncertainty permits robust localisation via the creating and active measurement of a sparse map of landmarks such that regions can be re-visited after periods of neglect and localisation can continue through periods when few features are visible. Results are presented of real-time localisation for a hand-waved camera with very sparse prior scene knowledge and all processing carried out on a desktop PC.
Andrew J. Davison
ICCV1
2003 Real-Time Localisation and Mapping with Wearable Active Vision
abstract
We present a general method for real-time, vision-only single-camera simultaneous localisation and mapping (SLAM) - an algorithm which is applicable to the localisation of any camera moving through a scene - and study its application to the localisation of a wearable robot with active vision. Starting from very sparse initial scene knowledge, a map of natural point features spanning a section of a room is generated on-the-fly as the motion of the camera is simultaneously estimated in full 3D. Naturally this permits the annotation of the scene with rigidly-registered graphics, but further it permits automatic control of the robot's active camera: for instance, fixation on a particular object can be maintained during extended periods of arbitrary user motion, then shifted at will to another object which has potentially been out of the field of view. This kind of functionality is the key to the understanding or "management" of a workspace which the robot needs to have in order to assist its wearer usefully in tasks. We believe that the techniques and technology developed are of particular immediate value in scenarios of remote collaboration, where a remote expert is able to annotate, through the robot, the environment the wearer is working in.
Andrew J. Davison, Walterio W. Mayol-Cuevas, David William Murray 0001
ISMAR1
2003 Real-Time Visual Workspace Localisation and Mapping for a Wearable Robot
abstract
This demo showcases breakthrough results in the general field real-time simultaneous localization and mapping (SLAM) using vision and in particular its vital role in enabling a wearable robot to assists its user. In our approach, a wearable active vision system ("wearable robot") is mounted at the shoulder. As the wearer moves around his environment, typically browsing a workspace in which a task must be completed, the robot acquires images continuously and generates a map of natural visual features on-the-fly while estimating its ego-motion.
Andrew J. Davison, Walterio W. Mayol-Cuevas, David William Murray 0001
ISMAR1
2003 Applying Active Vision and SLAM to Wearables
Walterio W. Mayol-Cuevas, Andrew J. Davison, Ben Tordoff, David William Murray 0001
ISRR2
2002 Simultaneous Localization and Map-Building Using Active Vision
abstract
An active approach to sensing can provide the focused measurement capability over a wide field of view which allows correctly formulated simultaneous localization and map-building (SLAM) to be implemented with vision, permitting repeatable longterm localization using only naturally occurring, automatically-detected features. In this paper, we present the first example of a general system for autonomous localization using active vision, enabled here by a high-performance stereo head, addressing such issues as uncertainty-based measurement selection, automatic map-maintenance, and goal-directed steering. We present varied real-time experiments in a complex environment.
Andrew J. Davison, David William Murray 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2001 3D Simultaneous Localisation and Map-Building Using Active Vision for a Robot Moving on Undulating Terrain
abstract
Work in simultaneous localisation and map-building ("SLAM") for mobile robots has focused on the simplified case in which a robot is considered to move in two dimensions on a ground plane. While this is often a good approximation, a large number of real-world applications require robots to move around terrain which has significant slopes and undulations, and it is desirable that these robots too should be able to estimate their locations by building maps of natural features. We describe a real-time EKF-based-SLAM system permitting unconstrained 3D localisation, and in particular develop models for the motion of a wheeled robot in the presence of unknown slope variations. In a fully automatic implementation, our robot observes visual point features using fixating stereo vision and builds a sparse map on-the-fly. Combining this visual measurement with information from odometry and a roll/pitch accelerometer sensor, the robot performs accurate, repeatable localisation while traversing an undulating course.
Andrew J. Davison, Nobuyuki Kita
CVPR (1)1
2001 Automatic Partitioning of High Dimensional Search Spaces Associated with Articulated Body Motion Capture
abstract
Particle filters have proven to be an effective tool for visual tracking in non-Gaussian, cluttered environments. Conventional particle filters, however, do not scale to the problem of human motion capture (HMC) because of the large number of degrees of freedom involved. Annealed Particle Filtering (APF), introduced by J. Deutscher et al. (2000), tackled this by layering the search space and was shown to be a very effective tool for HMC. We improve upon and extend the APF in two ways. First we develop a hierarchical search strategy which automatically partitions the search space without any explicit representation of the partitions. Then we introduce a crossover operator (similar to that found in genetic algorithms) which improves the ability of the tracker to search different partitions in parallel. We present results for a simple example to demonstrate the new algorithm's implementation and then apply it to the considerably more complex problem of human motion capture with 34 degrees of freedom.
Jonathan Deutscher, Andrew J. Davison, Ian D. Reid 0001
CVPR (2)2
2001 Towards constant time SLAM using postponement
abstract
Many recent approaches to simultaneous localisation and mapping (SLAM) use an extended Kalman filter (EKF) to update and maintain a map of vehicle location. and multiple feature positions as a sensor moves through a scene. Although it is a highly powerful and well-used tool, it suffers from a well-known complexity problem. In this paper we outline the postponement technique which allows for much greater flexibility on when to use the available processing time, while not affecting the optimality of the filter. It works by updating a constant-sized data set based on current measurements, which can be used to affect the updates on all unobserved parts of the map at a later stage. By expanding the set of updated features when each new feature is observed we show that the full map update can be postponed indefinitely. We also demonstrate how postponement can be used to improve the performance of sub-optimal algorithms by applying it to a simple constant time method.
Joss Knight, Andrew J. Davison, Ian D. Reid 0001
IROS2
2000 Active visual localisation for cooperating inspection robots
abstract
In the routine inspection of industrial or other areas, teams of robots with various sensors could operate together to great effect, but require reliable, accurate and flexible localisation capabilities to be able to move around safely. We demonstrate accurate localisation for an inspection team consisting of a robot with stereo active vision and its companion with an active lighting system, and show that in this case a single sensor can be used for measuring the position of known or unknown scene features, measuring the relative location of the two robots, and actually carrying out an inspection task.
Andrew J. Davison, Nobuyuki Kita
IROS1
1998 Mobile Robot Localisation Using Active Vision
Andrew J. Davison, David William Murray 0001
ECCV (2)1
1996 Steering and Navigation Behaviours Using Fixation
abstract
Steering a motor vehicle around a winding but otherwise uncluttered road has been observed by Land and Lee (1994) to involve repeated periods of visual fixation upon the tangent point of the inside of each bend. We demonstrate a similar use of `active' fixation in the autonomous navigation of a robot vehicle around an obstacle, and show how the control law devised for steering in the robotic example is applicable to the observed human performance data. We discuss the merits of fixation for mobile robot localization.
David William Murray 0001, Ian D. Reid 0001, Andrew J. Davison
BMVC3
1995 The Active Camera as a Projective Pointing Device
abstract
This paper demonstrates an approach which exploits an active camera as a projective pointing mechanism. The optical centre of a static camera is notionally substituted by the centre of rotation of the active camera, as is the image plane by a frontal plane, a plane perpendicular to the optical axis of the active camera in its resting direction. Algorithms devised for 3D motion and 3D structure recovery using a single passive camera become immediately applicable to the active camera without need for reformulation. Furthermore, because the active camera can access a panoramic field of view, instabilities which may arise when the field of view is small, or because the shared field of view between successive after movement is small, are lessened. Two quite different applications of the idea are presented. In the first, the homography between a planar surface in the scene and the frontal plane is recovered and used to recover scene trajectories. In the second, the essential matrix between points in two frontal-plane views is recovered and used to determine the motion of a mobile vehicle.
Andrew J. Davison, Ian D. Reid 0001, David William Murray 0001
BMVC1