VLDB 2026 Research / reviewers in the wild / expert
Paul Newman 0001
dblp:79/1187-1 · also Paul M. Newman
· DBLP profile ↗
129ranked-venue papers
9as first author
16since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 120 · 9 first-author · 13 since 2021Systems, architecture and hardware · 93 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation ModelsabstractFoundation models (FMs) trained with different objectives and data learn diverse representations, making some more effective than others for specific downstream tasks. Existing adaptation strategies, such as parameter-efficient fine-tuning, focus on individual models and do not exploit the complementary strengths across models. Probing methods offer a promising alternative by extracting information from frozen models, but current techniques do not scale well with large feature sets and often rely on dataset-specific hyperparameter tuning. We propose Combined backBones (ComBo), a simple and scalable probing-based adapter that effectively integrates features from multiple models and layers. ComBo compresses activations from layers of one or more FMs into compact token-wise representations and processes them with a lightweight transformer for task-specific prediction. Crucially, ComBo does not require dataset-specific tuning or backpropagation through the backbone models. However, not all models are equally relevant for all tasks. To address this, we introduce a mechanism that leverages ComBo’s joint multi-backbone probing to efficiently evaluate each backbone’s task-relevance, enabling both practical model comparison and improved performance through selective adaptation. On the 19 tasks of the VTAB-1k benchmark, ComBo outperforms previous probing methods, matches or surpasses more expensive alternatives, such as distillation-based model merging, and enables efficient probing of tuned models. Our results demonstrate that ComBo offers a practical and general-purpose framework for combining diverse representations from multiple FMs. Benjamin Ramtoula, Pierre-Yves Lajoie, Paul Newman 0001, Daniele De Martini |
NeurIPS | 3 |
| 2024 | That's My Point: Compact Object-centric LiDAR Pose Estimation for Large-scale Outdoor LocalisationabstractThis paper is about 3D pose estimation on LiDAR scans with extremely minimal storage requirements to enable scalable mapping and localisation. We achieve this by clustering all points of segmented scans into semantic objects and representing them only with their respective centroid and semantic class. In this way, each LiDAR scan is reduced to a compact collection of four-number vectors. This abstracts away important structural information from the scenes, which is crucial for traditional registration approaches. To mitigate this, we introduce an object-matching network based on self- and cross-correlation that captures geometric and semantic relationships between entities. The respective matches allow us to recover the relative transformation between scans through weighted Singular Value Decomposition (SVD) and RANdom SAmple Consensus (RANSAC). We demonstrate that such representation is sufficient for metric localisation by registering point clouds taken under different viewpoints on the KITTI dataset, and at different periods of time localising between KITTI and KITTI-360. We achieve accurate metric estimates comparable with state-of-the-art methods with almost half the representation size, specifically 1.33 kB on average. Georgi Pramatarov, Matthew Gadd, Paul Newman 0001, Daniele De Martini |
ICRA | 3 |
| 2024 | VDNA-PR: Using General Dataset Representations for Robust Sequential Visual Place RecognitionabstractThis paper adapts a general dataset representation technique to produce robust Visual Place Recognition (VPR) descriptors, crucial to enable real-world mobile robot localisation. Two parallel lines of work on VPR have shown, on one side, that general-purpose off-the-shelf feature representations can provide robustness to domain shifts, and, on the other, that fused information from sequences of images improves performance. In our recent work on measuring domain gaps between image datasets, we proposed a Visual Distribution of Neuron Activations (VDNA) representation to represent datasets of images. This representation can naturally handle image sequences and provides a general and granular feature representation derived from a general-purpose model. Moreover, our representation is based on tracking neuron activation values over the list of images to represent and is not limited to a particular neural network layer, therefore having access to high- and low-level concepts. This work shows how VDNAs can be used for VPR by learning a very lightweight and simple encoder to generate task-specific descriptors. Our experiments show that our representation can allow for better robustness than current solutions to serious domain shifts away from the training data distribution, such as to indoor environments and aerial imagery. Benjamin Ramtoula, Daniele De Martini, Matthew Gadd, Paul Newman 0001 |
ICRA | 4 |
| 2024 | Masked γ-SSL: Learning Uncertainty Estimation via Masked Image ModelingabstractThis work proposes a semantic segmentation network that produces high-quality uncertainty estimates in a single forward pass. We exploit general representations from foundation models and unlabelled datasets through a Masked Image Modeling (MIM) approach, which is robust to augmentation hyper-parameters and simpler than previous techniques. For neural networks used in safety-critical applications, bias in the training data can lead to errors; therefore it is crucial to understand a network’s limitations at run time and act accordingly. To this end, we test our proposed method on a number of test domains including the SAX Segmentation benchmark, which includes labelled test data from dense urban, rural and off-road driving domains. The proposed method consistently outperforms uncertainty estimation and Out-of-Distribution (OoD) techniques on this difficult benchmark. David S. W. Williams, Matthew Gadd, Paul Newman 0001, Daniele De Martini |
ICRA | 3 |
| 2024 | NeuralFloors++: Consistent Street-Level Scene Generation From BEV Semantic MapsabstractLearning autonomous driving capabilities requires diverse and realistic training data. This has led to exploring generative techniques as an alternative to real-world data collection. In this paper we propose a method for synthesising photo-realistic urban driving scenes, along with semantic, instance and depth ground-truth. Our model relies on Bird’s Eye View (BEV) representations due to their compositionality and scene content control capabilities, reducing the need for traditional simulators. We employ a two-stage process: first, a 3D scene representation is extracted from BEV semantic, instance and style maps using a neural field. After rendering the semantic, instance, depth and style maps from a ground-view perspective, a second stage based on a diffusion model is used to generate the photo-realistic scene. We extend our prior work - NeuralFloors, to include multiple-view outputs, style manipulation for finer control at the object level through instance-wise style maps and cross-frame consistency via auto-regressive training. The proposed system is evaluated extensively on the KITTI-360 dataset, showing improved realism and semantic alignment for generated images. Valentina Musat, Daniele De Martini, Matthew Gadd, Paul Newman 0001 |
IROS | 4 |
| 2024 | OORD: The Oxford Offroad Radar DatasetabstractThere is a growing academic interest as well as commercial exploitation of millimetre-wave scanning radar for autonomous vehicle localisation and scene understanding. Although several datasets to support this research area have been released, they are primarily focused on urban or semi-urban environments. Nevertheless, rugged offroad deployments are important application areas which also present unique challenges and opportunities for this sensor technology. Therefore, the Oxford Offroad Radar Dataset (OORD) presents data collected in the rugged Scottish highlands in extreme weather. The radar data we offer to the community are accompanied by GPS/INS reference – to further stimulate research in radar place recognition. In total we release over 90 GiB of radar scans as well as GPS and IMU readings by driving a diverse set of four routes over 11 forays, totalling approximately 154 km of rugged driving. This is an area increasingly explored in literature, and we therefore present and release examples of recent open-sourced radar place recognition systems and their performance on our dataset. This includes a learned neural network, the weights of which we also release. The data and tools are made freely available to the community at oxford-robotics-institute.github.io/oord-dataset Matthew Gadd, Daniele De Martini, Oliver Bartlett, Paul Murcutt, Matthew Towlson, Matthew Widojo, Valentina Musat, Luke Robinson, Efimia Panagiotaki, Georgi Pramatarov, Marc Alexander Kühn, Letizia Marchegiani, Paul Newman 0001, Lars Kunze |
IEEE Trans. Intell. Transp. Syst. | 13 |
| 2024 | Mitigating Distributional Shift in Semantic Segmentation via Uncertainty Estimation From Unlabeled DataabstractKnowing when a trained segmentation model is encountering data that is different to its training data is important. Understanding and mitigating the effects of this play an important part in their application from a performance and assurance perspective-this being a safety concern in applications such as autonomous vehicles (AVs). This work presents a segmentation network that can detect errors caused by challenging test domains without any additional annotation in a single forward pass. As annotation costs limit the diversity of labelled datasets, we use easy-to-obtain, uncurated and unlabelled data to learn to perform uncertainty estimation by selectively enforcing consistency over data augmentation. To this end, a novel segmentation benchmark based on the SAX Dataset is used, which includes labelledtestdata spanning three autonomous-driving domains, ranging in appearance from dense urban to off-road. The proposed method, named$\mathrm{\gamma }{-}\rm{SSL}$, consistently outperforms uncertainty estimation and Out-of-Distribution (OoD) techniques on this difficult benchmark-by up to 10.7% in area under the receiver operating characteristic (ROC) curve and 19.2% in area under the precision-recall (PR) curve in the most challenging of the three scenarios. David S. W. Williams, Daniele De Martini, Matthew Gadd, Paul Newman 0001 |
IEEE Trans. Robotics | 4 |
| 2023 | Visual DNA: Representing and Comparing Images Using Distributions of Neuron ActivationsabstractSelecting appropriate datasets is critical in modern computer vision. However, no general-purpose tools exist to evaluate the extent to which two datasets differ. For this, we propose representing images - and by extension datasets - using Distributions of Neuron Activations (DNAs). DNAsfit distributions, such as histograms or Gaussians, to activations of neurons in a pre-trained feature extractor through which we pass the imager s) to represent. This extractor is frozen for all datasets, and we rely on its generally expressive power in feature space. By comparing two DNAs, we can evaluate the extent to which two datasets differ with granular control over the comparison attributes of interest, providing the ability to customise the way distances are measured to suit the requirements of the task at hand. Furthermore, DNAs are compact, representing datasets of any size with less than 15 megabytes. We demonstrate the value of DNAs by evaluating their applicability on several tasks, including conditional dataset comparison, synthetic image evaluation, and transfer learning, and across diverse datasets, ranging from synthetic cat images to celebrity faces and urban driving scenes. Benjamin Ramtoula, Matthew Gadd, Paul Newman 0001, Daniele De Martini |
CVPR | 3 |
| 2023 | Visual Servoing on Wheels: Robust Robot Orientation Estimation in Remote Viewpoint ControlabstractThis work proposes a fast deployment pipeline for visually-servoed robots which does not assume anything about either the robot - e.g. sizes, colour or the presence of markers - or the deployment environment. Specifically, we apply a learning based approach to reliably estimate the pose of a robot in the image frame of a 2D camera upon which a visual servoing control system can be deployed. To alleviate the time-consuming process of labelling image data, we propose a weakly supervised pipeline that can produce a vast amount of data in a small amount of time. We evaluate our approach on a dataset of remote camera images captured in various indoor environments demonstrating high tracking performances when integrated into a fully-autonomous pipeline with a simple controller. With this, we then analyse the data requirement of our approach, showing how it is possible to deploy a new robot in a new environment in fewer than 30.00 min. Luke Robinson, Daniele De Martini, Matthew Gadd, Paul Newman 0001 |
IROS | 4 |
| 2023 | Off the Radar: Uncertainty-Aware Radar Place Recognition with Introspective Querying and Map MaintenanceabstractLocalisation with Frequency-Modulated Continuous-Wave (FMCW) radar has gained increasing interest due to its inherent resistance to challenging environments. However, complex artefacts of the radar measurement process require appropriate uncertainty estimation - to ensure the safe and reliable application of this promising sensor modality. In this work, we propose a multi-session map management system which constructs the “best” maps for further localisation based on learned variance properties in an embedding space. Using the same variance properties, we also propose a new way to introspectively reject localisation queries that are likely to be incorrect. For this, we apply robust noise-aware metric learning, which both leverages the short-timescale variability of radar data along a driven path (for data augmentation) and predicts the downstream uncertainty in metric-space-based place recognition. We prove the effectiveness of our method over extensive cross-validated tests of the Oxford Radar RobotCar and MulRan dataset. In this, we outperform the current state-of-the-art in radar place recognition and other uncertainty-aware methods when using only single nearest-neighbour queries. We also show consistent performance increases when rejecting queries based on uncertainty over a difficult test environment, which we did not observe for a competing uncertainty-aware place recognition system. Jianhao Yuan, Paul Newman 0001, Matthew Gadd |
IROS | 2 |
| 2022 | Depth-SIMS: Semi-Parametric Image and Depth SynthesisabstractIn this paper we present a compositing image synthesis method that generates RGB canvases with well aligned segmentation maps and sparse depth maps, coupled with an in-painting network that transforms the RGB canvases into high quality RGB images and the sparse depth maps into pixel-wise dense depth maps. We benchmark our method in terms of structural alignment and image quality, showing an increase in mIoU over SOTA by 3.7 percentage points and a highly competitive FID. Furthermore, we analyse the quality of the generated data as training data for semantic segmentation and depth completion, and show that our approach is more suited for this purpose than other methods. Valentina Musat, Daniele De Martini, Matthew Gadd, Paul Newman 0001 |
ICRA | 4 |
| 2022 | Fast-MbyM: Leveraging Translational Invariance of the Fourier Transform for Efficient and Accurate Radar OdometryabstractMasking by Moving (MByM), provides robust and accurate radar odometry measurements through an exhaustive correlative search across discretised pose candidates. However, this dense search creates a significant computational bottleneck which hinders real-time performance when high-end GPUs are not available. Utilising the translational invariance of the Fourier Transform, in our approach, Fast Masking by Moving (f-MByM), we decouple the search for angle and translation. By maintaining end-to-end differentiability a neural network is used to mask scans and trained by supervising pose prediction directly. Training faster and with less memory, utilising a decoupled search allows f-MbyM to achieve significant run-time performance improvements on a CPU (168 %) and to run in real-time on embedded devices, in stark contrast to MbyM. Throughout, our approach remains accurate and competitive with the best radar odometry variants available in the literature – achieving an end-point drift of 2.01 % in translation and 6.3 deg /km on the Oxford Radar RobotCar Dataset. Rob Weston, Matthew Gadd, Daniele De Martini, Paul Newman 0001, Ingmar Posner |
ICRA | 4 |
| 2022 | BoxGraph: Semantic Place Recognition and Pose Estimation from 3D LiDARabstractThis paper is about extremely robust and lightweight localisation using LiDAR point clouds based on instance segmentation and graph matching. We model 3D point clouds as fully-connected graphs of semantically identified components where each vertex corresponds to an object instance and encodes its shape. Optimal vertex association across graphs allows for full 6-Degree-of-Freedom (DoF) pose estimation and place recognition by measuring similarity. This representation is very concise, condensing the size of maps by a factor of 25 against the state-of-the-art, requiring only 3 kB to represent a 1.4 MB laser scan. We verify the efficacy of our system on the SemanticKITTI dataset, where we achieve a new state-of-the-art in place recognition, with an average of 88.4 % recall at 100 % precision where the next closest competitor follows with 64.9 %. We also show accurate metric pose estimation performance - estimating 6-DoF pose with median errors of 10cm and 0.33 deg. Georgi Pramatarov, Daniele De Martini, Matthew Gadd, Paul Newman 0001 |
IROS | 4 |
| 2022 | Listening for Sirens: Locating and Classifying Acoustic Alarms in City ScenesabstractThis paper is about acoustic event detection and sound source localisation in urban scenarios. Specifically, we are interested in detecting and localising horns and sirens of emergency vehicles. Urban scenarios, though, can be characterised by copious, unstructured and unpredictable traffic noise, which can severely compromise the performance and effectiveness of traditional filtering techniques. By analysing the spectrograms of incoming stereo signals as images, we can leverage image processing techniques and obtain a demonstrably robust system. Indeed, image processing methods, such as convolutional neural networks, which do not operate locally, offer interesting mechanisms for background foreground separation. When applied to spectrograms, those mechanisms allow using the entire context of the soundscape to discover and learn correlations both in the time and frequency domains, de facto implementing noise detection through semantic segmentation. In a multi-task learning scheme, together with signal denoising, we perform acoustic event classification to identify the nature of the alerting sound. Lastly, we use the denoised signals to localise the acoustic source on the ground plane, by regressing the direction of arrival of the sound. Our experimental evaluation shows an average classification rate of 94%, and a median absolute error on the localisation of 7.5° when operating on audio frames of 0.5 s, and of 2.5° when operating on frames of 2.5 s. The system offers excellent performance in particularly challenging scenarios, where the noise level is remarkably high. Letizia Marchegiani, Paul Newman 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Fool Me Once: Robust Selective Segmentation via Out-of-Distribution Detection with Contrastive LearningabstractIn this work, a neural network is trained to simultaneously perform segmentation and pixel-wise Out-of-Distribution (OoD) detection, such that the segmentation of unknown regions of scenes can be rejected. This is made possible by leveraging an OoD dataset with a novel contrastive objective and data augmentation scheme. By including unknown classes in the training data, a more robust feature representation is learned with known classes represented distinctly from those unknown. In comparison, when presented with unknown classes or conditions, many current approaches for segmentation frequently exhibit high confidence in their inaccurate segmentations and cannot be trusted in many operational environments. We validate our system on a real-world dataset of unusual driving scenes, and show that by selectively segmenting scenes based on what is predicted as OoD, we can increase the segmentation accuracy by an IoU of 0.2 with respect to alternative techniques. David S. W. Williams, Matthew Gadd, Daniele De Martini, Paul Newman 0001 |
ICRA | 4 |
| 2021 | Look Here: Learning Geometrically Consistent Refinement of Inverse-Depth Images for 3D ReconstructionabstractBuilding good 3D maps is a challenging and expensive task, which requires high-quality sensors and careful, time-consuming scanning. We seek to reduce the cost of building good reconstructions by correcting views of existing low-quality ones in a post-hoc fashion using learnt priors over surfaces and appearance. We train a convolutional neural network model to predict the difference in inverse-depth from varying viewpoints of two meshes — one of low-quality that we wish to correct, and one of high-quality that we use as a reference. Our full model runs at 11.3[Formula: see text]Hz when aggregating four input views. In contrast to previous work, we pay attention to the problem of excessive smoothing in corrected meshes. We address this with a suitable network architecture, and introduce a loss-weighting mechanism that emphasizes edges in the prediction. Furthermore, smooth predictions result in geometrical inconsistencies. To deal with this issue, we present a loss function which penalizes re-projection differences that are not due to occlusions. Future applications of this work will incorporate semantic scene understanding in a multi-task learning setting. We explore the efficacy of the proposed system in terms of gross error correction and generalization capability by showing its performance in practice on a subset of the Kitti Odometry dataset, complete with a component-wise ablation study. We evaluate correctness and completeness measures of surface reconstruction across viewpoints and show that the proposed system is introspective in regions lacking sufficient high-quality supervision — indeed, models trained with geometric consistency loss create a lot more surface in areas that were not supervised, in one case filling in 67.97% or 8010[Formula: see text]m2 of an unlabeled input region. Finally, we assess the practical applicability of our method at large-scale by experiments over the full scope of the Kitti Odometry dataset. Broadly, as a measure of effectiveness, our model reduces gross errors by 45.3–77.5%, up to five times more than previous work. We also assess the practical applicability of our method to 3D reconstruction at large scales and find that compared to the baseline our model shows better stability in correctness when improving completeness of surfaces, and is effective in reducing median total error by up to 21.8[Formula: see text]cm. Stefan Saftescu, Matthew Gadd, Paul Newman 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2020 | The Oxford Radar RobotCar Dataset: A Radar Extension to the Oxford RobotCar DatasetabstractIn this paper we present The Oxford Radar RobotCar Dataset, a new dataset for researching scene understanding using Millimetre-Wave FMCW scanning radar data. The target application is autonomous vehicles where this modality is robust to environmental conditions such as fog, rain, snow, or lens flare, which typically challenge other sensor modalities such as vision and LIDAR.(/P)(P)The data were gathered in January 2019 over thirty-two traversals of a central Oxford route spanning a total of 280 km of urban driving. It encompasses a variety of weather, traffic, and lighting conditions. This 4.7 TB dataset consists of over 240,000 scans from a Navtech CTS350-X radar and 2.4 million scans from two Velodyne HDL-32E 3D LIDARs; along with six cameras, two 2D LIDARs, and a GPS/INS receiver. In addition we release ground truth optimised radar odometry to provide an additional impetus to research in this domain. The full dataset is available for download at: ori.ox.ac.uk/datasets/radar-robotear-dataset. Dan Barnes, Matthew Gadd, Paul Murcutt, Paul Newman 0001, Ingmar Posner |
ICRA | 4 |
| 2020 | Radar as a Teacher: Weakly Supervised Vehicle Detection using Radar LabelsabstractIt has been demonstrated that the performance of an object detector degrades when it is used outside the domain of the data used to train it. However, obtaining training data for a new domain can be time consuming and expensive. In this work we demonstrate how a radar can be used to generate plentiful (but noisy) training data for image-based vehicle detection. We then show that the performance of a detector trained using the noisy labels can be considerably improved through a combination of noise-aware training techniques and relabelling of the training data using a second viewpoint. In our experiments, using our proposed process improves average precision by more than 17 percentage points when training from scratch and 10 percentage points when fine-tuning a pre-trained model. Simon Chadwick, Paul Newman 0001 |
ICRA | 2 |
| 2020 | Kidnapped Radar: Topological Radar Localisation using Rotationally-Invariant Metric LearningabstractThis paper presents a system for robust, large-scale topological localisation using Frequency-Modulated Continuous-Wave scanning radar which extends the state-of-the-art by an efficient, learning-based approach to handle radar data for localisation. We learn a metric space for embedding polar radar scans using CNN and NetVLAD architectures traditionally applied to the visual domain. However, we tailor the feature extraction for more suitability to the polar nature of radar scan formation using cylindrical convolutions, anti-aliasing blurring, and azimuth-wise max-pooling; all in order to bolster the rotational invariance. The enforced metric space is then used to encode a reference trajectory, serving as a map, which is queried for nearest neighbour for recognition of places at run-time. We demonstrate the performance of our topological localisation system over the course of many repeat forays using the largest radar-focused mobile autonomy dataset released to date, totalling 280 km of urban driving, a small portion of which we also use to learn the weights of the modified architecture. As this work represents a novel application for radar, we analyse the utility of the proposed method via a comprehensive set of metrics which provide insight into the efficacy when used in a realistic system, showing improved performance over the root architecture even in the face of random rotational perturbation. Stefan Saftescu, Matthew Gadd, Daniele De Martini, Dan Barnes, Paul Newman 0001 |
ICRA | 5 |
| 2020 | Sense-Assess-eXplain (SAX): Building Trust in Autonomous Vehicles in Challenging Real-World Driving ScenariosabstractThis paper discusses ongoing work in demonstrating research in mobile autonomy in challenging driving scenarios. In our approach, we address fundamental technical issues to overcome critical barriers to assurance and regulation for large-scale deployments of autonomous systems. To this end, we present how we build robots that (1) can robustly sense and interpret their environment using traditional as well as unconventional sensors; (2) can assess their own capabilities; and (3), vitally in the purpose of assurance and trust, can provide causal explanations of their interpretations and assessments. As it is essential that robots are safe and trusted, we design, develop, and demonstrate fundamental technologies in real-world applications to overcome critical barriers which impede the current deployment of robots in economically and socially important areas. Finally, we describe ongoing work in the collection of an unusual, rare, and highly valuable dataset. Matthew Gadd, Daniele De Martini, Letizia Marchegiani, Paul Newman 0001, Lars Kunze |
IV | 4 |
| 2020 | RSS-Net: Weakly-Supervised Multi-Class Semantic Segmentation with FMCW RadarabstractThis paper presents an efficient annotation procedure and an application thereof to end-to-end, rich semantic segmentation of the sensed environment using Frequency-Modulated Continuous-Wave scanning radar. We advocate radar over the traditional sensors used for this task as it operates at longer ranges and is substantially more robust to adverse weather and illumination conditions. We avoid laborious manual labelling by exploiting the largest radar-focused urban autonomy dataset collected to date, correlating radar scans with RGB cameras and LiDAR sensors, for which semantic segmentation is an already consolidated procedure. The training procedure leverages a state-of-the-art natural image segmentation system which is publicly available and as such, in contrast to previous approaches, allows for the production of copious labels for the radar stream by incorporating four camera and two LiDAR streams. Additionally, the losses are computed taking into account labels to the radar sensor horizon by accumulating LiDAR returns along a pose-chain ahead and behind of the current vehicle position. Finally, we present the network with multi-channel radar scan inputs in order to deal with ephemeral and dynamic scene objects. Prannay Kaul, Daniele De Martini, Matthew Gadd, Paul Newman 0001 |
IV | 4 |
| 2019 | Fast Radar Motion Estimation with a Learnt Focus of Attention using Weak SupervisionabstractThis paper is about fast motion estimation with scanning radar. We use weak supervision to train a focus of attention policy which actively down-samples the measurement stream before data association steps are undertaken. At training, we avoid laborious manual labelling by exploiting short-term sensor coherence from multiple poses in the presence of an external ego-motion estimator (for example, wheel odometry). In this way, we generate copious annotated measurements which can be used for training a learning algorithm in a weakly-supervised fashion. We demonstrate the validity of the approach in the context of a Radar Odometry (RO) task, pre-filtering raw data with a popular image segmentation network trained as presented. We evaluate our system against 26 km of data collected in Central Oxford and show consistent motion estimation with greatly reduced radar processing times (by a factor of 2.36). Roberto Aldera, Daniele De Martini, Matthew Gadd, Paul Newman 0001 |
ICRA | 4 |
| 2019 | Radar-only ego-motion estimation in difficult settings via graph matchingabstractRadar detects stable, long-range objects under variable weather and lighting conditions, making it a reliable and versatile sensor well suited for ego-motion estimation. In this work, we propose a radar-only odometry pipeline that is highly robust to radar artifacts (e.g., speckle noise and false positives) and requires only one input parameter. We demonstrate its ability to adapt across diverse settings, from urban UK to off-road Iceland, achieving a scan matching accuracy of approximately 5.20 cm and 0.0929 deg when using GPS as ground truth (compared to visual odometry's 5.77 cm and 0.1032 deg). We present algorithms for key point extraction and data association, framing the latter as a graph matching optimization problem, and provide an in-depth system analysis. Sarah H. Cen, Paul Newman 0001 |
ICRA | 2 |
| 2019 | Distant Vehicle Detection Using Radar and VisionabstractFor autonomous vehicles to be able to operate successfully they need to be aware of other vehicles with sufficient time to make safe, stable plans. Given the possible closing speeds between two vehicles, this necessitates the ability to accurately detect distant vehicles. Many current image-based object detectors using convolutional neural networks exhibit excellent performance on existing datasets such as KITTI. However, the performance of these networks falls when detecting small (distant) objects. We demonstrate that incorporating radar data can boost performance in these difficult situations. We also introduce an efficient automated method for training data generation using cameras of different focal lengths. Simon Chadwick, Will Maddern, Paul Newman 0001 |
ICRA | 3 |
| 2019 | I Can See Clearly Now: Image Restoration via De-RainingabstractWe present a method for improving segmentation tasks on images affected by adherent rain drops and streaks. We introduce a novel stereo dataset recorded using a system that allows one lens to be affected by real water droplets while keeping the other lens clear. We train a denoising generator using this dataset and show that it is effective at removing the effect of real water droplets, in the context of image reconstruction and road marking segmentation. To further test our de-noising approach, we describe a method of adding computer-generated adherent water droplets and streaks to any images, and use this technique as a proxy to demonstrate the effectiveness of our model in the context of general semantic segmentation. We benchmark our results using the CamVid road marking segmentation dataset, Cityscapes semantic segmentation datasets and our own realrain dataset, and show significant improvement on all tasks. Horia Porav, Tom Bruls, Paul Newman 0001 |
ICRA | 3 |
| 2019 | Probably Unknown: Deep Inverse Sensor Modelling RadarabstractRadar presents a promising alternative to lidar and vision in autonomous vehicle applications, able to detect objects at long range under a variety of weather conditions. However, distinguishing between occupied and free space from raw radar power returns is challenging due to complex interactions between sensor noise and occlusion. To counter this we propose to learn an Inverse Sensor Model (ISM) converting a raw radar scan to a grid map of occupancy probabilities using a deep neural network. Our network is selfsupervised using partial occupancy labels generated by lidar, allowing a robot to learn about world occupancy from past experience without human supervision. We evaluate our approach on five hours of data recorded in a dynamic urban environment. By accounting for the scene context of each grid cell our model is able to successfully segment the world into occupied and free space, outperforming standard CFAR filtering approaches. Additionally by incorporating heteroscedastic uncertainty into our model formulation, we are able to quantify the variance in the uncertainty throughout the sensor observation. Through this mechanism we are able to successfully identify regions of space that are likely to be occluded. Rob Weston, Sarah H. Cen, Paul Newman 0001, Ingmar Posner |
ICRA | 3 |
| 2019 | The Right (Angled) Perspective: Improving the Understanding of Road Scenes Using Boosted Inverse Perspective MappingabstractMany tasks performed by autonomous vehicles such as road marking detection, object tracking, and path planning are simpler in bird's-eye view. Hence, Inverse Perspective Mapping (IPM) is often applied to remove the perspective effect from a vehicle's front-facing camera and to remap its images into a 2D domain, resulting in a top-down view. Unfortunately, however, this leads to unnatural blurring and stretching of objects at further distance, due to the resolution of the camera, limiting applicability. In this paper, we present an adversarial learning approach for generating a significantly improved IPM from a single camera image in real time. The generated bird'seye-view images contain sharper features (e.g, road markings) and a more homogeneous illumination, while (dynamic) objects are automatically removed from the scene, thus revealing the underlying road layout in an improved fashion. We demonstrate our framework using real-world data from the Oxford Robot-Car Dataset and show that scene understanding tasks directly benefit from our boosted IPM approach. Tom Bruls, Horia Porav, Lars Kunze, Paul Newman 0001 |
IV | 4 |
| 2019 | Training Object Detectors With Noisy DataabstractThe availability of a large quantity of labelled training data is crucial for the training of modern object detectors. Hand labelling training data is time consuming and expensive while automatic labelling methods inevitably add unwanted noise to the labels. We examine the effect of different types of label noise on the performance of an object detector. We then show how co-teaching, a method developed for handling noisy labels and previously demonstrated on a classification problem, can be improved to mitigate the effects of label noise in an object detection setting. We illustrate our results using simulated noise on the KITTI dataset and on a vehicle detection task using automatically labelled data. Simon Chadwick, Paul Newman 0001 |
IV | 2 |
| 2018 | Geometric Multi-Model Fitting With a Convex Relaxation AlgorithmabstractWe propose a novel method for fitting multiple geometric models to multi-structural data via convex relaxation. Unlike greedy methods - which maximise the number of inliers - our approach efficiently searches for a soft assignment of points to geometric models by minimising the energy of the overall assignment. The inherently parallel nature of our approach, as compared to the sequential approach found in state-of-the-art energy minimisation techniques, allows for the elegant treatment of a scaling factor that occurs as the number of features in the data increases. This results in an energy minimisation that, per iteration, is as much as two orders of magnitude faster on comparable architectures thus bringing real-time, robust performance to a wider set of geometric multi-model fitting problems. We demonstrate the versatility of our approach on two canonical problems in estimating structure from images: plane extraction from RGB-D images and homography estimation from pairs of images. Our approach seamlessly adapts to the different metrics brought forth in these distinct problems. In both cases, we report results on publicly available data-sets that in most instances outperform the state-of-the-art while simultaneously presenting run-times that are as much as an order of magnitude faster. Paul Amayo, Pedro Pinies, Lina María Paz, Paul Newman 0001 |
CVPR | 4 |
| 2018 | Fast Global Labelling for Depth-Map Improvement Via Architectural PriorsabstractDepth map estimation techniques from cameras often struggle to accurately estimate the depth of large textureless regions. In this work we present a vision-only method that accurately extracts planar priors from a viewed scene without making any assumptions of the underlying scene layout. Through a fast global labelling, these planar priors can be associated to the individual pixels leading to more complete depth-maps specifically over large, plain and planar regions that tend to dominate the urban environment. When these depth-maps are deployed to the creation of a vision only dense reconstruction over large scales, we demonstrate reconstructions that yield significantly better results in terms of coverage while still maintaining high accuracy. Paul Amayo, Pedro Pinies, Lina María Paz, Paul Newman 0001 |
ICRA | 4 |
| 2018 | Surface Edge Explorer (see): Planning Next Best Views Directly from 3D ObservationsabstractSurveying 3D scenes is a common task in robotics. Systems can do so autonomously by iteratively obtaining measurements. This process of planning observations to improve the model of a scene is called Next Best View (NBV) planning. NBV planning approaches often use either volumetric (e.g., voxel grids) or surface (e.g., triangulated meshes) representations. Volumetric approaches generalise well between scenes as they do not depend on surface geometry but do not scale to high-resolution models of large scenes. Surface representations can obtain high-resolution models at any scale but often require tuning of unintuitive parameters or multiple survey stages. This paper presents a scene-model-free NBV planning approach with a density representation. The Surface Edge Explorer (SEE) uses the density of current measurements to detect and explore observed surface boundaries. This approach is shown experimentally to provide better surface coverage in lower computation time than the evaluated state-of-the-art volumetric approaches while moving equivalent distances. Rowan Border, Jonathan D. Gammell, Paul Newman 0001 |
ICRA | 3 |
| 2018 | Mark Yourself: Road Marking Segmentation via Weakly-Supervised Annotations from Multimodal DataabstractThis paper presents a weakly-supervised learning system for real-time road marking detection using images of complex urban environments obtained from a monocular camera. We avoid expensive manual labelling by exploiting additional sensor modalities to generate large quantities of annotated images in a weakly-supervised way, which are then used to train a deep semantic segmentation network. At run time, the road markings in the scene are detected in real time in a variety of traffic situations and under different lighting and weather conditions without relying on any preprocessing steps or predefined models. We achieve reliable qualitative performance on the Oxford RobotCar dataset, and demonstrate quantitatively on the CamVid dataset that exploiting these annotations significantly reduces the required labelling effort and improves performance. Tom Bruls, Will Maddern, Akshay A. Morye, Paul Newman 0001 |
ICRA | 4 |
| 2018 | Precise Ego-Motion Estimation with Millimeter-Wave Radar Under Diverse and Challenging ConditionsabstractIn contrast to cameras, lidars, GPS, and proprioceptive sensors, radars are affordable and efficient systems that operate well under variable weather and lighting conditions, require no external infrastructure, and detect long-range objects. In this paper, we present a reliable and accurate radar-only motion estimation algorithm for mobile autonomous systems. Using a frequency-modulated continuous-wave (FMCW) scanning radar, we first extract landmarks with an algorithm that accounts for unwanted effects in radar returns. To estimate relative motion, we then perform scan matching by greedily adding point correspondences based on unary descriptors and pairwise compatibility scores. Our radar odometry results are robust under a variety of conditions, including those under which visual odometry and GPS/INS fail. Sarah H. Cen, Paul Newman 0001 |
ICRA | 2 |
| 2018 | Adversarial Training for Adverse Conditions: Robust Metric Localisation Using Appearance TransferabstractWe present a method of improving visual place recognition and metric localisation under very strong appearance change. We learn an invertable generator that can transform the conditions of images, e.g. from day to night, summer to winter etc. This image transforming filter is explicitly designed to aid and abet feature-matching using a new loss based on SURF detector and dense descriptor maps. A network is trained to output synthetic images optimised for feature matching given only an input RGB image, and these generated images are used to localize the robot against a previously built map using traditional sparse matching approaches. We benchmark our results using multiple traversals of the Oxford RobotCar Dataset over a year-long period, using one traversal as a map and the other to localise. We show that this method significantly improves place recognition and localisation under changing and adverse conditions, while reducing the number of mapping runs needed to successfully achieve reliable localisation. Horia Porav, Will Maddern, Paul Newman 0001 |
ICRA | 3 |
| 2018 | Meshed Up: Learnt Error Correction in 3D ReconstructionsabstractDense reconstructions often contain errors that prior work has so far minimised using high quality sensors and regularising the output. Nevertheless, errors still persist. This paper proposes a machine learning technique to identify errors in three dimensional (3D) meshes. Beyond simply identifying errors, our method quantifies both the magnitude and the direction of depth estimate errors when viewing the scene. This enables us to Improve the reconstruction accuracy. We train a suitably deep network architecture with two 3D meshes: a high-quality laser reconstruction, and a lower quality stereo image reconstruction. The network predicts the amount of error in the lower quality reconstruction with respect to the high-quality one, having only view the former through its input. We evaluate our approach by correcting two dimensional (2D) inverse-depth images extracted from the 3D model, and show that our method improves the quality of these depth reconstructions by up to a relative 10% RMSE. Michael Tanner, Stefan Saftescu, Alex Bewley, Paul Newman 0001 |
ICRA | 4 |
| 2018 | Multimotion Visual Odometry (MVO): Simultaneous Estimation of Camera and Third-Party MotionsabstractEstimating motion from images is a well-studied problem in computer vision and robotics. Previous work has developed techniques to estimate the motion of a moving camera in a largely static environment (e.g., visual odometry) and to segment or track motions in a dynamic scene using known camera motions (e.g., multiple object tracking). It is more challenging to estimate the unknown motion of the camera and the dynamic scene simultaneously. Most previous work requires a priori object models (e.g., tracking-by-detection), motion constraints (e.g., planar motion), or fails to estimate the full SE (3) motions of the scene (e.g., scene flow). While these approaches work well in specific application domains, they are not generalizable to unconstrained motions. This paper extends the traditional visual odometry (VO) pipeline to estimate the full SE (3) motion of both a stereo/RGB-D camera and the dynamic scene. This multimotion visual odometry (MVO) pipeline requires no a priori knowledge of the environment or the dynamic objects. Its performance is evaluated on a real-world dynamic dataset with ground truth for all motions from a motion capture system. Kevin M. Judd, Jonathan D. Gammell, Paul Newman 0001 |
IROS | 3 |
| 2017 | NID-SLAM: Robust Monocular SLAM Using Normalised Information DistanceabstractWe propose a direct monocular SLAM algorithm based on the Normalised Information Distance (NID) metric. In contrast to current state-of-the-art direct methods based on photometric error minimisation, our information-theoretic NID metric provides robustness to appearance variation due to lighting, weather and structural changes in the scene. We demonstrate successful localisation and mapping across changes in lighting with a synthetic indoor scene, and across changes in weather (direct sun, rain, snow) using real-world data collected from a vehicle-mounted camera. Our approach runs in real-time on a consumer GPU using OpenGL, and provides comparable localisation accuracy to state-of-the-art photometric methods but significantly outperforms both direct and feature-based methods in robustness to appearance changes. Geoffrey Pascoe, Will Maddern, Michael Tanner, Pedro Pinies, Paul Newman 0001 |
CVPR | 5 |
| 2017 | Modelling scene change for large-scale long term laser localisationabstractThis paper addresses a difficulty in large-scale long term laser localisation - how to deal with scene change. We pose this as a distraction suppression problem. Urban driving environments are frequently subject to large dynamic outliers, such as buses, trucks etc. These objects can mask the static elements of the prior map that we rely on for localisation. At the same time some objects change shape in a way that is less dramatic but equally pernicious during localisation - for example trees over seasons and in wind, shop fronts and doorways. In this paper, we show how we can learn in high resolution, the areas of our map that are subject to such distractions (low value data) in a place-dependent approach. We demonstrate how to utilise this model to select individual laser measurements for localisation. Specifically, by leveraging repeated operation over weeks and months, for each point in our map pointcloud we build distributions of the errors associated with that point for multiple localisation passes. These distributions are then used to determine the legitimacy of laser measurements prior to their use in localisation. We demonstrate distraction suppression as a front-end process to large scale localiser by incrementally adding 50km of error data to our base map and show that robustness is improved over the base system with a further 10km of urban driving. Dan Withers, Paul Newman 0001 |
ICRA | 2 |
| 2017 | Principles of robotics: regulating robots in the real worldabstractThis paper proposes a set of five ethical principles, together with seven high-level messages, as a basis for responsible robotics. The Principles of Robotics were drafted in 2010 and published online in 2011. Since then the principles have influenced, and continue to influence, a number of initiatives in robot ethics but have not, to date, been formally published. This paper remedies that omission. Margaret A. Boden, Joanna Bryson, Darwin G. Caldwell, Kerstin Dautenhahn, Lilian Edwards, Sarah Kember, Paul Newman 0001, Vivienne Parry, Geoff Pegman, Tom Rodden, Tom Sorrell, Mick Wallis, Blay Whitby, Alan F. T. Winfield |
Connect. Sci. | 7 |
| 2016 | A unified representation for application of architectural constraints in large-scale mappingabstractThis paper is about discovering and leveraging architectural constraints in large scale 3D reconstructions using laser. Our contribution is to offer a formulation of the problem which naturally and in a unified way, captures the variety of architectural constraints that can be discovered and applied in urban reconstructions. We focus in particular on the case of survey construction with a push broom laser + VO system. Here visual odometry is combined with vertical 2D scans to create a 3D picture of the environment. A key characteristic here is that the sensors pass/sweep swiftly through the environment such that elements of the scene are seen only briefly by cameras and scanned just once by the laser. These qualities make for a an ill-constrained optimisation problem which is greatly aided if architectural constraints can be discovered and appropriately applied. We demonstrate our approach in an end-to-end implementation which discovers salient architectural constraints and rejects false loop closures before invoking an optimisation to return a 3D model of the workspace. We evaluate the precision of this model by comparison to a ground truth provided by a 3rd party professional survey using highend (static) 3D laser scanners. Paul Amayo, Pedro Pinies, Lina María Paz, Paul Newman 0001 |
ICRA | 4 |
| 2016 | Made to measure: Bespoke landmarks for 24-hour, all-weather localisation with a cameraabstractThis paper is about camera-only localisation in challenging outdoor environments, where changes in lighting, weather and season cause traditional localisation systems to fail. Conventional approaches to the localisation problem rely on point-features such as SIFT, SURF or BRIEF to associate landmark observations in the live image with landmarks stored in the map; however, these features are brittle to the severe appearance change routinely encountered in outdoor environments. In this paper, we propose an alternative to traditional point-features: we train place-specific linear SVM classifiers to recognise distinctive elements in the environment. The core contribution of this paper is an unsupervised mining algorithm which operates on a single mapping dataset to extract distinct elements from the environment for localisation. We evaluate our system on 205km of data collected from central Oxford over a period of six months in bright sun, night, rain, snow and at all times of the day. Our experiment consists of a comprehensive N-vs-N analysis on 22 laps of the approximately 10km route in central Oxford. With our proposed system, the portion of the route where localisation fails is reduced by a factor of 6, from 33.3% to 5.5%. Chris Linegar, Winston Churchill, Paul Newman 0001 |
ICRA | 3 |
| 2016 | Choosing a time and place for calibration of lidar-camera systemsabstractWe propose a calibration method that automatically estimates the extrinsic calibration between a sensor pose-graph from natural scenes. The sensor pose-graph represents a system of sensors comprising of lidars and cameras, without sensor co-visibility constraints. The method addresses the fact that each scene contributes differently to the calibration problem by introducing a diligent scene selection scheme. The algorithm searches over all scenes to extract a subset of exemplars, whose joint optimisation yields progressively better calibration estimates. This non-parametric method requires no knowledge of the physical world, and continuously finds scenes that better constrain the optimisation parameters. We explain the theory, implement the method, and provide detailed performance analyses with experiments on real-world data. Terry Scott 0002, Akshay A. Morye, Pedro Pinies, Lina María Paz, Ingmar Posner, Paul Newman 0001 |
ICRA | 6 |
| 2016 | What lies behind: Recovering hidden shape in dense mappingabstractIn mobile robotics applications, generation of accurate static maps is encumbered by the presence of ephemeral objects such as vehicles, pedestrians, or bicycles. We propose a method to process a sequence of laser point clouds and back-fill dense surfaces into gaps caused by removing objects from the scene - a valuable tool in scenarios where resource constraints permit only one mapping pass in a particular region. Our method processes laser scans in a three-dimensional voxel grid using the Truncated Signed Distance Function (TSDF) and then uses a Total Variation (TV) regulariser with a Kernel Conditional Density Estimation (KCDE) “soft” data term to interpolate missing surfaces. Using four scenarios captured with a push-broom 2D laser, our technique infills approximately 20 m2 of missing surface area for each removed object. Our reconstruction's median error ranges between 5.64 cm – 9.24 cm with standard deviations between 4.57 cm – 6.08 cm. Michael Tanner, Pedro Pinies, Lina María Paz, Paul Newman 0001 |
ICRA | 4 |
| 2016 | Checkout my map: Version control for fleetwide visual localisationabstractThis paper is about underpinning long-term operations of fleets of vehicles using visual localisation. In particular it examines ways in which vehicles, considered as independent agents, can share, update and leverage each others' visual experiences in a mutually beneficial way. We draw on our previous work in Experience-based Navigation (EBN) [1], in which a visual map supporting multiple representations of the same place is built, yielding real-time localisation capability for a solitary vehicle. We now consider how any number of such agents might operate in concert via data sharing policies that are germane to the shared task of lifelong localisation. We rapidly construct considerable maps by the conjoining of work distributed to asynchronous processes, and share expertise amongst the team by the selective dispensing of mission-specific map contents. We demonstrate and evaluate our system against 100km of data collected in North Oxford over a period of a month featuring diverse deviation in appearance due to atmospheric, lighting, and structural dynamics. We show that our framework is capable of creating maps in a fraction of the time required by single-agent EBN, with no significant loss in localisation robustness, and is able to furnish robots on real-world forays with maps which require much less storage. Matthew Gadd, Paul Newman 0001 |
IROS | 2 |
| 2016 | Real-time probabilistic fusion of sparse 3D LIDAR and dense stereoabstractReal-time 3D perception is critical for localisation, mapping, path planning and obstacle avoidance for mobile robots and autonomous vehicles. For outdoor operation in real-world environments, 3D perception is often provided by sparse 3D LIDAR scanners, which provide accurate but low-density depth maps, and dense stereo approaches, which require significant computational resources for accurate results. Here, taking advantage of the complementary error characteristics of LIDAR range sensing and dense stereo, we present a probabilistic method for fusing sparse 3D LIDAR data with stereo images to provide accurate dense depth maps and uncertainty estimates in real-time. We evaluate the method on data collected from a small urban autonomous vehicle and the KITTI dataset, providing accuracy results competitive with state-of-the-art stereo approaches and credible uncertainty estimates that do not misrepresent the true errors, and demonstrate real-time operation on a range of low-power GPU systems. Will Maddern, Paul Newman 0001 |
IROS | 2 |
| 2016 | The path less taken: A fast variational approach for scene segmentation used for closed loop controlabstractIn this paper we propose an on-line system that discovers and drives collision-free traversable paths, using a variational approach to dense stereo vision. Our system is light weight, can be run on low cost hardware and is remarkably quick to predict the semantics. In addition to the scene's path affordance it yields a segmentation of the local scene as a composite of distinctive labels - e.g, ground, sky, obstacles and vegetation. To estimate the labels, we combine a very fast and light weight (shallow) image classifier which considers informative feature channels derived from colour images and dense depth maps estimates. Unlike other approaches, we do not use local descriptors around pixel features. Instead, we encompass label-predicted probabilities with a variational approach for image segmentation. Akin to dense depth map estimation, we obtain semantically segmented images by means of convex regularisation. We show how our system can rapidly obtain the required semantics and paths at VGA resolution. Extensive experiments on the KITTI dataset support the robustness of our system to derive collision-free local routes. An accompanied video supports the robustness of the system at live execution in an outdoor experiment. Tarlan Suleymanov, Lina María Paz, Pedro Pinies, Geoff Hester, Paul Newman 0001 |
IROS | 5 |
| 2016 | Automated valet parking and charging for e-mobilityabstractAutomated valet parking services provide great potential to increase the attractiveness of electric vehicles by mitigating their two main current deficiencies: reduced driving ranges and prolonged refueling times. The European research project V-Charge aims at providing this service on designated parking lots using close-to-market sensors only. For this purpose the project developed a prototype capable of performing fully automated navigation in mixed traffic on designated parking lots and GPS-denied parking garages with cameras and ultrasonic sensors only. This paper summarizes the work of the project, comprising advances in network communication and parking space scheduling, multi-camera calibration, semantic mapping concepts, visual localization and motion planning. The project pushed visual localization, environment perception and automated parking to centimetre precision. The developed infrastructure-based camera calibration and semi-supervised semantic mapping concepts greatly reduce maintenance efforts. Results are presented from extensive month-long field tests. Ulrich Schwesinger, Mathias Bürki, Julian Timpner, Stephan Rottmann, Lars C. Wolf, Lina María Paz, Hugo Grimmett, Ingmar Posner, Paul Newman 0001, Christian Häne, Lionel Heng, Gim Hee Lee, Torsten Sattler, Marc Pollefeys, Marco Allodi, Francesco Valenti, Keiji Mimura, Bernd Goebelsmann, Wojciech Derendarz, Peter Mühlfellner, Stefan Wonneberger, Rene Waldmann, Sebastian Grysczyk, Carsten Last, Stefan Bruning, Sven Horstmann, Marc Bartholomaus, Clemens Brummer, Martin Stellmacher, Fabian Pucks, Marcel Nicklas, Roland Siegwart |
Intelligent Vehicles Symposium | 9 |
| 2016 | Visual Place Recognition: A SurveyabstractVisual place recognition is a challenging problem due to the vast range of ways in which the appearance of real-world places can vary. In recent years, improvements in visual sensing capabilities, an ever-increasing focus on long-term mobile robot autonomy, and the ability to draw on state-of-the-art research in other disciplines-particularly recognition in computer vision and animal navigation in neuroscience-have all contributed to significant advances in visual place recognition systems. This paper presents a survey of the visual place recognition research landscape. We start by introducing the concepts behind place recognition-the role of place recognition in the animal kingdom, how a “place” is defined in a robotics context, and the major components of a place recognition system. Long-term robot operations have revealed that changing appearance can be a significant factor in visual place recognition failure; therefore, we discuss how place recognition solutions can implicitly or explicitly account for appearance change within the environment. Finally, we close with a discussion on the future of visual place recognition, in particular with respect to the rapid advances being made in the related fields of deep learning, semantic scene understanding, and video description. Stephanie M. Lowry, Niko Sünderhauf, Paul Newman 0001, John J. Leonard, David D. Cox, Peter I. Corke, Michael Milford |
IEEE Trans. Robotics | 3 |
| 2015 | Robust Direct Visual Localisation using Normalised Information DistanceabstractWe present an information-theoretic approach for direct localisation of a monocular camera within a 3D appearance prior. In contrast to existing direct visual localisation methods based on minimising photometric error, an information-theoretic metric allows us to compare the whole image without relying on individual pixel values, yielding robustness to changes in the appearance of the scene due to lighting, camera motion, occlusions and sensor modality. Using a low-fidelity textured 3D model of the environment, we synthesise virtual images at a candidate pose within the model. We use the Normalised Information Distance (NID) metric to evaluate the appearance match between the camera image and the virtual image, and present a derivation of analytical NID derivatives for the SE(3) direct localisation problem, along with an efficient GPGPU implementation capable of online processing. We present results showing successful online visual localisation under significant appearance change both in a synthetic indoor environment and outdoors with real-world data from a vehicle-mounted camera. Geoffrey Pascoe, Will Maddern, Paul Newman 0001 |
BMVC | 3 |
| 2015 | Opportunistic Radio Assisted Navigation for Autonomous Ground VehiclesabstractNavigating autonomous ground vehicles with visual sensors has many advantages - it does not rely on global maps, yet is accurate and reliable even in GPS-denied environments. However, due to the limitation of the camera field of view, one typically has to record a large number of visual experiences for practical navigation. In this paper, we explore new avenues in linking together visual experiences, by opportunistically harvesting and sharing a variety of radio signals emitted by surrounding stationary access points and mobile devices. We propose a novel navigation approach, which exploits side-channel information of co-location to thread up visually-separated experiences with short exploration phases. The proposed approach empowers users to trade travel time for manual navigation effort, allowing them to choose the itinerary that best serves their needs. We evaluate the proposed approach with data collected from a typical urban area, and show that it achieves much better navigation performance in both reach ability and cost, comparing with the state of the arts that only use visual information. Hongkai Wen 0001, Yiran Shen 0001, Savvas Papaioannou, Winston Churchill, Agathoniki Trigoni, Paul Newman 0001 |
DCOSS | 6 |
| 2015 | Know your limits: Embedding localiser performance models in teach and repeat mapsabstractThis paper is about building maps which not only contain the traditional information useful for localising — such as point features — but also embeds a spatial model of expected localiser performance. This often overlooked second-order information provides vital context when it comes to map use and planning. Our motivation here is to improve the performance of the popular Teach and Repeat paradigm [1] which has been shown to enable truly large-scale field operation. When using the taught route for localisation, it is often assumed the robot is following exactly, or is sufficiently close to, the original path, enabling successful localisation. However, what happens if it is not possible, or not desirable to exactly follow the mapped path? How far off the beaten track can the robot travel before it gets lost? We present an approach for assessing this localisation area around a taught route, which we refer to as the localisation envelope. Using a combination of physical sampling and a Gaussian Process model, we are able to accurately predict the localisation performance at unseen points. Winston Churchill, Chi Hay Tong, Corina Gurau, Ingmar Posner, Paul Newman 0001 |
ICRA | 5 |
| 2015 | A framework for infrastructure-free warehouse navigationabstractThis paper presents a universally applicable graph-based framework for the navigation of warehouse robots equipped with only monocular cameras. We strongly advocate the use of relative pose information stored in a topological map, rather than a globally consistent metric representation of the environment. We show how multiple traversals of adjacent workspaces can be naturally “stitched” together in the course of a typical warehouse picking and shelving schedule to create a network of reusable paths in which the robot can efficiently localise and plan new routes. This allows us to command the robot to return to any of the previously visited locations not necessarily through the same route that we taught it. Unlike state-of-the-art teach and repeat systems using stereo vision, our approach exploits the strongly planar nature of the data obtained from a downward-facing camera, and creates odometric constraints by tracking the perceived texture of the floor and computing a simple homography. To demonstrate the robustness of our system, we validate our approach on datasets collected over a week-long period within a challenging and representative environment in the form of a warehouse shelving area. Matthew Gadd, Paul Newman 0001 |
ICRA | 2 |
| 2015 | Integrating metric and semantic maps for vision-only automated parkingabstractWe present a framework for integrating two layers of map which are often required for fully automated operation: metric and semantic. Metric maps are likely to improve with subsequent visitations to the same place, while semantic maps can comprise both permanent and fluctuating features of the environment. However, it is not clear from the state of the art how to update the semantic layer as the metric map evolves. The strengths of our method are threefold: the framework allows for the unsupervised evolution of both maps as the environment is revisited by the robot; it uses vision-only sensors, making it appropriate for production cars; and the human labelling effort is minimised as far as possible while maintaining high fidelity. We evaluate this on two different car parks with a fully automated car, performing repeated automated parking manoeuvres to demonstrate the robustness of the system. Hugo Grimmett, Mathias Bürki, Lina María Paz, Pedro Pinies, Paul Timothy Furgale, Ingmar Posner, Paul Newman 0001 |
ICRA | 7 |
| 2015 | Work smart, not hard: Recalling relevant experiences for vast-scale but time-constrained localisationabstractThis paper is about life-long vast-scale localisation in spite of changes in weather, lighting and scene structure. Building upon our previous work in Experience-based Navigation [1], we continually grow and curate a visual map of the world that explicitly supports multiple representations of the same place. We refer to these representations as experiences, where a single experience captures the appearance of an environment under certain conditions. Pedagogically, an experience can be thought of as a visual memory. By accumulating experiences we are able to handle cyclic appearance change (diurnal lighting, seasonal changes, and extreme weather conditions) and also adapt to slow structural change. This strategy, although elegant and effective, poses a new challenge: In a region with many stored representations - which one(s) should we try to localise against given finite computational resources? By learning from our previous use of the experience-map, we can make predictions about which memories we should consider next, conditioned on how the robot is currently localised in the experience-map. During localisation, we prioritise the loading of past experiences in order to minimise the expected computation required. We do this in a probabilistic way and show that this memory policy significantly improves localisation efficiency, enabling long-term autonomy on robots with limited computational resources. We demonstrate and evaluate our system over three challenging datasets, totalling 206km of outdoor travel. We demonstrate the system in a diverse range of lighting and weather conditions, scene clutter, camera occlusions, and permanent structural change in the environment. Chris Linegar, Winston Churchill, Paul Newman 0001 |
ICRA | 3 |
| 2015 | Leveraging experience for large-scale LIDAR localisation in changing citiesabstractRecent successful approaches to autonomous vehicle localisation and navigation typically involve 3D LIDAR scanners and a static, curated 3D map, both of which are expensive to acquire and maintain. In this paper we propose an experience-based approach to matching a local 3D swathe built using a push-broom 2D LIDAR to a number of prior 3D maps, each of which has been collected during normal driving in different conditions. Local swathes are converted to a combined 2D height and reflectance representation, and we exploit the GPU rendering pipeline to densely sample the localisation cost function to provide robustness and a wide basin of convergence. Prior maps are incrementally built into an experience-based framework from multiple traversals of the same environment, capturing changes in environment structure and appearance over time. The LIDAR localisation solutions from each prior map are fused with vehicle odometry in a probabilistic framework to provide a single pose solution suitable for automated driving. Using this framework we demonstrate real-time centimetre-level localisation using LIDAR data collected in a dynamic city environment over a period of a year. Will Maddern, Geoffrey Pascoe, Paul Newman 0001 |
ICRA | 3 |
| 2015 | From dusk till dawn: Localisation at night using artificial light sourcesabstractThis paper is about localising at night in urban environments using vision. Despite it being dark exactly half of the time, surprisingly little attention has been given to this problem. A defining aspect of night-time urban scenes is the presence and effect of artificial lighting - be that in the form of street or interior lighting through windows. By building a model of the environment which includes a representation of the spatial location of every light source, localisation becomes possible using monocular cameras. One of the challenges we face is the gross change in light appearance as a function of distance due to flare, saturation and bleeding - city lights certainly do not appear as point features. To overcome this, we model the appearance of each light as a function of vehicle location, using this to inform our data-association decisions and to regularise the cost function which is used to infer vehicle pose. In this way we develop a place-dependent but stable sensor model which is customised for the particular environment in which we are operating. We demonstrate that our system is able to localise successfully at night over 12 km in situations where a traditional point feature based system fails. Peter Nelson, Winston Churchill, Ingmar Posner, Paul Newman 0001 |
ICRA | 4 |
| 2015 | FARLAP: Fast robust localisation using appearance priorsabstractThis paper is concerned with large-scale localisation at city scales with monocular cameras. Our primary motivation lies with the development of autonomous road vehicles - an application domain in which low-cost sensing is particularly important. Here we present a method for localising against a textured 3-dimensional prior mesh using a monocular camera. We first present a system for generating and texturing the prior using a LIDAR scanner and camera. We then describe how we can localise against that prior with a single camera, using an information-theoretic measure of image similarity. This process requires dealing with the distortions induced by a wide-angle camera. We present and justify an interesting approach to this issue in which we distort the prior map into the image rather than vice-versa. Finally we explain how the general purpose computation functionality of a modern GPU is particularly apt for our task, allowing us to run the system in real time. We present results showing centimetre-level localisation accuracy through a city over six kilometres. Geoffrey Pascoe, Will Maddern, Alexander D. Stewart, Paul Newman 0001 |
ICRA | 4 |
| 2015 | A variational approach to online road and path segmentation with monocular visionabstractIn this paper we present an online approach to segmenting roads on large scale trajectories using only a monocular camera mounted on a car. We differ from popular 2D segmentation solutions which use single colour images and machine learning algorithms that require supervised training on huge image databases. Instead, we propose a novel approach that fuses 3D geometric data with appearance-based segmentation of 2D information in an automatic system. Our contribution is twofold: first, we propagate labels from frame to frame using depth priors of the segmented road avoiding user interaction most of the time; second, we transfer the segmented road labels to 3D laser point clouds. This reduces the complexity of state-of-the-art segmentation algorithms running on 3D Lidar data. Segmentation fails is in only 3% of the cases over a sequence of 13,600 monocular images spanning an urban trajectory of more than 10km. Lina María Paz, Pedro Pinies, Paul Newman 0001 |
ICRA | 3 |
| 2015 | Too much TV is bad: Dense reconstruction from sparse laser with non-convex regularisationabstractIn this paper we address the problem of dense depth map estimation from sparse noisy range data to reconstruct large heterogeneous outdoor scenes. We propose a surface inpainting solution through energy minimisation with an adaptive selection of surface regularisers among a set of well known convex and non-convex regularisers. In fact, the selection of norm is pivotal with respect to the intrinsic surface characteristics. Our goal is to show how dense interpolation of sparse range data can be leveraged of more exotic and non-convex regularisers such as the log and logTGV [1] which can better capture the scene geometry. In contrast to state of the art solutions, we do not restrict ourselves to this set of norms, instead we search for the most apt norm for each semantically segmented part of the scene. Our energy model selection use Bayesian optimisation to learn the best choice of free parameters. This results in an adaptive model selection and the generalisation of well studied regularisation norms. We conclude with a detailed experimental analysis of our approach using a basis of four norms over a set of challenging outdoor scenes. Pedro Pinies, Lina María Paz, Paul Newman 0001 |
ICRA | 3 |
| 2015 | Dense mono reconstruction: Living with the pain of the plain planeabstractThis paper is about dense depthmap estimation using a monocular camera in workspaces with extensive textureless surfaces. Current state of the art techniques have been shown to work in real time with an admirable performance in desktop-size environments. Unfortunately, as we show in this paper, when applied to larger indoor environments, performance often degrades. A common cause is the presence of large affine texture-less areas like by walls, floors, ceilings and drab objects such as chairs and tables. These produce noisy and worse still, grossly erroneous initial seeds for the depthmap that greatly impede successful optimisation. We solve this problem via the introduction of a new non-local higher-order regularisation term that enforces piecewise affine constraints between image pixels that are far apart in the image. This property leverages the observation that the depth at the edges of bland regions are often well estimated whereas their inner pixels are deeply problematic. A welcome by-product of our proposed technique is an estimate of the surface normals at each pixel. We will show that in terms of implementation, our algorithm is a natural extension of the often used variational approaches. We evaluate the proposed technique using real datasets for which we have ground truth models. Pedro Pinies, Lina María Paz, Paul Newman 0001 |
ICRA | 3 |
| 2015 | Exploiting known unknowns: Scene induced cross-calibration of lidar-stereo systemsabstractWe propose an automatic, targetless, data-driven, extrinsic calibration method to calibrate push-broom 2D lidars with a multi-camera system. The calibration problem is decoupled into alternating optimisers over two hierarchical levels, where both levels are linked with a penalty term. The lower-level optimises the six degrees-of-freedom (DoF) rigid-body transforms between the lidar and each camera of the multi-camera unit by minimising the Normalised Information Distance between intensity measurements obtained from both sensor modalities. The upper-level minimises a nonlinear least squares error between the lower-level solutions. We describe the theory, implement the method, and provide a detailed performance analysis with experiments on real-world data. Terry Scott 0002, Akshay A. Morye, Pedro Pinies, Lina María Paz, Ingmar Posner, Paul Newman 0001 |
IROS | 6 |
| 2015 | Reading the Road: Road Marking Classification and InterpretationabstractRoad markings embody the rules of the road whilst capturing the upcoming road layout. These rules are diligently studied and applied to driving situations by human drivers who have read Highway Traffic driving manuals (road marking interpretation). An autonomous vehicle must however be taught to read the road, as a human might. This paper addresses the problem of automatically reading the rules encoded in road markings, by classifying them into seven distinct classes: single boundary, double boundary, separator, zig-zag, intersection, boxed junction and special lane. Our method employs a unique set of geometric feature functions within a probabilistic RUSBoost and Conditional Random Field (CRF) classification framework. This allows us to jointly classify extracted road markings. Furthermore, we infer the semantics of road scenes (pedestrian approaches and no drive regions) based on marking classification results. Finally, our algorithms are evaluated on a large real-life ground truth annotated dataset from our vehicle. Bonolo Mathibela, Paul Newman 0001, Ingmar Posner |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2014 | Shady dealings: Robust, long-term visual localisation using illumination invarianceabstractThis paper is about extending the reach and endurance of outdoor localisation using stereo vision. At the heart of the localisation is the fundamental task of discovering feature correspondences between recorded and live images. One aspect of this problem involves deciding where to look for correspondences in an image and the second is deciding what to look for. This latter point, which is the main focus of our paper, requires understanding how and why the appearance of visual features can change over time. In particular, such knowledge allows us to better deal with abrupt and challenging changes in lighting. We show how by instantiating a parallel image processing stream which operates on illumination-invariant images, we can substantially improve the performance of an outdoor visual navigation system. We will demonstrate, explain and analyse the effect of the RGB to illumination-invariant transformation and suggest that for little cost it becomes a viable tool for those concerned with having robots operate for long periods outdoors. Colin McManus, Winston Churchill, Will Maddern, Alexander D. Stewart, Paul Newman 0001 |
ICRA | 5 |
| 2014 | Visual precis generation using coresetsabstractGiven an image stream, our on-line algorithm will select the semantically-important images that summarize the visual experience of a mobile robot. Our approach consists of data pre-clustering using coresets followed by a graph based incremental clustering procedure using a topic based image representation. A coreset for an image stream is a set of representative images that semantically compresses the data corpus, in the sense that every frame has a similar representative image in the coreset. We prove that our algorithm efficiently computes the smallest possible coreset under natural well-defined similarity metric and up to provably small approximation factor. The output visual summary is computed via a hierarchical tree of coresets for different parts of the image stream. This allows multi-resolution summarization (or a video summary of specified duration) in the batch setting and a memory-efficient incremental summary for the streaming case. Rohan Paul, Dan Feldman, Daniela Rus, Paul Newman 0001 |
ICRA | 4 |
| 2014 | Lighting invariant urban street classificationabstractIn this paper we propose the hybrid use of illuminant invariant and RGB images to perform image classification of urban scenes despite challenging variation in lighting conditions. Coping with lighting change (and the shadows thereby invoked) is a non-negotiable requirement for long term autonomy using vision. One aspect of this is the ability to reliably classify scene components in the presence of marked and often sudden changes in lighting. This is the focus of this paper. Posed with the task of classifying all parts in a scene from a full colour image, we propose that lighting invariant transforms can reduce the variability of the scene, resulting in a more reliable classification. We leverage the ideas of “data transfer” for classification, beginning with full colour images for obtaining candidate scene-level matches using global image descriptors. This is commonly followed by superpixellevel matching with local features. However, we show that if the RGB images are subjected to an illuminant invariant transform before computing the superpixel-level features, classification is significantly more robust to scene illumination effects. The approach is evaluated using three datasets. The first being our own dataset and the second being the KITTI dataset using manually generated ground truth for quantitative analysis. We qualitatively evaluate the method on a third custom dataset over a 750m trajectory. Ben Upcroft, Colin McManus, Winston Churchill, Will Maddern, Paul Newman 0001 |
ICRA | 5 |
| 2014 | LAPS-II: 6-DoF day and night visual localisation with prior 3D structure for autonomous road vehiclesabstractRobust and reliable visual localisation at any time of day is an essential component towards low-cost autonomy for road vehicles. We present a method to perform online 6-DoF visual localisation across a wide range of outdoor illumination conditions throughout the day and night using a 3D scene prior collected by a survey vehicle. We propose the use of a one-dimensional illumination invariant colour space which stems from modelling the spectral properties of the camera and scene illumination in conjunction. We combine our previous work on Localisation with Appearance of Prior Structure (LAPS) with this illumination invariant colour space to demonstrate a marked improvement in our ability to localise throughout the day compared to using a conventional RGB colour space. Our ultimate goal is robust and reliable any-time localisation — an attractive proposition for low-cost autonomy for road vehicles. Accordingly, we demonstrate our technique using 32km of data collected over a full 24-hour period from a road vehicle. Will Maddern, Alexander D. Stewart, Paul Newman 0001 |
Intelligent Vehicles Symposium | 3 |
| 2013 | Distraction suppression for vision-based pose estimation at city scalesabstractThis paper is concerned with the problem of egomotion estimation in highly dynamic, heavily cluttered urban environments over long periods of time. This is a challenging problem for vision-based systems because extreme scene movement caused by dynamic objects (e.g., enormous buses) can result in erroneous motion estimates. We describe two methods that combine 3D scene priors with vision sensors to generate background-likelihood images, which act as probability masks for objects that are not part of the scene prior. This results in a system that is able to cope with extreme scene motion, even when most of the image is obscured. We present results on real data collected in central London during rush hour and demonstrate the benefits of our techniques on a core navigation system - visual odometry. Colin McManus, Winston Churchill, Ashley Napier, Ben Davis, Paul Newman 0001 |
ICRA | 5 |
| 2013 | Cross-calibration of push-broom 2D LIDARs and cameras in natural scenesabstractThis paper addresses the problem of automatically estimating the relative pose between a push-broom LIDAR and a camera without the need for artificial calibration targets or other human intervention. Further we do not require the sensors to have an overlapping field of view, it is enough that they observe the same scene but at different times from a moving platform. Matching between sensor modalities is achieved without feature extraction. We present results from field trials which suggest that this new approach achieves an extrinsic calibration accuracy of millimeters in translation and deci-degrees in rotation. Ashley Napier, Peter I. Corke, Paul Newman 0001 |
ICRA | 3 |
| 2013 | Dealing with shadows: Capturing intrinsic scene appearance for image-based outdoor localisationabstractIn outdoor environments shadows are common. These typically strong visual features cause considerable change in the appearance of a place, and therefore confound vision-based localisation approaches. In this paper we describe how to convert a colour image of the scene to a greyscale invariant image where pixel values are a function of underlying material property not lighting. We summarise the theory of shadow invariant images and discuss the modelling and calibration issues which are important for non-ideal off-the-shelf colour cameras. We evaluate the technique with a commonly used robotic camera and an autonomous car operating in an outdoor environment, and show that it can outperform the use of ordinary greyscale images for the task of visual localisation. Peter I. Corke, Rohan Paul, Winston Churchill, Paul Newman 0001 |
IROS | 4 |
| 2013 | A roadwork scene signature based on the opponent colour modelabstractThe presence of roadworks greatly affects the validity of prior maps used for navigation by autonomous vehicles. This paper addresses the problem of quickly and robustly assessing the gist of traffic scenes for whether roadworks might be present. Without explicitly modelling individual roadwork indicators such as traffic cones, construction barriers or traffic signs, our method instead only exploits the engineered visual saliency of such objects. We draw inspiration from opponent colour vision in humans to formulate a novel roadwork scene signature based on an opponent spatial prior combined with gradient information. Finally, we apply our roadwork scene signature to the task of roadwork scene recognition, within a classification framework based on soft assignment vec-torization and RUSBoost. We evaluate our roadwork signature on real life data from our autonomous vehicle. Bonolo Mathibela, Ingmar Posner, Paul Newman 0001 |
IROS | 3 |
| 2013 | Continuous vehicle localisation using sparse 3D sensing, kernelised rényi distance and fast Gauss transformsabstractThis paper is about estimating a smooth, continuous-time trajectory of a vehicle relative to a prior 3D laser map. We pose the estimation problem as that of finding a sequence of Catmull-Rom splines which optimise the Kernelised Rényi Distance (KRD) between the prior map and live measurements from a 3D laser sensor. Our approach treats the laser measurements as a continual stream of data from a smoothly moving vehicle. We side-step entirely the segmentation and feature matching problems incumbent in traditional point cloud matching algorithms, relying instead on a smooth and well behaved objective function. Importantly our approach admits the exploitation of sensors with modest sampling rates - sensors that take seconds to densely sample the workspace. We show how by appropriate use of the Improved Fast Gauss Transform we can reduce the order of the estimation problem from quadratic (straight forward application of the KRD) to linear. Although in this paper we use 3D laser, our approach is also applicable to vehicles using 2D laser sensing or dense stereo. We demonstrate and evaluate the performance of our approach when estimating the full 6DOF continuous time pose of a road vehicle undertaking over 2.7km of outdoor travel. Mark Sheehan, Alastair Harrison, Paul Newman 0001 |
IROS | 3 |
| 2013 | A New Approach to Model-Free Tracking with 2D Lidar
Dominic Zeng Wang, Ingmar Posner, Paul Newman 0001 |
ISRR | 3 |
| 2013 | Toward automated driving in cities using close-to-market sensors: An overview of the V-Charge ProjectabstractFuture requirements for drastic reduction of CO2production and energy consumption will lead to significant changes in the way we see mobility in the years to come. However, the automotive industry has identified significant barriers to the adoption of electric vehicles, including reduced driving range and greatly increased refueling times. Automated cars have the potential to reduce the environmental impact of driving, and increase the safety of motor vehicle travel. The current state-of-the-art in vehicle automation requires a suite of expensive sensors. While the cost of these sensors is decreasing, integrating them into electric cars will increase the price and represent another barrier to adoption. The V-Charge Project, funded by the European Commission, seeks to address these problems simultaneously by developing an electric automated car, outfitted with close-to-market sensors, which is able to automate valet parking and recharging for integration into a future transportation system. The final goal is the demonstration of a fully operational system including automated navigation and parking. This paper presents an overview of the V-Charge system, from the platform setup to the mapping, perception, and planning sub-systems. Paul Timothy Furgale, Ulrich Schwesinger, Martin Rufli, Wojciech Derendarz, Hugo Grimmett, Peter Mühlfellner, Stefan Wonneberger, Julian Timpner, Stephan Rottmann, Bo Li 0018, Bastian Schmidt, Thien-Nghia Nguyen, Elena Cardarelli, Stefano Cattani, Stefan Bruning, Sven Horstmann, Martin Stellmacher, Holger Mielenz, Kevin Köser, Markus Beermann, Christian Häne, Lionel Heng, Gim Hee Lee, Friedrich Fraundorfer, René Iser, Rudolph Triebel, Ingmar Posner, Paul Newman 0001, Lars C. Wolf, Marc Pollefeys, Stefan Brosig, Jan Effertz, Cédric Pradalier, Roland Siegwart |
Intelligent Vehicles Symposium | 28 |
| 2013 | Risky Planning on Probabilistic Costmaps for Path Planning in Outdoor EnvironmentsabstractThis paper presents a framework for path planning over probabilistic costmaps of outdoor terrain that is compatible with fast grid-based planners such as A$^*$and D$^*$. We begin with an exemplar of how probabilistic costmaps may be constructed and then show how the a priori availability of such maps lends itself to the precomputation of exact probabilistic heuristics. In turn, the probabilistic nature of these heuristics allow the user to employ a bounded speed–accuracy tradeoff that characterizes the risk of paths returned not being of optimal shortest-path length. Results are shown which demonstrate that the method is able to closely approximate a probability distribution over the underlying exact distance and that efficiency increases on the order of 90% in terms of nodes expanded, and 60% in terms of search time over Euclidean distance heuristics, can be achieved. Liz Murphy, Paul Newman 0001 |
IEEE Trans. Robotics | 2 |
| 2012 | Parsing Outdoor Scenes from Streamed 3D Laser Data Using Online Clustering and Incremental Belief UpdatesabstractIn this paper, we address the problem of continually parsing a stream of 3D point cloud data acquired from a laser sensor mounted on a road vehicle. We leverage an online star clustering algorithm coupled with an incremental belief update in an evolving undirected graphical model. The fusion of these techniques allows the robot to parse streamed data and to continually improve its understanding of the world. The core competency produced is an ability to infer object classes from similarities based on appearance and shape features, and to concurrently combine that with a spatial smoothing algorithm incorporating geometric consistency. This formulation of feature-space star clustering modulating the potentials of a spatial graphical model is entirely novel. In our method, the two sources of information: feature similarity and geometrical consistency are fed continu- ally into the system, improving the belief over the class distributions as new data arrives. The algorithm obviates the need for hand-labeled training data and makes no apriori assumptions on the number or characteristics of object categories. Rather, they are learnt incrementally over time from streamed input data. In experiments per- formed on real 3D laser data from an outdoor scene, we show that our approach is capable of obtaining an ever- improving unsupervised scene categorization. Rudolph Triebel, Rohan Paul, Daniela Rus, Paul Newman 0001 |
AAAI | 4 |
| 2012 | Road vehicle localization with 2D push-broom LIDAR and 3D priorsabstractIn this paper we describe and demonstrate a method for precisely localizing a road vehicle using a single push-broom 2D laser scanner while leveraging a prior 3D survey. In contrast to conventional scan matching, our laser is oriented downwards, thus causing continual ground strike. Our method exploits this to produce a small 3D swathe of laser data which can be matched statistically within the 3D survey. This swathe generation is predicated upon time varying estimates of vehicle velocity. While in theory this data could be obtained from vehicle speedometers, in reality these instruments are biased and so we also provide a way to estimate this bias from survey data. We show that our low cost system consistently outperforms a high caliber integrated DGPS/IMU system over 26 km of driven path around a test site. Ian A. Baldwin, Paul Newman 0001 |
ICRA | 2 |
| 2012 | Practice makes perfect? Managing and leveraging visual experiences for lifelong navigationabstractThis paper is about long-term navigation in environments whose appearance changes over time - suddenly or gradually. We describe, implement and validate an approach which allows us to incrementally learn a model whose complexity varies naturally in accordance with variation of scene appearance. It allows us to leverage the state of the art in pose estimation to build over many runs, a world model of sufficient richness to allow simple localisation despite a large variation in conditions. As our robot repeatedly traverses its workspace, it accumulates distinct visual experiences that in concert, implicitly represent the scene variation - each experience captures a visual mode. When operating in a previously visited area, we continually try to localise in these previous experiences while simultaneously running an independent vision based pose estimation system. Failure to localise in a sufficient number of prior experiences indicates an insufficient model of the workspace and instigates the laying down of the live image sequence as a new distinct experience. In this way, over time we can capture the typical time varying appearance of an environment and the number of experiences required tends to a constant. Although we focus on vision as a primary sensor throughout, the ideas we present here are equally applicable to other sensor modalities. We demonstrate our approach working on a road vehicle operating over a three month period at different times of day, in different weather and lighting conditions. In all, we process over 136,000 frames captured from 37km of driving. Winston Churchill, Paul Newman 0001 |
ICRA | 2 |
| 2012 | Lost in translation (and rotation): Rapid extrinsic calibration for 2D and 3D LIDARsabstractThis paper describes a novel method for determining the extrinsic calibration parameters between 2D and 3D LIDAR sensors with respect to a vehicle base frame. To recover the calibration parameters we attempt to optimize the quality of a 3D point cloud produced by the vehicle as it traverses an unknown, unmodified environment. The point cloud quality metric is derived from Rényi Quadratic Entropy and quantifies the compactness of the point distribution using only a single tuning parameter. We also present a fast approximate method to reduce the computational requirements of the entropy evaluation, allowing unsupervised calibration in vast environments with millions of points. The algorithm is analyzed using real world data gathered in many locations, showing robust calibration performance and substantial speed improvements from the approximations. Will Maddern, Alastair Harrison, Paul Newman 0001 |
ICRA | 3 |
| 2012 | How was your day? Online visual workspace summaries using incremental clustering in topic spaceabstractSomeday mobile robots will operate continually. Day after day, they will be in receipt of a never ending stream of images. In anticipation of this, this paper is about having a mobile robot generate apt and compact summaries of its life experience. We consider a robot moving around its environment both revisiting and exploring, accruing images as it goes. We describe how we can choose a subset of images to summarise the robot's cumulative visual experience. Moreover we show how to do this such that the time cost of generating an summary is largely independent of the total number of images processed. No one day is harder to summarise than any other. Rohan Paul, Daniela Rus, Paul Newman 0001 |
ICRA | 3 |
| 2012 | LAPS - localisation using appearance of prior structure: 6-DoF monocular camera localisation using prior pointcloudsabstractThis paper is about pose estimation using monocular cameras with a 3D laser pointcloud as a workspace prior. We have in mind autonomous transport systems in which low cost vehicles equipped with monocular cameras are furnished with preprocessed 3D lidar workspaces surveys. Our inherently cross-modal approach offers robustness to changes in scene lighting and is computationally cheap. At the heart of our approach lies inference of camera motion by minimisation of the Normalised Information Distance (NID) between the appearance of 3D lidar data reprojected into overlapping images. Results are presented which demonstrate the applicability of this approach to the localisation of a camera against a lidar pointcloud using data gathered from a road vehicle. Alexander D. Stewart, Paul Newman 0001 |
ICRA | 2 |
| 2012 | What could move? Finding cars, pedestrians and bicyclists in 3D laser dataabstractThis paper tackles the problem of segmenting things that could move from 3D laser scans of urban scenes. In particular, we wish to detect instances of classes of interest in autonomous driving applications - cars, pedestrians and bicyclists - amongst significant background clutter. Our aim is to provide the layout of an end-to-end pipeline which, when fed by a raw stream of 3D data, produces distinct groups of points which can be fed to downstream classifiers for categorisation. We postulate that, for the specific classes considered in this work, solving a binary classification task (i.e. separating the data into foreground and background first) outperforms approaches that tackle the multi-class problem directly. This is confirmed using custom and third-party datasets gathered of urban street scenes. While our system is agnostic to the specific clustering algorithm deployed we explore the use of a Euclidean Minimum Spanning Tree for an end-to-end segmentation pipeline and devise a RANSAC-based edge selection criterion. Dominic Zeng Wang, Ingmar Posner, Paul Newman 0001 |
ICRA | 3 |
| 2012 | Laser-only road-vehicle localization with dual 2D push-broom LIDARS and 3D priorsabstractWe demonstrate the viability of using 2D LIDAR data as the sole means for accurate, robust, long-term road-vehicle localization within a prior map in a complex, dynamic real-world setting. We utilize a dual-LIDAR system - one oriented horizontally, in order to infer vehicle linear and rotational velocity, and one declined to capture a dense view of the surrounds - that allows us to estimate both velocity and position within a prior map. We show how probabilistically modelling the noisy local velocity estimates from the horizontal laser feed, fusing these estimates with data from the declined LIDAR to form a dense 3D swathe and matching this swathe statistically within a map will allow for robust, long-term position estimation. We accommodate estimation errors induced by passing vehicles, pedestrians, ground-strike etc., by learning a positional-dependent sensor model - that is, a sensor-model that varies spatially - and show that learning such a model for LIDAR data allows us to deal gracefully with the complexities of realworld data. We validate the concept over more than 9 kilometres of driven distance in and around the town of Woodstock, Oxfordshire. Ian A. Baldwin, Paul Newman 0001 |
IROS | 2 |
| 2012 | Semantic categorization of outdoor scenes with uncertainty estimates using multi-class gaussian process classificationabstractThis paper presents a novel semantic categorization method for 3D point cloud data using supervised, multiclass Gaussian Process (GP) classification. In contrast to other approaches, and particularly Support Vector Machines, which probably are the most used method for this task to date, GPs have the major advantage of providing informative uncertainty estimates about the resulting class labels. As we show in experiments, these uncertainty estimates can either be used to improve the classification by neglecting uncertain class labels or - more importantly - they can serve as an indication of the under-representation of certain classes in the training data. This means that GP classifiers are much better suited in a lifelong learning framework, where not all classes are represented initially, but instead new training data arrives during the operation of the robot. Rohan Paul, Rudolph Triebel, Daniela Rus, Paul Newman 0001 |
IROS | 4 |
| 2012 | Generation and exploitation of local orthographic imagery for road vehicle localisationabstractThis paper is about road vehicle localisation based on vision using synthesised local orthographic imagery. We exploit state of the art stereo visual odometry (VO) on our survey vehicle to generate high precision synthetic orthographic images of the road surface as would be seen from overhead. The fidelity and detail of these images far exceeds that of aerial photographs. When undertaking subsequent passes of the same route, the vehicle is localised against the survey vehicle's trajectory by maximising the mutual information between the synthetic orthographic images and live image streams. Thus we explicitly leverage the gross appearance of the workspace rather than a discrete set of point features. We test our technique on data gathered from a road vehicle and show that centimeter-level precision is possible without the complexity and instability of contemporary feature based techniques. Ashley Napier, Paul Newman 0001 |
Intelligent Vehicles Symposium | 2 |
| 2011 | TICSync: Knowing when things happenedabstractModern robotic systems are composed of many distributed processes sharing a common communications infrastructure. High bandwidth sensor data is often collected on one computer and served to many consumers. It is vital that every device on the network agrees on how time is measured. If not, sensor data may be at best inconsistent and at worst useless. Typical clocks in consumer grade PCs are highly inaccurate and temperature sensitive. We argue that traditional approaches to clock synchronization, such as the use of NTP are inappropriate in the robotics context. We present an extremely efficient algorithm for learning the mapping between distributed clocks, which typically achieves better than millisecond accuracy within just a few seconds. We also give a probabilistic analysis providing an upper-bound error estimate. Alastair Harrison, Paul Newman 0001 |
ICRA | 2 |
| 2011 | Hidden view synthesis using real-time visual SLAM for simplifying video surveillance analysisabstractUnderstanding and analysing video data from static or mobile surveillance cameras often requires knowledge of the scene and the camera placement. In this article, we provide a way to simplify the user's task of understanding the scene by rendering the camera view as if observed from the user's perspective by estimating his position using a real-time visual SLAM system. Augmenting the view is referred to as hidden view synthesis. Compared to previous work, the current approach improves by simplifying the setup and requiring minimal user input. This is achieved by building a map of the environment using a visual SLAM system and then registering the surveillance camera in this map. By exploiting the map, a different moving camera can render hidden views in real-time at 30Hz. We discuss some of the challenges remaining for full automation. Results are shown in an indoor environment for surveillance applications and outdoors with application to improved safety in transport. Christopher Mei, Eric Sommerlade, Gabe Sibley, Paul Newman 0001, Ian D. Reid 0001 |
ICRA | 4 |
| 2011 | Risky planning: Path planning over costmaps with a probabilistically bounded speed-accuracy tradeoffabstractThis paper is about generating plans over uncertain maps quickly. Our approach combines the ALT (A* search, landmarks and the triangle inequality) algorithm and risk heuristics to guide search over probabilistic cost maps. We build on previous work which generates probabilistic cost maps from aerial imagery and use these cost maps to precompute heuristics for searches such as A* and D* using the ALT technique. The resulting heuristics are probability distributions. We can speed up and direct search by characterising the risk we are prepared to take in gaining search efficiency while sacrificing optimal path length. Results are shown which demonstrate that ALT provides a good approximation to the true distribution of the heuristic, and which show efficiency increases in excess of 70% over normal heuristic search methods. Liz Murphy, Paul Newman 0001 |
ICRA | 2 |
| 2011 | Self help: Seeking out perplexing images for ever improving navigationabstractThis paper is a demonstration of how a robot can, through introspection and then targeted data retrieval, improve its own performance. It is a step in the direction of lifelong learning and adaptation and is motivated by the desire to build robots that have plastic competencies which are not baked in. They should react to and benefit from use. We consider a particular instantiation of this problem in the context of place recognition. Based on a topic based probabilistic model of images, we use a measure of perplexity to evaluate how well a working set of background images explain the robot's online view of the world. Offline, the robot then searches an external resource to seek out additional background images that bolster its ability to localise in its environment when used next. In this way the robot adapts and improves performance through use. Rohan Paul, Paul Newman 0001 |
ICRA | 2 |
| 2011 | Choosing where to go: Complete 3D exploration with stereoabstractThis paper is about the autonomous acquisition of detailed 3D maps of a-priori unknown environments using a stereo camera - it is about choosing where to go. Our approach hinges upon a boundary value constrained partial differential equation (PDE) - the solution of which provides a scalar field guaranteed to have no local minima. This scalar field is trivially transformed into a vector field in which following lines of max flow causes provably complete exploration of the environment in full 6 degrees of freedom (6-DOF). We use a SLAM system to infer the position of a stereo pair in real time and fused stereo depth maps to generate the boundary conditions which drive exploration. Our exploration algorithm is parameter free, is as applicable to 3D laser data as it is to stereo, is real time and is guaranteed to deliver complete exploration. We show empirically that it performs better than oft-used frontier based approaches and demonstrate our system working with real and simulated data. Robert Shade, Paul Newman 0001 |
ICRA | 2 |
| 2011 | Adaptive Data Compression for Robot Perception
Mike Smith 0002, Ingmar Posner, Paul Newman 0001 |
IJCAI | 3 |
| 2011 | Choosing landmarks for risky planningabstractThis work examines the effect of landmark placement on the efficiency and accuracy of risk-bounded searches over probabilistic costmaps for mobile robot path planning. In previous work, risk-bounded searches were shown to offer in excess of 70% efficiency increases over normal heuristic search methods. The technique relies on precomputing distance estimates to landmarks which are then used to produce probability distributions over exact heuristics for use in heuristic searches such as A* and D*. The location and number of these landmarks therefore influence greatly the efficiency of the search and the quality of the risk bounds. Here four new methods of selecting landmarks for risk based search are evaluated. Results are shown which demonstrate that landmark selection needs to take into account the centrality of the landmark, and that diminishing rewards are obtained from using large numbers of landmarks. Liz Murphy, Peter I. Corke, Paul Newman 0001 |
IROS | 3 |
| 2011 | RSLAM: A System for Large-Scale Mapping in Constant-Time Using Stereo
Christopher Mei, Gabe Sibley, Mark Joseph Cummins, Paul Newman 0001, Ian D. Reid 0001 |
Int. J. Comput. Vis. | 4 |
| 2010 | FAB-MAP: Appearance-Based Place Recognition and Mapping using a Learned Visual Vocabulary Model
Mark Joseph Cummins, Paul Newman 0001 |
ICML | 2 |
| 2010 | Planning most-likely paths from overhead imageryabstractThis paper is about planning paths from overhead imagery, the novelty of which is taking explicit account of uncertainty in terrain classification and spatial variation in terrain cost. The image is first classified using a multi-class Gaussian Process Classifier which provides probabilities of class membership at each location in the image. The probability of class membership at a particular grid location is then combined with a terrain cost evaluated at that location using a spatial Gaussian process. The resulting cost function is, in turn, passed to a planner. This allows both the uncertainty in terrain classification and spatial variations in terrain costs to be incorporated into the planned path. Because the cost of traversing a grid cell is now a probability density rather than a single scalar value, we can produce not only the most-likely shortest path between points on the map, but also sample from the cost map to produce a distribution of paths between the points. Results are shown in the form of planned paths over aerial maps, these paths are shown to vary in response to local variations in terrain cost. Elizabeth Murphy, Paul Newman 0001 |
ICRA | 2 |
| 2010 | FAB-MAP 3D: Topological mapping with spatial and visual appearanceabstractThis paper describes a probabilistic framework for appearance based navigation and mapping using spatial and visual appearance data. Like much recent work on appearance based navigation we adopt a bag-of-words approach in which positive or negative observations of visual words in a scene are used to discriminate between already visited and new places. In this paper we add an important extra dimension to the approach. We explicitly model the spatial distribution of visual words as a random graph in which nodes are visual words and edges are distributions over distances. Care is taken to ensure that the spatial model is able to capture the multi-modal distributions of inter-word spacing and account for sensor errors both in word detection and distances. Crucially, these inter-word distances are viewpoint invariant and collectively constitute strong place signatures and hence the impact of using both spatial and visual appearance is marked. We provide results illustrating a tremendous increase in precision-recall area compared to a state-of-the-art visual appearance only systems. Rohan Paul, Paul Newman 0001 |
ICRA | 2 |
| 2010 | Discovering and mapping complete surfaces with stereoabstractThis paper is about the automated discovery and mapping of surfaces using a stereo pair. We begin with the observation that for any workspace which is topologically connected (i.e. does not contain free flying islands) there exists a single surface that covers the entirety of the workspace. We call this surface the covering surface. We assume that while this surface is complex and self intersecting every point on it can be imaged from a suitable camera pose and furthermore that it is locally smooth at some finite scale - it is a manifold. We show how by representing the covering surface as a non-planar graph of observed pixels we are able to plan new views and importantly fuse disparity maps from multiple views. The resulting graph can be lifted to 3D to yield a full scene reconstruction. Robert Shade, Paul Newman 0001 |
ICRA | 2 |
| 2010 | Planes, trains and automobiles - autonomy for the modern robotabstractWe are concerned with enabling truly large scale autonomous navigation in typical human environments. To this end we describe the acquisition and modeling of large urban spaces from data that reflects human sensory input. Over 181GB of image and inertial data are captured using head-mounted stereo cameras. This data is processed into a relative map covering 121 km of Southern England. We point out the numerous challenges we encounter, and highlight in particular the problem of undetected ego-motion, which occurs when the robot finds itself on-or-within a moving frame of reference. In contrast to global-frame representations, we find that the continuous relative representation naturally accommodates moving-reference-frames - without having to identify them first, and without inconsistency. Within a moving-reference-frame, and without drift-less global exteroceptive sensing, motion with respect to the global-frame is effectively unobservable. This underlying truth drives us towards relative topometric solutions like relative bundle adjustment (RBA), which has no problem representing distance and metric Euclidean structure, yet does not suffer inconsistency introduced by the attempt to solve in the global-frame. Gabe Sibley, Christopher Mei, Ian D. Reid 0001, Paul Newman 0001 |
ICRA | 4 |
| 2010 | Non-parametric learning for natural plan generationabstractWe present a novel way to learn sampling distributions for sampling-based motion planners by making use of expert data. We learn an estimate (in a non-parametric setting) of sample densities around semantic regions of interest, and incorporate these learned distributions into a sampling-based planner to produce natural plans. Our motivation is that certain aspects of the workspace have a local influence on planning strategies, which is dependent both on where, and what, they are. In the event that learning the density estimate of the training data is impractical in the original feature space, we utilize a non-linear dimensionality-reduction technique and perform density estimation on a lower-dimensional embedding. Samples are then lifted from this embedded density into the original feature space, producing samples that still well approximate the original distribution. A goal of this work is to learn how various features in the environment influence the behavior of experts - for example, how pedestrian crossings, traffic signals and so on affect drivers. We show that learning sampling distributions from expert trajectory data around these semantic regions leads to more natural paths that are measurably closer to those of an expert. We demonstrate the feasibility of the technique in various scenarios for a virtual car-like robotic vehicle and a simple manipulator, contrasting the differences in planned trajectories of the semantically-biased distributions with conventional techniques. Ian A. Baldwin, Paul Newman 0001 |
IROS | 2 |
| 2010 | Closing loops without placesabstractThis paper proposes a new topo-metric representation of the world based on co-visibility that simplifies data association and improves the performance of appearance-based recognition. We introduce the concept of dynamic bagof-words, which is a novel form of query expansion based on finding cliques in the landmark co-visibility graph. The proposed approach avoids the - often arbitrary - discretisation of space from the robot's trajectory that is common to most image-based loop closure algorithms. Instead we show that reasoning on sets of co-visible landmarks leads to a simple model that out-performs pose-based or view-based approaches. Using real and simulated imagery, we demonstrate that dynamic bag-of-words query expansion can improve precision and recall for appearance-based localisation. Christopher Mei, Gabe Sibley, Paul Newman 0001 |
IROS | 3 |
| 2010 | Using text-spotting to query the worldabstractThe world we live in is labeled extensively for the benefit of humans. Yet, to date, robots have made little use of human readable text as a resource. In this paper we aim to draw attention to text as a readily available source of semantic information in robotics by implementing a system which allows robots to read visible text in natural scene images and to use this knowledge to interpret the content of a given scene. The reliable detection and parsing of text in natural scene images is an active area of research and remains a non-trivial problem. We extend a commonly adopted approach based on boosting for the detection and optical character recognition (OCR) for the parsing of text by a probabilistic error correction scheme incorporating a sensor-model for our pipeline. In order to interpret the scene content we introduce a generative model which explains spotted text in terms of arbitrary search terms. This allows the robot to estimate the relevance of a given scene with respect to arbitrary queries such as, for example, whether it is looking at a bank or a restaurant. We present results from images recorded by a robot in a busy cityscape. Ingmar Posner, Peter I. Corke, Paul Newman 0001 |
IROS | 3 |
| 2010 | Accelerating FAB-MAP With Concentration InequalitiesabstractWe outline an approach for using concentration inequalities to perform rapid approximate multi-hypothesis testing. In a scenario where multiple hypotheses are ranked according to a large set of features, our scheme improves the efficiency of selecting the best hypothesis by providing a “bail-out threshold” at which unpromising hypotheses can be excluded from further evaluation. We show how concentration inequalities can be used to derive principled bail-out thresholds, subject to a user-specified error tolerance. The technique is similar to the sequential probability ratio test, but is applicable in more general conditions. We apply the technique to improve the speed of the fast-appearance-based mapping system for appearance-based place recognition and mapping. The speed increase provided by the new approach is data dependent, but we demonstrate speed improvements of between 25x - 50x on real data, with only a slight degradation in accuracy. Mark Joseph Cummins, Paul Newman 0001 |
IEEE Trans. Robotics | 2 |
| 2009 | A Constant-Time Efficient Stereo SLAM SystemabstractContinuous, real-time mapping of an environment using a camera requires a constant-time estimation engine. This rules out optimal global solving such as bundle adjustment. In this article, we investigate the precision that can be achieved with only local estimation of motion and structure provided by a stereo pair. We introduce a simple but novel representation of the environment in terms of a sequence of relative locations. We demonstrate precise local mapping and easy navigation using the relative map, and importantly show that this can be done without requiring a global minimisation after loop closure. We discuss some of the issues that arise from using a relative representation, and evaluate our system on long sequences processed at a constant 30-45 Hz, obtaining precisions down to a few metres over distances of a few kilometres. Christopher Mei, Gabe Sibley, Mark Joseph Cummins, Paul Newman 0001, Ian D. Reid 0001 |
BMVC | 4 |
| 2008 | Accelerated appearance-only SLAMabstractThis paper describes a probabilistic bail-out condition for multihypothesis testing based on Bennett's inequality. We investigate the use of the test for increasing the speed of an appearance-only SLAM system where locations are recognised on the basis of their sensory appearance. The bail-out condition yields speed increases between 25x-50x on real data, with only slight degradation in accuracy. We demonstrate the system performing real-time loop closure detection on a mobile robot over multiple-kilometre paths in initially unknown outdoor environments. Mark Joseph Cummins, Paul Newman 0001 |
ICRA | 2 |
| 2008 | High quality 3D laser ranging under general vehicle motionabstractThis paper describes an end-to-end system capable of generating high-quality 3D point clouds from the popular LMS200 laser on a continuously moving platform. We describe the hardware, data capture, calibration and data stream processing we have developed which yields remarkable detail in the generated point clouds of urban scenes. Given the increasing interest in outdoor 3D navigation and scene reconstruction by mobile platforms, our aim is to provide a level of hardware and algorithmic detail suitable for replication of our system by interested parties who do not wish to invest in dedicated 3D laser rangers. Alastair Harrison, Paul Newman 0001 |
ICRA | 2 |
| 2008 | Using incomplete online metric maps for topological exploration with the Gap Navigation TreeabstractThis paper presents a general, global approach to the problem of robot exploration, utilizing a topological data structure to guide an underlying Simultaneous Localization and Mapping (SLAM) process. A Gap Navigation Tree (GNT) is used to motivate global target selection and occluded regions of the environment (called "gaps") are tracked probabilistically. The process of map construction and the motion of the vehicle alters both the shape and location of these regions. The use of online mapping is shown to reduce the difficulties in implementing the GNT. Elizabeth Murphy, Paul Newman 0001 |
ICRA | 2 |
| 2008 | An image-to-map loop closing method for monocular SLAMabstractIn this paper we present a loop closure method for a handheld single-camera SLAM system based on our previous work on relocalization. By finding correspondences between the current image and the map, our system is able to reliably detect loop closures. We compare our algorithm to existing techniques for loop closure in single-camera SLAM based on both image-to-image and map-to-map correspondences and discuss both the reliability and suitability of each algorithm in the context of monocular SLAM. Brian Patrick Williams, Mark Joseph Cummins, José Neira, Paul Newman 0001, Ian D. Reid 0001, Juan D. Tardós |
IROS | 4 |
| 2007 | Probabilistic Appearance Based Navigation and Loop ClosingabstractThis paper describes a probabilistic framework for navigation using only appearance data. By learning a generative model of appearance, we can compute not only the similarity of two observations, but also the probability that they originate from the same location, and hence compute a pdf over observer location. We do not limit ourselves to the kidnapped robot problem (localizing in a known map), but admit the possibility that observations may come from previously unvisited places. The principled probabilistic approach we develop allows us to explicitly account for the perceptual aliasing in the environment - identical but indistinctive observations receive a low probability of having come from the same place. Our algorithm complexity is linear in the number of places, and is particularly suitable for online loop closure detection in mobile robotics. Mark Joseph Cummins, Paul Newman 0001 |
ICRA | 2 |
| 2007 | Describing Composite Urban WorkspacesabstractIn this paper we present an appearance-based method for augmenting maps of outdoor urban environments with higher-order, semantic labels. Our motivation is to increase the value and utility of the typically low-level representations built by contemporary SLAM algorithms. A supervised learning scheme is employed to train a set of classifiers to respond to common scene attributes given a mixture of geometric and visual scene information. The union of classifier responses yields a composite description of the local workspace. We apply our method to three large data sets Ingmar Posner, Derik Schröter, Paul Newman 0001 |
ICRA | 3 |
| 2007 | Describing, Navigating and Recognising Urban Spaces - Building an End-to-End SLAM System
Paul Newman 0001, Manjari Chandran-Ramesh, Dave Cole, Mark Joseph Cummins, Alastair Harrison, Ingmar Posner, Derik Schröter |
ISRR | 1 |
| 2007 | Detecting Loop Closure with Scene Sequences
Kin Leong Ho, Paul Newman 0001 |
Int. J. Comput. Vis. | 2 |
| 2006 | Navigation of Unmanned Marine Vehicles in Accordance with the Rules of the RoadabstractThis paper is concerned with the in-field autonomous operation of unmanned marine vehicles in accordance with convention for safe and proper collision avoidance as prescribed by the coast guard collision regulations (COLREGS). These rules are written to train and guide safe human operation of marine vehicles and are heavily dependent on human common sense in determining rule applicability as well as rule execution, especially when multiple rules apply simultaneously. To capture the flexibility exploited by humans, this work applies a novel method of multi-objective optimization, interval programming, in a behavior-based control framework for representing the navigation rules, as well as task behaviors, in a way that achieves simultaneous optimal satisfaction. We present experimental validation of this approach using multiple autonomous surface craft. This work represents the first in-field demonstration of multiobjective optimization applied to autonomous COLREGS-based marine vehicle navigation Michael R. Benjamin, Joseph A. Curcio, John J. Leonard, Paul Newman 0001 |
ICRA | 4 |
| 2006 | Multi-objective Optimization of Sensor Quality with Efficient Marine Vehicle Task ExecutionabstractThis paper describes the in-field operation of two interacting autonomous marine vehicles to demonstrate the suitability of interval programming (IvP), a novel mathematical model for multiple-objective optimization. Broadly speaking, IvP coordinates competing control needs such as primary task execution that depends on a sufficient position estimate, and vehicle maneuvers that will improve that position estimate. In this work, vehicles cooperate to improve their position estimates using a sequence of vehicle-to-vehicle range estimates from acoustic modems. Coordinating primary task execution and sensor quality maintenance is a ubiquitous problem, especially in underwater marine vehicles. This work represents the first use of multiobjective optimization in a behavior-based architecture to address this problem Michael R. Benjamin, Matthew Grund, Paul Newman 0001 |
ICRA | 3 |
| 2006 | Using Laser Range Data for 3D SLAM in Outdoor EnvironmentsabstractTraditional simultaneous localization and mapping (SLAM) algorithms have been used to great effect in flat, indoor environments such as corridors and offices. We demonstrate that with a few augmentations, existing 2D SLAM technology can be extended to perform full 3D SLAM in less benign, outdoor, undulating environments. In particular, we use data acquired with a 3D laser range finder. We use a simple segmentation algorithm to separate the data stream into distinct point clouds, each referenced to a vehicle position. The SLAM technique we then adopt inherits much from 2D delayed state (or scan-matching) SLAM in that the state vector is an ever growing stack of past vehicle positions and inter-scan registrations are used to form measurements between them. The registration algorithm used is a novel combination of previous techniques carefully balancing the need for maximally wide convergence basins, robustness and speed. In addition, we introduce a novel post-registration classification technique to detect matches which have converged to incorrect local minima David M. Cole, Paul Newman 0001 |
ICRA | 2 |
| 2006 | Outdoor SLAM using Visual Appearance and Laser RangingabstractThis paper describes a 3D SLAM system using information from an actuated laser scanner and camera installed on a mobile robot. The laser samples the local geometry of the environment and is used to incrementally build a 3D point-cloud map of the workspace. Sequences of images from the camera are used to detect loop closure events (without reference to the internal estimates of vehicle location) using a novel appearance-based retrieval system. The loop closure detection is robust to repetitive visual structure and provides a probabilistic measure of confidence. The images suggesting loop closure are then further processed with their corresponding local laser scans to yield putative Euclidean image-image transformations. We show how naive application of this transformation to effect the loop closure can lead to catastrophic linearization errors and go on to describe a way in which gross, pre-loop closing errors can be successfully annulled. We demonstrate our system working in a challenging, outdoor setting containing substantial loops and beguiling, gently curving traversals. The results are overlaid on an aerial image to provide a ground truth comparison with the estimated map. The paper concludes with an extension into the multi-robot domain in which 3D maps resulting from distinct SLAM sessions (no common reference frame) are combined without recourse to mutual observation Paul Newman 0001, David M. Cole, Kin Leong Ho |
ICRA | 1 |
| 2006 | Motion Estimation from Map Quality with Millimeter Wave RadarabstractSimultaneous localization and mapping (SLAM) builds maps of a priori unknown environments. Whilst this key mobile robotic competency continues to receive substantial attention, less attention has been paid to assessing the quality of the resulting maps. This paper proposes a way to quantify the intrinsic quality of point-cloud maps built from a stream of range bearing measurements. It does so by considering both the temporal and spatial distribution of the points within the map. One of the causes of unsatisfactory maps is the execution of unmodelled or poorly sensed vehicle manoeuvres. In this paper we show that by maximizing the quality of the map as a function of a motion parameterization, the vehicle motion can be recovered while correcting the map at the same time. In contrast to typical scan matching techniques, we do not rely on segmentation of the measurement stream into two separate "scans"; Instead we treat the measurement sequence as a continuous signal. We illustrate the efficacy of this approach by processing range data from a 77 GHz millimeter wave radar that completes 2 rotations per second. We show that despite this acquisition speed being commensurate with vehicle rotation rates, we are able to extract the underlying vehicle motion and yield crisp, well aligned point clouds Manjari Chandran, Paul Newman 0001 |
IROS | 2 |
| 2005 | SLAM-Loop Closing with Visually Salient FeaturesabstractWithin the context of Simultaneous Localisation and Mapping (SLAM), “loop closing” is the task of deciding whether or not a vehicle has, after an excursion of arbitrary length, returned to a previously visited area. Reliable loop closing is both essential and hard. It is without doubt one of the greatest impediments to long term, robust SLAM. This paper illustrates how visual features, used in conjunction with scanning laser data, can be used to a great advantage. We use the notion of visual saliency to focus the selection of suitable (affine invariant) image-feature descriptors for storage in a database. When queried with a recently taken image the database returns the capture time of matching images. This time information is used to discover loop closing events. Crucially this is achieved independently of estimated map and vehicle location. We integrate the above technique into a SLAM algorithm using delayed vehicle states and scan matching to form interpose geometric constraints. We present initial results using this system to close loops (around 100m) in an indoor environment. Paul Newman 0001, Kin Leong Ho |
ICRA | 1 |
| 2005 | Session Overview Simultaneous Localisation and Mapping
Paul Newman 0001, Henrik I. Christensen |
ISRR | 1 |
| 2003 | An atlas framework for scalable mappingabstractThis paper describes Atlas, a hybrid metrical/topological approach to SLAM that achieves efficient mapping of large-scale environments. The representation is a graph of coordinate frames, with each vertex in the graph representing a local frame, and each edge representing the transformation between adjacent frames. In each frame, we build a map that captures the local environment and the current robot pose along with the uncertainties of each. Each map's uncertainties are modeled with respect to its own frame. Probabilities of entities with respect to arbitrary frames are generated by following a path formed by the edges between adjacent frames, computed via Dijkstra's shortest path algorithm. Loop closing is achieved via an efficient map matching algorithm. We demonstrate the technique running in real-time in a large indoor structured environment (2.2 km path length) with multiple nested loops using laser or ultrasonic ranging sensors. Michael Bosse, Paul Newman 0001, John J. Leonard, Martin Soika, Wendelin Feiten, Seth J. Teller |
ICRA | 2 |
| 2003 | Autonomous feature-based explorationabstractThis paper presents an algorithm for feature-based exploration of a priori unknown environments. We aim to build a robot that, unsupervised, plans its motion such that it continually increases both the spatial extent and detail of its world model - its map. We present a method by which the planned motion at any instant is motivated by the geometric, spatial and stochastic characteristics of the current map. In particular each feature within the map is responsible for determining nearby unexplored areas that if visited are likely to constitute exploration. We assume that the location of the features is uncertain and represented by a set of probability distribution functions (pdfs). These distributions are used in conjunction with the robot path history to determine a robot trajectory suited to exploration. We show results that demonstrate the algorithm providing real-time exploration of a mobile robot in an unknown environment. Paul Newman 0001, Michael Bosse, John J. Leonard |
ICRA | 1 |
| 2003 | Pure range-only sub-sea SLAMabstractThis paper is about using range-only data to navigate an autonomous underwater vehicle (AUV). We assume the vehicle is equipped with conventional long base line (LBL) transceiver which measures acoustic time of flights (TOFs) between vehicle and small submerged transponders. Using only range data and no prior information other than approximate water column depth, we solve for both transponder location and vehicle trajectory. Results are given using data from a AUV operating in shallow water. A ground truth comparison is made with surveyed transponder locations and trajectory estimates from an on-board Doppler/Compass/LBL derived navigation filter. Paul Newman 0001, John J. Leonard |
ICRA | 1 |
| 2003 | Consistent, Convergent, and Constant-Time SLAM
John J. Leonard, Paul Newman 0001 |
IJCAI | 2 |
| 2003 | Towards Constant-Time SLAM on an Autonomous Underwater Vehicle Using Synthetic Aperture Sonar
Paul Newman 0001, John J. Leonard, Richard J. Rikoski |
ISRR | 1 |
| 2002 | Cooperative Concurrent Mapping and LocalizationabstractAutonomous vehicles require the ability to build maps of an unknown environment while concurrently using these maps for navigation. Current algorithms for this concurrent mapping and localization (CML) problem have been implemented for single vehicles, but do not account for extra positional information available when multiple vehicles operate simultaneously. Multiple vehicles have the potential to map an environment more quickly and robustly than a single vehicle. This paper presents a cooperative CML algorithm that merges sensor and navigation information from multiple autonomous vehicles. The algorithm presented is based on stochastic estimation and uses a feature-based approach to extract landmarks from the environment. The theoretical framework for the collaborative CML algorithm is presented, and a convergence theorem central to the cooperative CML problem. is proved for the first time. This theorem quantifies the performance gains of collaboration, allowing for determination of the number of cooperating vehicles required to accomplish a task. A simulated implementation of the collaborative CML algorithm demonstrates substantial performance improvement over non-cooperative CML. John W. Fenwick, Paul Newman 0001, John J. Leonard |
ICRA | 2 |
| 2002 | Explore and Return: Experimental Validation of Real-Time Concurrent Mapping and LocalizationabstractThis paper describes a real-time implementation of feature-based concurrent mapping and localization (CML) running on a mobile robot in a dynamic indoor environment. Novel characteristics of this work include: (1) a hierarchical representation of uncertain geometric relationships that extends the SPMap framework, (2) use of robust statistics to perform extraction of line segments from laser data in real-time, and (3) the integration of CML with a "roadmap" path planning method for autonomous trajectory execution. These Innovations are combined to demonstrate the ability for a mobile robot to autonomously return back to its starting position within a few centimeters of precision, despite the presence of numerous people walking through the environment. Paul Newman 0001, John J. Leonard, Juan D. Tardós, José Neira |
ICRA | 1 |
| 2002 | Stochastic Mapping FrameworksabstractStochastic mapping is an approach to the concurrent mapping and localization problem. The approach is powerful because the feature and robot states are explicitly correlated. Improving the estimate of any state automatically improves the estimates of correlated states. This paper describes a number of extensions to the stochastic mapping framework, which are made possible by the incorporation of past vehicle states into the state vector to explicitly represent the robot's trajectory. Having access to past robot states simplifies the mapping, navigation, and cooperation. Experimental results using sonar data are presented. Richard J. Rikoski, John J. Leonard, Paul Newman 0001 |
ICRA | 3 |
| 2001 | Towards Robust Data Association and Feature Modeling for Concurrent Mapping and Localization
John J. Leonard, Paul Newman 0001, Richard J. Rikoski, José Neira, Juan D. Tardós |
ISRR | 2 |
| 2001 | A solution to the simultaneous localization and map building (SLAM) problemabstractThe simultaneous localization and map building (SLAM) problem asks if it is possible for an autonomous vehicle to start in an unknown location in an unknown environment and then to incrementally build a map of this environment while simultaneously using this map to compute absolute vehicle location. Starting from estimation-theoretic foundations of this problem, the paper proves that a solution to the SLAM problem is indeed possible. The underlying structure of the SLAM problem is first elucidated. A proof that the estimated map converges monotonically to a relative map with zero uncertainty is then developed. It is then shown that the absolute accuracy of the map and the vehicle location reach a lower bound defined only by the initial vehicle uncertainty. Together, these results show that it is possible for an autonomous vehicle to start in an unknown location in an unknown environment and, using relative observations only, incrementally build a perfect map of the world and to compute simultaneously a bounded estimate of vehicle location. The paper also describes a substantial implementation of the SLAM algorithm on a vehicle operating in an outdoor environment using millimeter-wave radar to provide relative map observations. This implementation is used to demonstrate how some key issues such as map management and data association can be handled in a practical environment. The results obtained are cross-compared with absolute locations of the map landmarks obtained by surveying. In conclusion, the paper discusses a number of key issues raised by the solution to the SLAM problem including suboptimal map-building algorithms and map management. Gamini Dissanayake, Paul Newman 0001, Steve Clark, Hugh F. Durrant-Whyte, M. Csorba |
IEEE Trans. Robotics Autom. | 2 |
| 2000 | Autonomous Underwater Simultaneous Localisation and Map BuildingabstractWe present results of the application of a simultaneous localisation and map building (SLAM) algorithm to estimate the motion of a submersible vehicle. Scans obtained from an on-board sonar are processed to extract stable point features in the environment. These point features are then used to build up a map of the environment while simultaneously providing estimates of the vehicle location. Results are shown from deployment in a swimming pool at the University of Sydney as well as from field trials in a natural environment along Sydney's coast. This work represents the first instance of a deployable underwater implementation of the SLAM algorithm. Stefan B. Williams, Paul Newman 0001, Gamini Dissanayake, Hugh F. Durrant-Whyte |
ICRA | 2 |
| 1998 | Using Sonar in Terrain-Aided Underwater NavigationabstractIn many ways autonomous navigation of underwater vehicles is a 'holy grail' of subsea robotics. For an AUV to continually estimate its pose (position and orientation) and operate within an initially unknown environment a solution is required to the simultaneous map building and localisation or 'SLAM' problem. The paper describes and investigates an inertial and terrain based approach to this problem. The fusion of body frame inertial and sonar based world frame feature information can form part of a robust navigation algorithm suitable for implementation in both artificial and natural environments. Particular attention is given to the role of the sonar and its ability to detect and track terrain features. Paul Newman 0001, Hugh F. Durrant-Whyte |
ICRA | 1 |