VLDB 2026 Research / reviewers in the wild / expert
Matthew Gadd
dblp:164/8450
· DBLP profile ↗
21ranked-venue papers
4as first author
14since 2021 · last 2024
0000-0001-9447-8619ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 3 first-author · 12 since 2021Systems, architecture and hardware · 15 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | That's My Point: Compact Object-centric LiDAR Pose Estimation for Large-scale Outdoor LocalisationabstractThis paper is about 3D pose estimation on LiDAR scans with extremely minimal storage requirements to enable scalable mapping and localisation. We achieve this by clustering all points of segmented scans into semantic objects and representing them only with their respective centroid and semantic class. In this way, each LiDAR scan is reduced to a compact collection of four-number vectors. This abstracts away important structural information from the scenes, which is crucial for traditional registration approaches. To mitigate this, we introduce an object-matching network based on self- and cross-correlation that captures geometric and semantic relationships between entities. The respective matches allow us to recover the relative transformation between scans through weighted Singular Value Decomposition (SVD) and RANdom SAmple Consensus (RANSAC). We demonstrate that such representation is sufficient for metric localisation by registering point clouds taken under different viewpoints on the KITTI dataset, and at different periods of time localising between KITTI and KITTI-360. We achieve accurate metric estimates comparable with state-of-the-art methods with almost half the representation size, specifically 1.33 kB on average. Georgi Pramatarov, Matthew Gadd, Paul Newman 0001, Daniele De Martini |
ICRA | 2 |
| 2024 | VDNA-PR: Using General Dataset Representations for Robust Sequential Visual Place RecognitionabstractThis paper adapts a general dataset representation technique to produce robust Visual Place Recognition (VPR) descriptors, crucial to enable real-world mobile robot localisation. Two parallel lines of work on VPR have shown, on one side, that general-purpose off-the-shelf feature representations can provide robustness to domain shifts, and, on the other, that fused information from sequences of images improves performance. In our recent work on measuring domain gaps between image datasets, we proposed a Visual Distribution of Neuron Activations (VDNA) representation to represent datasets of images. This representation can naturally handle image sequences and provides a general and granular feature representation derived from a general-purpose model. Moreover, our representation is based on tracking neuron activation values over the list of images to represent and is not limited to a particular neural network layer, therefore having access to high- and low-level concepts. This work shows how VDNAs can be used for VPR by learning a very lightweight and simple encoder to generate task-specific descriptors. Our experiments show that our representation can allow for better robustness than current solutions to serious domain shifts away from the training data distribution, such as to indoor environments and aerial imagery. Benjamin Ramtoula, Daniele De Martini, Matthew Gadd, Paul Newman 0001 |
ICRA | 3 |
| 2024 | Masked γ-SSL: Learning Uncertainty Estimation via Masked Image ModelingabstractThis work proposes a semantic segmentation network that produces high-quality uncertainty estimates in a single forward pass. We exploit general representations from foundation models and unlabelled datasets through a Masked Image Modeling (MIM) approach, which is robust to augmentation hyper-parameters and simpler than previous techniques. For neural networks used in safety-critical applications, bias in the training data can lead to errors; therefore it is crucial to understand a network’s limitations at run time and act accordingly. To this end, we test our proposed method on a number of test domains including the SAX Segmentation benchmark, which includes labelled test data from dense urban, rural and off-road driving domains. The proposed method consistently outperforms uncertainty estimation and Out-of-Distribution (OoD) techniques on this difficult benchmark. David S. W. Williams, Matthew Gadd, Paul Newman 0001, Daniele De Martini |
ICRA | 2 |
| 2024 | NeuralFloors++: Consistent Street-Level Scene Generation From BEV Semantic MapsabstractLearning autonomous driving capabilities requires diverse and realistic training data. This has led to exploring generative techniques as an alternative to real-world data collection. In this paper we propose a method for synthesising photo-realistic urban driving scenes, along with semantic, instance and depth ground-truth. Our model relies on Bird’s Eye View (BEV) representations due to their compositionality and scene content control capabilities, reducing the need for traditional simulators. We employ a two-stage process: first, a 3D scene representation is extracted from BEV semantic, instance and style maps using a neural field. After rendering the semantic, instance, depth and style maps from a ground-view perspective, a second stage based on a diffusion model is used to generate the photo-realistic scene. We extend our prior work - NeuralFloors, to include multiple-view outputs, style manipulation for finer control at the object level through instance-wise style maps and cross-frame consistency via auto-regressive training. The proposed system is evaluated extensively on the KITTI-360 dataset, showing improved realism and semantic alignment for generated images. Valentina Musat, Daniele De Martini, Matthew Gadd, Paul Newman 0001 |
IROS | 3 |
| 2024 | OORD: The Oxford Offroad Radar DatasetabstractThere is a growing academic interest as well as commercial exploitation of millimetre-wave scanning radar for autonomous vehicle localisation and scene understanding. Although several datasets to support this research area have been released, they are primarily focused on urban or semi-urban environments. Nevertheless, rugged offroad deployments are important application areas which also present unique challenges and opportunities for this sensor technology. Therefore, the Oxford Offroad Radar Dataset (OORD) presents data collected in the rugged Scottish highlands in extreme weather. The radar data we offer to the community are accompanied by GPS/INS reference – to further stimulate research in radar place recognition. In total we release over 90 GiB of radar scans as well as GPS and IMU readings by driving a diverse set of four routes over 11 forays, totalling approximately 154 km of rugged driving. This is an area increasingly explored in literature, and we therefore present and release examples of recent open-sourced radar place recognition systems and their performance on our dataset. This includes a learned neural network, the weights of which we also release. The data and tools are made freely available to the community at oxford-robotics-institute.github.io/oord-dataset Matthew Gadd, Daniele De Martini, Oliver Bartlett, Paul Murcutt, Matthew Towlson, Matthew Widojo, Valentina Musat, Luke Robinson, Efimia Panagiotaki, Georgi Pramatarov, Marc Alexander Kühn, Letizia Marchegiani, Paul Newman 0001, Lars Kunze |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Mitigating Distributional Shift in Semantic Segmentation via Uncertainty Estimation From Unlabeled DataabstractKnowing when a trained segmentation model is encountering data that is different to its training data is important. Understanding and mitigating the effects of this play an important part in their application from a performance and assurance perspective-this being a safety concern in applications such as autonomous vehicles (AVs). This work presents a segmentation network that can detect errors caused by challenging test domains without any additional annotation in a single forward pass. As annotation costs limit the diversity of labelled datasets, we use easy-to-obtain, uncurated and unlabelled data to learn to perform uncertainty estimation by selectively enforcing consistency over data augmentation. To this end, a novel segmentation benchmark based on the SAX Dataset is used, which includes labelledtestdata spanning three autonomous-driving domains, ranging in appearance from dense urban to off-road. The proposed method, named$\mathrm{\gamma }{-}\rm{SSL}$, consistently outperforms uncertainty estimation and Out-of-Distribution (OoD) techniques on this difficult benchmark-by up to 10.7% in area under the receiver operating characteristic (ROC) curve and 19.2% in area under the precision-recall (PR) curve in the most challenging of the three scenarios. David S. W. Williams, Daniele De Martini, Matthew Gadd, Paul Newman 0001 |
IEEE Trans. Robotics | 3 |
| 2023 | Visual DNA: Representing and Comparing Images Using Distributions of Neuron ActivationsabstractSelecting appropriate datasets is critical in modern computer vision. However, no general-purpose tools exist to evaluate the extent to which two datasets differ. For this, we propose representing images - and by extension datasets - using Distributions of Neuron Activations (DNAs). DNAsfit distributions, such as histograms or Gaussians, to activations of neurons in a pre-trained feature extractor through which we pass the imager s) to represent. This extractor is frozen for all datasets, and we rely on its generally expressive power in feature space. By comparing two DNAs, we can evaluate the extent to which two datasets differ with granular control over the comparison attributes of interest, providing the ability to customise the way distances are measured to suit the requirements of the task at hand. Furthermore, DNAs are compact, representing datasets of any size with less than 15 megabytes. We demonstrate the value of DNAs by evaluating their applicability on several tasks, including conditional dataset comparison, synthetic image evaluation, and transfer learning, and across diverse datasets, ranging from synthetic cat images to celebrity faces and urban driving scenes. Benjamin Ramtoula, Matthew Gadd, Paul Newman 0001, Daniele De Martini |
CVPR | 2 |
| 2023 | Visual Servoing on Wheels: Robust Robot Orientation Estimation in Remote Viewpoint ControlabstractThis work proposes a fast deployment pipeline for visually-servoed robots which does not assume anything about either the robot - e.g. sizes, colour or the presence of markers - or the deployment environment. Specifically, we apply a learning based approach to reliably estimate the pose of a robot in the image frame of a 2D camera upon which a visual servoing control system can be deployed. To alleviate the time-consuming process of labelling image data, we propose a weakly supervised pipeline that can produce a vast amount of data in a small amount of time. We evaluate our approach on a dataset of remote camera images captured in various indoor environments demonstrating high tracking performances when integrated into a fully-autonomous pipeline with a simple controller. With this, we then analyse the data requirement of our approach, showing how it is possible to deploy a new robot in a new environment in fewer than 30.00 min. Luke Robinson, Daniele De Martini, Matthew Gadd, Paul Newman 0001 |
IROS | 3 |
| 2023 | Off the Radar: Uncertainty-Aware Radar Place Recognition with Introspective Querying and Map MaintenanceabstractLocalisation with Frequency-Modulated Continuous-Wave (FMCW) radar has gained increasing interest due to its inherent resistance to challenging environments. However, complex artefacts of the radar measurement process require appropriate uncertainty estimation - to ensure the safe and reliable application of this promising sensor modality. In this work, we propose a multi-session map management system which constructs the “best” maps for further localisation based on learned variance properties in an embedding space. Using the same variance properties, we also propose a new way to introspectively reject localisation queries that are likely to be incorrect. For this, we apply robust noise-aware metric learning, which both leverages the short-timescale variability of radar data along a driven path (for data augmentation) and predicts the downstream uncertainty in metric-space-based place recognition. We prove the effectiveness of our method over extensive cross-validated tests of the Oxford Radar RobotCar and MulRan dataset. In this, we outperform the current state-of-the-art in radar place recognition and other uncertainty-aware methods when using only single nearest-neighbour queries. We also show consistent performance increases when rejecting queries based on uncertainty over a difficult test environment, which we did not observe for a competing uncertainty-aware place recognition system. Jianhao Yuan, Paul Newman 0001, Matthew Gadd |
IROS | 3 |
| 2022 | Depth-SIMS: Semi-Parametric Image and Depth SynthesisabstractIn this paper we present a compositing image synthesis method that generates RGB canvases with well aligned segmentation maps and sparse depth maps, coupled with an in-painting network that transforms the RGB canvases into high quality RGB images and the sparse depth maps into pixel-wise dense depth maps. We benchmark our method in terms of structural alignment and image quality, showing an increase in mIoU over SOTA by 3.7 percentage points and a highly competitive FID. Furthermore, we analyse the quality of the generated data as training data for semantic segmentation and depth completion, and show that our approach is more suited for this purpose than other methods. Valentina Musat, Daniele De Martini, Matthew Gadd, Paul Newman 0001 |
ICRA | 3 |
| 2022 | Fast-MbyM: Leveraging Translational Invariance of the Fourier Transform for Efficient and Accurate Radar OdometryabstractMasking by Moving (MByM), provides robust and accurate radar odometry measurements through an exhaustive correlative search across discretised pose candidates. However, this dense search creates a significant computational bottleneck which hinders real-time performance when high-end GPUs are not available. Utilising the translational invariance of the Fourier Transform, in our approach, Fast Masking by Moving (f-MByM), we decouple the search for angle and translation. By maintaining end-to-end differentiability a neural network is used to mask scans and trained by supervising pose prediction directly. Training faster and with less memory, utilising a decoupled search allows f-MbyM to achieve significant run-time performance improvements on a CPU (168 %) and to run in real-time on embedded devices, in stark contrast to MbyM. Throughout, our approach remains accurate and competitive with the best radar odometry variants available in the literature – achieving an end-point drift of 2.01 % in translation and 6.3 deg /km on the Oxford Radar RobotCar Dataset. Rob Weston, Matthew Gadd, Daniele De Martini, Paul Newman 0001, Ingmar Posner |
ICRA | 2 |
| 2022 | BoxGraph: Semantic Place Recognition and Pose Estimation from 3D LiDARabstractThis paper is about extremely robust and lightweight localisation using LiDAR point clouds based on instance segmentation and graph matching. We model 3D point clouds as fully-connected graphs of semantically identified components where each vertex corresponds to an object instance and encodes its shape. Optimal vertex association across graphs allows for full 6-Degree-of-Freedom (DoF) pose estimation and place recognition by measuring similarity. This representation is very concise, condensing the size of maps by a factor of 25 against the state-of-the-art, requiring only 3 kB to represent a 1.4 MB laser scan. We verify the efficacy of our system on the SemanticKITTI dataset, where we achieve a new state-of-the-art in place recognition, with an average of 88.4 % recall at 100 % precision where the next closest competitor follows with 64.9 %. We also show accurate metric pose estimation performance - estimating 6-DoF pose with median errors of 10cm and 0.33 deg. Georgi Pramatarov, Daniele De Martini, Matthew Gadd, Paul Newman 0001 |
IROS | 3 |
| 2021 | Fool Me Once: Robust Selective Segmentation via Out-of-Distribution Detection with Contrastive LearningabstractIn this work, a neural network is trained to simultaneously perform segmentation and pixel-wise Out-of-Distribution (OoD) detection, such that the segmentation of unknown regions of scenes can be rejected. This is made possible by leveraging an OoD dataset with a novel contrastive objective and data augmentation scheme. By including unknown classes in the training data, a more robust feature representation is learned with known classes represented distinctly from those unknown. In comparison, when presented with unknown classes or conditions, many current approaches for segmentation frequently exhibit high confidence in their inaccurate segmentations and cannot be trusted in many operational environments. We validate our system on a real-world dataset of unusual driving scenes, and show that by selectively segmenting scenes based on what is predicted as OoD, we can increase the segmentation accuracy by an IoU of 0.2 with respect to alternative techniques. David S. W. Williams, Matthew Gadd, Daniele De Martini, Paul Newman 0001 |
ICRA | 2 |
| 2021 | Look Here: Learning Geometrically Consistent Refinement of Inverse-Depth Images for 3D ReconstructionabstractBuilding good 3D maps is a challenging and expensive task, which requires high-quality sensors and careful, time-consuming scanning. We seek to reduce the cost of building good reconstructions by correcting views of existing low-quality ones in a post-hoc fashion using learnt priors over surfaces and appearance. We train a convolutional neural network model to predict the difference in inverse-depth from varying viewpoints of two meshes — one of low-quality that we wish to correct, and one of high-quality that we use as a reference. Our full model runs at 11.3[Formula: see text]Hz when aggregating four input views. In contrast to previous work, we pay attention to the problem of excessive smoothing in corrected meshes. We address this with a suitable network architecture, and introduce a loss-weighting mechanism that emphasizes edges in the prediction. Furthermore, smooth predictions result in geometrical inconsistencies. To deal with this issue, we present a loss function which penalizes re-projection differences that are not due to occlusions. Future applications of this work will incorporate semantic scene understanding in a multi-task learning setting. We explore the efficacy of the proposed system in terms of gross error correction and generalization capability by showing its performance in practice on a subset of the Kitti Odometry dataset, complete with a component-wise ablation study. We evaluate correctness and completeness measures of surface reconstruction across viewpoints and show that the proposed system is introspective in regions lacking sufficient high-quality supervision — indeed, models trained with geometric consistency loss create a lot more surface in areas that were not supervised, in one case filling in 67.97% or 8010[Formula: see text]m2 of an unlabeled input region. Finally, we assess the practical applicability of our method at large-scale by experiments over the full scope of the Kitti Odometry dataset. Broadly, as a measure of effectiveness, our model reduces gross errors by 45.3–77.5%, up to five times more than previous work. We also assess the practical applicability of our method to 3D reconstruction at large scales and find that compared to the baseline our model shows better stability in correctness when improving completeness of surfaces, and is effective in reducing median total error by up to 21.8[Formula: see text]cm. Stefan Saftescu, Matthew Gadd, Paul Newman 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2020 | The Oxford Radar RobotCar Dataset: A Radar Extension to the Oxford RobotCar DatasetabstractIn this paper we present The Oxford Radar RobotCar Dataset, a new dataset for researching scene understanding using Millimetre-Wave FMCW scanning radar data. The target application is autonomous vehicles where this modality is robust to environmental conditions such as fog, rain, snow, or lens flare, which typically challenge other sensor modalities such as vision and LIDAR.(/P)(P)The data were gathered in January 2019 over thirty-two traversals of a central Oxford route spanning a total of 280 km of urban driving. It encompasses a variety of weather, traffic, and lighting conditions. This 4.7 TB dataset consists of over 240,000 scans from a Navtech CTS350-X radar and 2.4 million scans from two Velodyne HDL-32E 3D LIDARs; along with six cameras, two 2D LIDARs, and a GPS/INS receiver. In addition we release ground truth optimised radar odometry to provide an additional impetus to research in this domain. The full dataset is available for download at: ori.ox.ac.uk/datasets/radar-robotear-dataset. Dan Barnes, Matthew Gadd, Paul Murcutt, Paul Newman 0001, Ingmar Posner |
ICRA | 2 |
| 2020 | Kidnapped Radar: Topological Radar Localisation using Rotationally-Invariant Metric LearningabstractThis paper presents a system for robust, large-scale topological localisation using Frequency-Modulated Continuous-Wave scanning radar which extends the state-of-the-art by an efficient, learning-based approach to handle radar data for localisation. We learn a metric space for embedding polar radar scans using CNN and NetVLAD architectures traditionally applied to the visual domain. However, we tailor the feature extraction for more suitability to the polar nature of radar scan formation using cylindrical convolutions, anti-aliasing blurring, and azimuth-wise max-pooling; all in order to bolster the rotational invariance. The enforced metric space is then used to encode a reference trajectory, serving as a map, which is queried for nearest neighbour for recognition of places at run-time. We demonstrate the performance of our topological localisation system over the course of many repeat forays using the largest radar-focused mobile autonomy dataset released to date, totalling 280 km of urban driving, a small portion of which we also use to learn the weights of the modified architecture. As this work represents a novel application for radar, we analyse the utility of the proposed method via a comprehensive set of metrics which provide insight into the efficacy when used in a realistic system, showing improved performance over the root architecture even in the face of random rotational perturbation. Stefan Saftescu, Matthew Gadd, Daniele De Martini, Dan Barnes, Paul Newman 0001 |
ICRA | 2 |
| 2020 | Sense-Assess-eXplain (SAX): Building Trust in Autonomous Vehicles in Challenging Real-World Driving ScenariosabstractThis paper discusses ongoing work in demonstrating research in mobile autonomy in challenging driving scenarios. In our approach, we address fundamental technical issues to overcome critical barriers to assurance and regulation for large-scale deployments of autonomous systems. To this end, we present how we build robots that (1) can robustly sense and interpret their environment using traditional as well as unconventional sensors; (2) can assess their own capabilities; and (3), vitally in the purpose of assurance and trust, can provide causal explanations of their interpretations and assessments. As it is essential that robots are safe and trusted, we design, develop, and demonstrate fundamental technologies in real-world applications to overcome critical barriers which impede the current deployment of robots in economically and socially important areas. Finally, we describe ongoing work in the collection of an unusual, rare, and highly valuable dataset. Matthew Gadd, Daniele De Martini, Letizia Marchegiani, Paul Newman 0001, Lars Kunze |
IV | 1 |
| 2020 | RSS-Net: Weakly-Supervised Multi-Class Semantic Segmentation with FMCW RadarabstractThis paper presents an efficient annotation procedure and an application thereof to end-to-end, rich semantic segmentation of the sensed environment using Frequency-Modulated Continuous-Wave scanning radar. We advocate radar over the traditional sensors used for this task as it operates at longer ranges and is substantially more robust to adverse weather and illumination conditions. We avoid laborious manual labelling by exploiting the largest radar-focused urban autonomy dataset collected to date, correlating radar scans with RGB cameras and LiDAR sensors, for which semantic segmentation is an already consolidated procedure. The training procedure leverages a state-of-the-art natural image segmentation system which is publicly available and as such, in contrast to previous approaches, allows for the production of copious labels for the radar stream by incorporating four camera and two LiDAR streams. Additionally, the losses are computed taking into account labels to the radar sensor horizon by accumulating LiDAR returns along a pose-chain ahead and behind of the current vehicle position. Finally, we present the network with multi-channel radar scan inputs in order to deal with ephemeral and dynamic scene objects. Prannay Kaul, Daniele De Martini, Matthew Gadd, Paul Newman 0001 |
IV | 3 |
| 2019 | Fast Radar Motion Estimation with a Learnt Focus of Attention using Weak SupervisionabstractThis paper is about fast motion estimation with scanning radar. We use weak supervision to train a focus of attention policy which actively down-samples the measurement stream before data association steps are undertaken. At training, we avoid laborious manual labelling by exploiting short-term sensor coherence from multiple poses in the presence of an external ego-motion estimator (for example, wheel odometry). In this way, we generate copious annotated measurements which can be used for training a learning algorithm in a weakly-supervised fashion. We demonstrate the validity of the approach in the context of a Radar Odometry (RO) task, pre-filtering raw data with a popular image segmentation network trained as presented. We evaluate our system against 26 km of data collected in Central Oxford and show consistent motion estimation with greatly reduced radar processing times (by a factor of 2.36). Roberto Aldera, Daniele De Martini, Matthew Gadd, Paul Newman 0001 |
ICRA | 3 |
| 2016 | Checkout my map: Version control for fleetwide visual localisationabstractThis paper is about underpinning long-term operations of fleets of vehicles using visual localisation. In particular it examines ways in which vehicles, considered as independent agents, can share, update and leverage each others' visual experiences in a mutually beneficial way. We draw on our previous work in Experience-based Navigation (EBN) [1], in which a visual map supporting multiple representations of the same place is built, yielding real-time localisation capability for a solitary vehicle. We now consider how any number of such agents might operate in concert via data sharing policies that are germane to the shared task of lifelong localisation. We rapidly construct considerable maps by the conjoining of work distributed to asynchronous processes, and share expertise amongst the team by the selective dispensing of mission-specific map contents. We demonstrate and evaluate our system against 100km of data collected in North Oxford over a period of a month featuring diverse deviation in appearance due to atmospheric, lighting, and structural dynamics. We show that our framework is capable of creating maps in a fraction of the time required by single-agent EBN, with no significant loss in localisation robustness, and is able to furnish robots on real-world forays with maps which require much less storage. Matthew Gadd, Paul Newman 0001 |
IROS | 1 |
| 2015 | A framework for infrastructure-free warehouse navigationabstractThis paper presents a universally applicable graph-based framework for the navigation of warehouse robots equipped with only monocular cameras. We strongly advocate the use of relative pose information stored in a topological map, rather than a globally consistent metric representation of the environment. We show how multiple traversals of adjacent workspaces can be naturally “stitched” together in the course of a typical warehouse picking and shelving schedule to create a network of reusable paths in which the robot can efficiently localise and plan new routes. This allows us to command the robot to return to any of the previously visited locations not necessarily through the same route that we taught it. Unlike state-of-the-art teach and repeat systems using stereo vision, our approach exploits the strongly planar nature of the data obtained from a downward-facing camera, and creates odometric constraints by tracking the perceived texture of the floor and computing a simple homography. To demonstrate the robustness of our system, we validate our approach on datasets collected over a week-long period within a challenging and representative environment in the form of a warehouse shelving area. Matthew Gadd, Paul Newman 0001 |
ICRA | 1 |