VLDB 2026 Research / reviewers in the wild / expert
Ingmar Posner
dblp:59/542
· DBLP profile ↗
71ranked-venue papers
2as first author
24since 2021 · last 2026
0000-0001-6270-700XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 68 · 2 first-author · 22 since 2021Systems, architecture and hardware · 38 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TactGen: Tactile Sensory Data Generation via Zero-Shot Sim-to-Real Transfer (Abstract Reprint)abstractRecent advances in machine learning have driven a step-change in robot perception with modalities such as vision, where large amounts of training data are readily available or cheap to collect. However, in tactile perception, the relatively high cost of data collection still largely impedes the adoption of such data-driven learning solutions. In this article, we introduce TactGen, a novel, cross-modal framework to tackle this challenge. In particular, using a two-step data generation pipeline, we leverage easily accessible vision data to synthesise artificial tactile data for downstream classifier training. Specifically, we use readily collected video data of objects of interest to efficiently learn neural radiance field (NeRF) representations. The NeRF models are then used to render red–green–blue-depth (RGBD) images from any desired vantage points. In the second stage, the RGBD images are translated into corresponding tactile images typically generated by camera-based tactile sensors using a conditional generative adversarial network (cGAN). The cGAN model is itself trained with a large set of visuo-tactile images collected in simulation, and can be transferred into the real world without fine-tuning or additional data collection. We extensively validate this approach in the context of tactile object classification, showing that it effectively reduces data collection time by a factor of 20 while achieving similar performance to training a classifier on manually collected real data. Shaohong Zhong, Alessandro Albini, Perla Maiolino, Ingmar Posner |
AAAI | 4 |
| 2025 | Group Decision-Making in Robot Teleoperation: Two Heads are Better Than OneabstractOperators working with robots in safety-critical domains have to make decisions under uncertainty, which remains a challenging problem for a single human operator. An open question is whether two human operators can make better decisions jointly, as compared to a single operator alone. While prior work has shown that two heads are better than one, such studies have been mostly limited to static and passive tasks. We investigate joint decision-making in a dynamic task involving humans teleoperating robots. We conduct a human-subject experiment with N = 100 participants where each participant performed a navigation task with two mobiles robots in simulation. We find that joint decision-making through confidence sharing improves dyad performance beyond the better-performing individual (p <0.0001). Further, we find that the extent of this benefit is regulated both by the skill level of each individual, as well as how well-calibrated their confidence estimates are. Finally, we present findings on characterising the human-human dyad's confidence calibration based on the individuals constituting the dyad. Our findings demonstrate for the first time that two heads are better than one, even on a spatiotemporal task which includes active operator control of robots. Duc-An Nguyen, Raunak P. Bhattacharyya, Clara Colombatto, Steve Fleming, Ingmar Posner, Nick Hawes |
HRI | 5 |
| 2025 | Kaputt: A Large-Scale Dataset for Visual Defect DetectionabstractWe present a novel large-scale dataset for defect detection in a logistics setting. Recent work on industrial anomaly detection has primarily focused on manufacturing scenarios with highly controlled poses and a limited number of object categories. Existing benchmarks like MVTec-AD [6] and VisA [33] have reached saturation, with state-of-the-art methods achieving up to 99.9% AUROC scores. In contrast to manufacturing, anomaly detection in retail logistics faces new challenges, particularly in the diversity and variability of object pose and appearance. Leading anomaly detection methods fall short when applied to this new setting. To bridge this gap, we introduce a new benchmark that overcomes the current limitations of existing datasets. With over 230,000 images (and more than 29,000 defective instances), it is 40 times larger than MVTec-AD and contains more than 48,000 distinct objects. To validate the difficulty of the problem, we conduct an extensive evaluation of multiple state-of-the-art anomaly detection methods, demonstrating that they do not surpass 56.96% AUROC on our dataset. Further qualitative analysis confirms that existing methods struggle to leverage normal samples under heavy pose and appearance variation. With our large-scale dataset, we set a new benchmark and encourage future research towards solving this challenging problem in retail logistics anomaly detection. The dataset is available for download under https://www.kaputt-dataset.com. Sebastian Höfer, Dorian Henning, Artemij Amiranashvili, Douglas Morrison, Mariliza Tzes, Ingmar Posner, Marc Matvienko, Alessandro Rennola, Anton Milan |
ICCV | 6 |
| 2025 | LUMOS: Language-Conditioned Imitation Learning with World ModelsabstractWe introduce LUMOS, a language-conditioned multi-task imitation learning framework for robotics. LUMOS learns skills by practicing them over many long-horizon rollouts in the latent space of a learned world model and transfers these skills zero-shot to a real robot. By learning on-policy in the latent space of the learned world model, our algorithm mitigates policy-induced distribution shift which most offline imitation learning methods suffer from. LUMOS learns from unstructured play data with fewer than 1 % hindsight language annotations but is steerable with language commands at test time. We achieve this coherent long-horizon performance by combining latent planning with both image-and language-based hindsight goal relabeling during training, and by optimizing an intrinsic reward defined in the latent space of the world model over multiple time steps, effectively reducing covariate shift. In experiments on the difficult long-horizon CALVIN benchmark, LUMOS outperforms prior learning-based methods with com-parable approaches on chained multi-task evaluations. To the best of our knowledge, we are the first to learn a language-conditioned continuous visuomotor control for a real-world robot within an offline world model. Videos, dataset and code are available at http://lumos.cs.uni-freiburg.de. Iman Nematollahi, Branton DeMoss, Akshay L. Chandra, Nick Hawes, Wolfram Burgard, Ingmar Posner |
ICRA | 6 |
| 2025 | Offline Adaptation of Quadruped Locomotion Using Diffusion ModelsabstractWe present a diffusion-based approach to quadrupedal locomotion that simultaneously addresses the limitations of learning and interpolating between multiple skills (modes) and of offline adapting to new locomotion behaviours after training. This is the first framework to apply classifier-free guided diffusion to quadruped locomotion and demonstrate its efficacy by extracting goal-conditioned behaviour from an originally unlabelled dataset. We show that these capabilities are compatible with a multi-skill policy and can be applied with little modification and minimal compute overhead, i.e., running entirely on the robot's onboard CPU. We verify the validity of our approach with hardware experiments on the ANYmal quadruped platform. Reece O'Mahoney, Alexander L. Mitchell, Wanming Yu, Ingmar Posner, Ioannis Havoutis |
ICRA | 4 |
| 2025 | SPARTAN: A Sparse Transformer World Model Attending to What MattersabstractCapturing the interactions between entities in a structured way plays a central role in world models that flexibly adapt to changes in the environment. Recent works motivate the benefits of models that explicitly represent the structure of interactions and formulate the problem as discovering local causal structures. In this work, we demonstrate that reliably capturing these relationships in complex settings remains challenging. To remedy this shortcoming, we postulate that sparsity is a critical ingredient for the discovery of such local structures. To this end we present the SPARse TrANsformer World model (SPARTAN), a Transformer-based world model that learns context-dependent interaction structures between entities in a scene. By applying sparsity regularisation on the attention patterns between object-factored tokens, SPARTAN learns sparse, context-dependent interaction graphs that accurately predict future object states. We further extend our model to adapt to sparse interventions with unknown targets on the dynamics of the environment. This results in a highly interpretable world model that can efficiently adapt to changes. Empirically, we evaluate SPARTAN against the current state-of-the-art in object-centric world models on observation-based environments and demonstrate that our model can learn local causal graphs that accurately reflects the underlying interactions between objects and achieve significantly improved few-shot adaptation to dynamics changes as well as robustness against distractors. Anson Lei, Bernhard Schölkopf, Ingmar Posner |
NeurIPS | 3 |
| 2025 | TactGen: Tactile Sensory Data Generation via Zero-Shot Sim-to-Real TransferabstractRecent advances in machine learning have driven a step-change in robot perception with modalities such as vision, where large amounts of training data are readily available or cheap to collect. However, in tactile perception, the relatively high cost of data collection still largely impedes the adoption of such data-driven learning solutions. In this article, we introduce TactGen, a novel, cross-modal framework to tackle this challenge. In particular, using a two-step data generation pipeline, we leverage easily accessible vision data to synthesise artificial tactile data for downstream classifier training. Specifically, we use readily collected video data of objects of interest to efficiently learn neural radiance field (NeRF) representations. The NeRF models are then used to render red–green–blue-depth (RGBD) images from any desired vantage points. In the second stage, the RGBD images are translated into corresponding tactile images typically generated by camera-based tactile sensors using a conditional generative adversarial network (cGAN). The cGAN model is itself trained with a large set of visuo-tactile images collected in simulation, and can be transferred into the real world without fine-tuning or additional data collection. We extensively validate this approach in the context of tactile object classification, showing that it effectively reduces data collection time by a factor of 20 while achieving similar performance to training a classifier on manually collected real data. Shaohong Zhong, Alessandro Albini, Perla Maiolino, Ingmar Posner |
IEEE Trans. Robotics | 4 |
| 2024 | Reward-Free Curricula for Training Robust World ModelsabstractThere has been a recent surge of interest in developing generally-capable agents that can adapt to new tasks without additional training in the environment. Learning world models from reward-free exploration is a promising approach, and enables policies to be trained using imagined experience for new tasks. However, achieving a general agent requires robustness across different environments. In this work, we address the novel problem of generating curricula in the reward-free setting to train robust world models. We consider robustness in terms of minimax regret over all environment instantiations and show that the minimax regret can be connected to minimising the maximum error in the world model across environment instances. This result informs our algorithm, WAKER: Weighted Acquisition of Knowledge across Environments for Robustness. WAKER selects environments for data collection based on the estimated error of the world model for each environment. Our experiments demonstrate that WAKER outperforms naı̈ve domain randomisation, resulting in improved robustness, efficiency, and generalisation. Marc Rigter, Minqi Jiang, Ingmar Posner |
ICLR | 3 |
| 2024 | TWIST: Teacher-Student World Model Distillation for Efficient Sim-to-Real TransferabstractModel-based RL is a promising approach for real-world robotics due to its improved sample efficiency and generalization capabilities compared to model-free RL. However, effective model-based RL solutions for vision-based real-world applications require bridging the sim-to-real gap for any world model learnt. Due to its significant computational cost, standard domain randomisation does not provide an effective solution to this problem. This paper proposes TWIST (Teacher-Student World Model Distillation for Sim-to-Real Transfer) to achieve efficient sim-to-real transfer of vision-based model-based RL using distillation. Specifically, TWIST leverages state observations as readily accessible, privileged information commonly garnered from a simulator to significantly accelerate sim-to-real transfer. Specifically, a teacher world model is trained efficiently on state information. At the same time, a matching dataset is collected of domain-randomised image observations. The teacher world model then supervises a student world model that takes the domain-randomised image observations as input. By distilling the learned latent dynamics model from the teacher to the student model, TWIST achieves efficient and effective sim-to-real transfer for vision-based model-based RL tasks. Experiments in simulated and real robotics tasks demonstrate that our approach outperforms naive domain randomisation and model-free methods in terms of sample efficiency and task performance of sim-to-real transfer. Jun Yamada, Marc Rigter, Jack Collins, Ingmar Posner |
ICRA | 4 |
| 2023 | Priors, Hierarchy, and Information Asymmetry for Skill Transfer in Reinforcement Learning
Sasha Salter, Kristian Hartikainen, Walter Goodwin, Ingmar Posner |
ICLR | 4 |
| 2023 | Leveraging Scene Embeddings for Gradient-Based Motion Planning in Latent SpaceabstractMotion planning framed as optimisation in structured latent spaces has recently emerged as competitive with traditional methods in terms of planning success while significantly outperforming them in terms of computational speed. However, the real-world applicability of recent work in this domain remains limited by the need to express obstacle information directly in state-space, involving simple geometric primitives. In this work we address this challenge by leveraging learned scene embeddings together with a generative model of the robot manipulator to drive the optimisation process. In addition, we introduce an approach for efficient collision checking which directly regularises the optimisation undertaken for planning. Using simulated as well as real-world experiments, we demonstrate that our approach, AMP-LS, is able to successfully plan in novel, complex scenes while outperforming traditional planning baselines in terms of computation speed by an order of magnitude. We show that the resulting system is fast enough to enable closed-loop planning in real-world dynamic scenes. Jun Yamada, Chia-Man Hung, Jack Collins, Ioannis Havoutis, Ingmar Posner |
ICRA | 5 |
| 2023 | Neural Latent Geometry Search: Product Manifold Inference via Gromov-Hausdorff-Informed Bayesian OptimizationabstractRecent research indicates that the performance of machine learning models can be improved by aligning the geometry of the latent space with the underlying data structure. Rather than relying solely on Euclidean space, researchers have proposed using hyperbolic and spherical spaces with constant curvature, or combinations thereof, to better model the latent space and enhance model performance. However, little attention has been given to the problem of automatically identifying the optimal latent geometry for the downstream task. We mathematically define this novel formulation and coin it as neural latent geometry search (NLGS). More specifically, we introduce an initial attempt to search for a latent geometry composed of a product of constant curvature model spaces with a small number of query evaluations, under some simplifying assumptions. To accomplish this, we propose a novel notion of distance between candidate latent geometries based on the Gromov-Hausdorff distance from metric geometry. In order to compute the Gromov-Hausdorff distance, we introduce a mapping function that enables the comparison of different manifolds by embedding them in a common high-dimensional ambient space. We then design a graph search space based on the notion of smoothness between latent geometries and employ the calculated distances as an additional inductive bias. Finally, we use Bayesian optimization to search for the optimal latent geometry in a query-efficient manner. This is a general method which can be applied to search for the optimal latent geometry for a variety of models and downstream tasks. We perform experiments on synthetic and real-world datasets to identify the optimal latent geometry for multiple machine learning problems. Haitz Sáez de Ocáriz Borde, Alvaro Arroyo, Ismael Morales, Ingmar Posner, Xiaowen Dong 0001 |
NeurIPS | 4 |
| 2023 | VAE-Loco: Versatile Quadruped Locomotion by Learning a Disentangled Gait RepresentationabstractQuadruped locomotion is rapidly maturing to a degree where robots are able to realize highly dynamic maneuvers. However, current planners are unable to vary key gait parameters of the in-swing feetmidair. In this article, we address this limitation and show that it is pivotal in increasing controller robustness by learning a latent space capturing the key stance phases constituting a particular gait. This is achieved via a generative model trained on a single trot style, which encourages disentanglement such that application of a drive signal to a single dimension of the latent state induces holistic plans synthesizing a continuous variety of trot styles. We demonstrate that specific properties of the drive signal map directly to gait parameters, such as cadence, footstep height, and full-stance duration. Due to the nature of our approach, these synthesized gaits are continuously variable online during robot operation. The use of a generative model facilitates the detection and mitigation of disturbances to provide a versatile and robust planning framework. We evaluate our approach on two versions of the real ANYmal quadruped robots and demonstrate that our method achieves a continuous blend of dynamic trot styles while being robust and reactive to external perturbations. Alexander L. Mitchell, Wolfgang Merkt, Mathieu Geisert, Siddhant Gangapurwala, Martin Engelcke, Oiwi Parker Jones, Ioannis Havoutis, Ingmar Posner |
IEEE Trans. Robotics | 8 |
| 2022 | Zero-Shot Category-Level Object Pose Estimation
Walter Goodwin, Sagar Vaze, Ioannis Havoutis, Ingmar Posner |
ECCV (39) | 4 |
| 2022 | Semantically Grounded Object Matching for Robust Robotic Scene RearrangementabstractObject rearrangement has recently emerged as a key competency in robot manipulation, with practical solutions generally involving object detection, recognition, grasping and high-level planning. Goal-images describing a desired scene configuration are a promising and increasingly used mode of instruction. A key outstanding challenge is the accurate inference of matches between objects in front of a robot, and those seen in a provided goal image, where recent works have struggled in the absence of object-specific training data. In this work, we explore the deterioration of existing methods' ability to infer matches between objects as the visual shift between observed and goal scenes increases. We find that a fundamental limitation of the current setting is that source and target images must contain the same instance of every object, which restricts practical deployment. We present a novel approach to object matching that uses a large pre-trained vision-language model to match objects in a cross-instance setting by leveraging semantics together with visual features as a more robust, and much more general, measure of similarity. We demonstrate that this provides considerably improved matching performance in cross-instance settings, and can be used to guide multi-object rearrangement with a robot manipulator from an image that shares no object instances with the robot's scene. Our code is available at https://github.com/applied-ai-lab/object_matching. Walter Goodwin, Sagar Vaze, Ioannis Havoutis, Ingmar Posner |
ICRA | 4 |
| 2022 | Next Steps: Learning a Disentangled Gait Representation for Versatile Quadruped LocomotionabstractQuadruped locomotion is rapidly maturing to a degree where robots now routinely traverse a variety of unstructured terrains. However, while gaits can be varied typically by selecting from a range of pre-computed styles, current planners are unable to vary key gait parameters continuously while the robot is in motion. The synthesis, on-the-fly, of gaits with unexpected operational characteristics or even the blending of dynamic manoeuvres lies beyond the capabilities of the current state-of-the-art. In this work we address this limitation by learning a latent space capturing the key stance phases of a particular gait, via a generative model trained on a single trot style. This encourages disentanglement such that application of a drive signal to a single dimension of the latent state induces holistic plans synthesising a continuous variety of trot styles. In fact properties of this drive signal map directly to gait parameters such as cadence, footstep height and full stance duration. The use of a generative model facilitates the detection and mitigation of disturbances to provide a versatile and robust planning framework. We evaluate our approach on a real ANYmal quadruped robot and demonstrate that our method achieves a continuous blend of dynamic trot styles whilst being robust and reactive to external perturbations. Alexander L. Mitchell, Wolfgang Merkt, Mathieu Geisert, Siddhant Gangapurwala, Martin Engelcke, Oiwi Parker Jones, Ioannis Havoutis, Ingmar Posner |
ICRA | 8 |
| 2022 | Fast-MbyM: Leveraging Translational Invariance of the Fourier Transform for Efficient and Accurate Radar OdometryabstractMasking by Moving (MByM), provides robust and accurate radar odometry measurements through an exhaustive correlative search across discretised pose candidates. However, this dense search creates a significant computational bottleneck which hinders real-time performance when high-end GPUs are not available. Utilising the translational invariance of the Fourier Transform, in our approach, Fast Masking by Moving (f-MByM), we decouple the search for angle and translation. By maintaining end-to-end differentiability a neural network is used to mask scans and trained by supervising pose prediction directly. Training faster and with less memory, utilising a decoupled search allows f-MbyM to achieve significant run-time performance improvements on a CPU (168 %) and to run in real-time on embedded devices, in stark contrast to MbyM. Throughout, our approach remains accurate and competitive with the best radar odometry variants available in the literature – achieving an end-point drift of 2.01 % in translation and 6.3 deg /km on the Oxford Radar RobotCar Dataset. Rob Weston, Matthew Gadd, Daniele De Martini, Paul Newman 0001, Ingmar Posner |
ICRA | 5 |
| 2022 | Universal Approximation of Functions on SetsabstractModelling functions of sets, or equivalently, permutation-invariant functions, is a long-standing challenge in machine learning. Deep Sets is a popular method which is known to be a universal approximator for continuous set functions. We provide a theoretical analysis of Deep Sets which shows that this universal approximation property is only guaranteed if the model's latent space is sufficiently high-dimensional. If the latent space is even one dimension lower than necessary, there exist piecewise-affine functions for which Deep Sets performs no better than a naïve constant baseline, as judged by worst-case error. Deep Sets may be viewed as the most efficient incarnation of the Janossy pooling paradigm. We identify this paradigm as encompassing most currently popular set-learning methods. Based on this connection, we discuss the implications of our results for set learning more broadly, and identify some open questions on the universality of Janossy pooling in general. Edward Wagstaff, Fabian Fuchs, Martin Engelcke, Michael A. Osborne, Ingmar Posner |
J. Mach. Learn. Res. | 5 |
| 2021 | Goal-Conditioned End-to-End Visuomotor Control for Versatile Skill PrimitivesabstractVisuomotor control (VMC) is an effective means of achieving basic manipulation tasks such as pushing or pick- and-place from raw images. Conditioning VMC on desired goal states is a promising way of achieving versatile skill primitives. However, common conditioning schemes either rely on task-specific fine tuning - e.g. using one-shot imitation learning (IL) - or on sampling approaches using a forward model of scene dynamics i.e. model-predictive control (MPC), leaving deployability and planning horizon severely limited. In this paper we propose a conditioning scheme which avoids these pitfalls by learning the controller and its conditioning in an end-to-end manner. Our model predicts complex action sequences based directly on a dynamic image representation of the robot motion and the distance to a given target observation. In contrast to related works, this enables our approach to efficiently perform complex manipulation tasks from raw image observations without predefined control primitives or test time demonstrations. We report significant improvements in task success over representative MPC and IL baselines. We also demonstrate our model's generalisation capabilities in challenging, unseen tasks featuring visual noise, cluttered scenes and unseen object geometries. Oliver Groth, Chia-Man Hung, Andrea Vedaldi, Ingmar Posner |
ICRA | 4 |
| 2021 | Introspective Visuomotor Control: Exploiting Uncertainty in Deep Visuomotor Control for Failure RecoveryabstractEnd-to-end visuomotor control is emerging as a compelling solution for robot manipulation tasks. However, imitation learning-based visuomotor control approaches tend to suffer from a common limitation, lacking the ability to recover from an out-of-distribution state caused by compounding errors. In this paper, instead of using tactile feedback or explicitly detecting the failure through vision, we investigate using the uncertainty of a policy neural network. We propose a novel uncertainty-based approach to detect and recover from failure cases. Our hypothesis is that policy uncertainties can implicitly indicate the potential failures in the visuomotor control task and that robot states with minimum uncertainty are more likely to lead to task success. To recover from high uncertainty cases, the robot monitors its uncertainty along a trajectory and explores possible actions in the state-action space to bring itself to a more certain state. Our experiments verify this hypothesis and show a significant improvement on task success rate: 12% in pushing, 15% in pick-and-reach and 22% in pick-and-place. Chia-Man Hung, Li Sun 0005, Yizhe Wu, Ioannis Havoutis, Ingmar Posner |
ICRA | 5 |
| 2021 | There and Back Again: Learning to Simulate Radar Data for Real-World ApplicationsabstractSimulating realistic radar data has the potential to significantly accelerate the development of data-driven approaches to radar processing. However, it is fraught with difficulty due to the notoriously complex image formation process. Here we propose to learn a radar sensor model capable of synthesising faithful radar observations based on simulated elevation maps. In particular, we adopt an adversarial approach to learning a forward sensor model from unaligned radar examples. In addition, modelling the backward model encourages the output to remain aligned to the world state through a cyclical consistency criterion. The backward model is further constrained to predict elevation maps from real radar data that are grounded by partial measurements obtained from corresponding lidar scans. Both models are trained in a joint optimisation. We demonstrate the efficacy of our approach by evaluating a down-stream segmentation model trained purely on simulated data in a real-world deployment. This achieves performance within four percentage points of the same model trained entirely on real data. Rob Weston, Oiwi Parker Jones, Ingmar Posner |
ICRA | 3 |
| 2021 | APEX: Unsupervised, Object-Centric Scene Segmentation and Tracking for Robot ManipulationabstractRecent advances in unsupervised learning for object detection, segmentation, and tracking hold significant promise for applications in robotics. A common approach is to frame these tasks as inference in probabilistic latent-variable models. In this paper, however, we show that the current state-of-the-art struggles with visually complex scenes such as typically encountered in robot manipulation tasks. We propose APEX, a new latent-variable model which is able to segment and track objects in more realistic scenes featuring objects that vary widely in size and texture, including the robot arm itself. This is achieved by a principled mask normalisation algorithm and a high-resolution scene encoder. To evaluate our approach, we present results on the real-world Sketchy dataset. This dataset, however, does not contain ground truth masks and object IDs for a quantitative evaluation. We thus introduce the Panda Pushing Dataset (P2D) which shows a Panda arm interacting with objects on a table in simulation and which includes ground-truth segmentation masks and object IDs for tracking. In both cases, APEX comprehensively outperforms the current state-of-the-art in unsupervised object segmentation and tracking. We demonstrate the efficacy of our segmentations for robot skill execution on an object arrangement task, where we also achieve the best or comparable performance among all the baselines. Yizhe Wu, Oiwi Parker Jones, Martin Engelcke, Ingmar Posner |
IROS | 4 |
| 2021 | GENESIS-V2: Inferring Unordered Object Representations without Iterative RefinementabstractAdvances in unsupervised learning of object-representations have culminated in the development of a broad range of methods for unsupervised object segmentation and interpretable object-centric scene generation. These methods, however, are limited to simulated and real-world datasets with limited visual complexity. Moreover, object representations are often inferred using RNNs which do not scale well to large images or iterative refinement which avoids imposing an unnatural ordering on objects in an image but requires the a priori initialisation of a fixed number of object representations. In contrast to established paradigms, this work proposes an embedding-based approach in which embeddings of pixels are clustered in a differentiable fashion using a stochastic stick-breaking process. Similar to iterative refinement, this clustering procedure also leads to randomly ordered object representations, but without the need of initialising a fixed number of clusters a priori. This is used to develop a new model, GENESIS-v2, which can infer a variable number of object representations without using RNNs or iterative refinement. We show that GENESIS-v2 performs strongly in comparison to recent baselines in terms of unsupervised image segmentation and object-centric scene generation on established synthetic datasets as well as more complex real-world datasets. Martin Engelcke, Oiwi Parker Jones, Ingmar Posner |
NeurIPS | 3 |
| 2021 | E(n) Equivariant Normalizing FlowsabstractThis paper introduces a generative model equivariant to Euclidean symmetries: E(n) Equivariant Normalizing Flows (E-NFs). To construct E-NFs, we take the discriminative E(n) graph neural networks and integrate them as a differential equation to obtain an invertible equivariant function: a continuous-time normalizing flow. We demonstrate that E-NFs considerably outperform baselines and existing methods from the literature on particle systems such as DW4 and LJ13, and on molecules from QM9 in terms of log-likelihood. To the best of our knowledge, this is the first flow that jointly generates molecule features and positions in 3D. Victor Garcia Satorras, Emiel Hoogeboom, Fabian Fuchs, Ingmar Posner, Max Welling |
NeurIPS | 4 |
| 2020 | GENESIS: Generative Scene Inference and Sampling with Object-Centric Latent Representations
Martin Engelcke, Adam R. Kosiorek, Oiwi Parker Jones, Ingmar Posner |
ICLR | 4 |
| 2020 | Localising Faster: Efficient and precise lidar-based robot localisation in large-scale environmentsabstractThis paper proposes a novel approach for global localisation of mobile robots in large-scale environments. Our method leverages learning-based localisation and filtering-based localisation, to localise the robot efficiently and precisely through seeding Monte Carlo Localisation (MCL) with a deeplearned distribution. In particular, a fast localisation system rapidly estimates the 6-DOF pose through a deep-probabilistic model (Gaussian Process Regression with a deep kernel), then a precise recursive estimator refines the estimated robot pose according to the geometric alignment. More importantly, the Gaussian method (i.e. deep probabilistic localisation) and nonGaussian method (i.e. MCL) can be integrated naturally via importance sampling. Consequently, the two systems can be integrated seamlessly and mutually benefit from each other. To verify the proposed framework, we provide a case study in large-scale localisation with a 3D lidar sensor. Our experiments on the Michigan NCLT long-term dataset show that the proposed method is able to localise the robot in 1.94 s on average (median of 0.8 s) with precision 0.75 m in a largescale environment of approximately 0.5 km2. Li Sun 0005, Daniel Adolfsson, Martin Magnusson 0002, Henrik Andreasson, Ingmar Posner, Tom Duckett |
ICRA | 5 |
| 2020 | The Oxford Radar RobotCar Dataset: A Radar Extension to the Oxford RobotCar DatasetabstractIn this paper we present The Oxford Radar RobotCar Dataset, a new dataset for researching scene understanding using Millimetre-Wave FMCW scanning radar data. The target application is autonomous vehicles where this modality is robust to environmental conditions such as fog, rain, snow, or lens flare, which typically challenge other sensor modalities such as vision and LIDAR.(/P)(P)The data were gathered in January 2019 over thirty-two traversals of a central Oxford route spanning a total of 280 km of urban driving. It encompasses a variety of weather, traffic, and lighting conditions. This 4.7 TB dataset consists of over 240,000 scans from a Navtech CTS350-X radar and 2.4 million scans from two Velodyne HDL-32E 3D LIDARs; along with six cameras, two 2D LIDARs, and a GPS/INS receiver. In addition we release ground truth optimised radar odometry to provide an additional impetus to research in this domain. The full dataset is available for download at: ori.ox.ac.uk/datasets/radar-robotear-dataset. Dan Barnes, Matthew Gadd, Paul Murcutt, Paul Newman 0001, Ingmar Posner |
ICRA | 5 |
| 2020 | Under the Radar: Learning to Predict Robust Keypoints for Odometry Estimation and Metric Localisation in RadarabstractThis paper presents a self-supervised framework for learning to detect robust keypoints for odometry estimation and metric localisation in radar. By embedding a differentiable point-based motion estimator inside our architecture, we learn keypoint locations, scores and descriptors from localisation error alone. This approach avoids imposing any assumption on what makes a robust keypoint and crucially allows them to be optimised for our application. Furthermore the architecture is sensor agnostic and can be applied to most modalities. We run experiments on 280km of real world driving from the Oxford Radar RobotCar Dataset and improve on the state-of-the-art in point-based radar odometry, reducing errors by up to 45% whilst running an order of magnitude faster, simultaneously solving metric loop closures. Combining these outputs, we provide a framework capable of full mapping and localisation with radar in urban environments. Dan Barnes, Ingmar Posner |
ICRA | 2 |
| 2020 | First Steps: Latent-Space Control with Semantic Constraints for Quadruped LocomotionabstractTraditional approaches to quadruped control frequently employ simplified, hand-derived models. This significantly reduces the capability of the robot since its effective kinematic range is curtailed. In addition, kinodynamic constraints are often non-differentiable and difficult to implement in an optimisation approach. In this work, these challenges are addressed by framing quadruped control as optimisation in a structured latent space. A deep generative model captures a statistical representation of feasible joint configurations, whilst complex dynamic and terminal constraints are expressed via high-level, semantic indicators and represented by learned classifiers operating upon the latent space. As a consequence, complex constraints are rendered differentiable and evaluated an order of magnitude faster than analytical approaches. We validate the feasibility of locomotion trajectories optimised using our approach both in simulation and on a real-world ANY-mal quadruped. Our results demonstrate that this approach is capable of generating smooth and realisable trajectories. To the best of our knowledge, this is the first time latent space control has been successfully applied to a complex, real robot platform. Alexander L. Mitchell, Martin Engelcke, Oiwi Parker Jones, David Surovik, Siddhant Gangapurwala, Oliwier Melon, Ioannis Havoutis, Ingmar Posner |
IROS | 8 |
| 2020 | RELATE: Physically Plausible Multi-Object Scene Synthesis Using Structured Latent SpacesabstractWe present RELATE, a model that learns to generate physically plausible scenes and videos of multiple interacting objects. Similar to other generative approaches, RELATE is trained end-to-end on raw, unlabeled data. RELATE combines an object-centric GAN formulation with a model that explicitly accounts for correlations between individual objects. This allows the model to generate realistic scenes and videos from a physically-interpretable parameterization. Furthermore, we show that modeling the object correlation is necessary to learn to disentangle object positions and identity. We find that RELATE is also amenable to physically realistic scene editing and that it significantly outperforms prior art in object-centric scene generation in both synthetic (CLEVR, ShapeStacks) and real-world data (cars). In addition, in contrast to state-of-the-art methods in object-centric generative modeling, RELATE also extends naturally to dynamic scenes and generates videos of high visual fidelity. Source code, datasets and more results are available at http://geometry.cs.ucl.ac.uk/projects/2020/relate/. Sébastien Ehrhardt, Oliver Groth, Áron Monszpart, Martin Engelcke, Ingmar Posner, Niloy J. Mitra, Andrea Vedaldi |
NeurIPS | 5 |
| 2019 | Scrutinizing and De-Biasing Intuitive Physics with Neural Stethoscopes
Fabian Fuchs, Oliver Groth, Adam R. Kosiorek, Alex Bewley, Markus Wulfmeier, Andrea Vedaldi, Ingmar Posner |
BMVC | 7 |
| 2019 | On the Limitations of Representing Functions on SetsabstractRecent work on the representation of functions on sets has considered the use of summation in a latent space to enforce permutation invariance. In particular, it has been conjectured that the dimension of this latent space may remain fixed as the cardinality of the sets under consideration increases. However, we demonstrate that the analysis leading to this conjecture requires mappings which are highly discontinuous and argue that this is only of limited practical use. Motivated by this observation, we prove that an implementation of this model via continuous mappings (as provided by e.g. neural networks or Gaussian processes) actually imposes a constraint on the dimensionality of the latent space. Practical universal function representation for set inputs can only be achieved with a latent dimension at least the size of the maximum number of input elements. Edward Wagstaff, Fabian Fuchs, Martin Engelcke, Ingmar Posner, Michael A. Osborne |
ICML | 4 |
| 2019 | Probably Unknown: Deep Inverse Sensor Modelling RadarabstractRadar presents a promising alternative to lidar and vision in autonomous vehicle applications, able to detect objects at long range under a variety of weather conditions. However, distinguishing between occupied and free space from raw radar power returns is challenging due to complex interactions between sensor noise and occlusion. To counter this we propose to learn an Inverse Sensor Model (ISM) converting a raw radar scan to a grid map of occupancy probabilities using a deep neural network. Our network is selfsupervised using partial occupancy labels generated by lidar, allowing a robot to learn about world occupancy from past experience without human supervision. We evaluate our approach on five hours of data recorded in a dynamic urban environment. By accounting for the scene context of each grid cell our model is able to successfully segment the world into occupied and free space, outperforming standard CFAR filtering approaches. Additionally by incorporating heteroscedastic uncertainty into our model formulation, we are able to quantify the variance in the uncertainty throughout the sensor observation. Through this mechanism we are able to successfully identify regions of space that are likely to be occluded. Rob Weston, Sarah H. Cen, Paul Newman 0001, Ingmar Posner |
ICRA | 4 |
| 2018 | ShapeStacks: Learning Vision-Based Physical Intuition for Generalised Object Stacking
Oliver Groth, Fabian Fuchs, Ingmar Posner, Andrea Vedaldi |
ECCV (1) | 3 |
| 2018 | TACO: Learning Task Decomposition via Temporal Alignment for ControlabstractMany advanced Learning from Demonstration (LfD) methods consider the decomposition of complex, real-world tasks into simpler sub-tasks. By reusing the corresponding sub-policies within and between tasks, we can provide training data for each policy from different high-level tasks and compose them to perform novel ones. Existing approaches to modular LfD focus either on learning a single high-level task or depend on domain knowledge and temporal segmentation. In contrast, we propose a weakly supervised, domain-agnostic approach based on task sketches, which include only the sequence of sub-tasks performed in each demonstration. Our approach simultaneously aligns the sketches with the observed demonstrations and learns the required sub-policies. This improves generalisation in comparison to separate optimisation procedures. We evaluate the approach on multiple domains, including a simulated 3D robot arm control task using purely image-based observations. The results show that our approach performs commensurately with fully supervised approaches, while requiring significantly less annotation effort. Kyriacos Shiarlis, Markus Wulfmeier, Sasha Salter, Shimon Whiteson, Ingmar Posner |
ICML | 5 |
| 2018 | Driven to Distraction: Self-Supervised Distractor Learning for Robust Monocular Visual Odometry in Urban EnvironmentsabstractWe present a self-supervised approach to ignoring “distractors” in camera images for the purposes of robustly estimating vehicle motion in cluttered urban environments. We leverage offline multi-session mapping approaches to automatically generate a per-pixel ephemerality mask and depth map for each input image, which we use to train a deep convolutional network. At run-time we use the predicted ephemerality and depth as an input to a monocular visual odometry (VO) pipeline, using either sparse features or dense photometric matching. Our approach yields metric-scale VO using only a single camera and can recover the correct egomotion even when 90% of the image is obscured by dynamic, independently moving objects. We evaluate our robust VO methods on more than 400km of driving from the Oxford RobotCar Dataset and demonstrate reduced odometry drift and significantly improved egomotion estimation in the presence of large moving vehicles in urban traffic. Dan Barnes, Will Maddern, Geoffrey Pascoe, Ingmar Posner |
ICRA | 4 |
| 2018 | Incremental Adversarial Domain Adaptation for Continually Changing EnvironmentsabstractContinuous appearance shifts such as changes in weather and lighting conditions can impact the performance of deployed machine learning models. While unsupervised domain adaptation aims to address this challenge, current approaches do not utilise the continuity of the occurring shifts. In particular, many robotics applications exhibit these conditions and thus facilitate the potential to incrementally adapt a learnt model over minor shifts which integrate to massive differences over time. Our work presents an adversarial approach for lifelong, incremental domain adaptation which benefits from unsupervised alignment to a series of intermediate domains which successively diverge from the labelled source domain. We empirically demonstrate that our incremental approach improves handling of large appearance changes, e.g. day to night, on a traversable-path segmentation task compared with a direct, single alignment step approach. Furthermore, by approximating the feature distribution for the source domain with a generative adversarial network, the deployment module can be rendered fully independent of retaining potentially large amounts of the related source training data for only a minor reduction in performance. Markus Wulfmeier, Alex Bewley, Ingmar Posner |
ICRA | 3 |
| 2018 | Sequential Attend, Infer, Repeat: Generative Modelling of Moving ObjectsabstractWe present Sequential Attend, Infer, Repeat (SQAIR), an interpretable deep generative model for image sequences. It can reliably discover and track objects through the sequence; it can also conditionally generate future frames, thereby simulating expected motion of objects. This is achieved by explicitly encoding object numbers, locations and appearances in the latent variables of the model. SQAIR retains all strengths of its predecessor, Attend, Infer, Repeat (AIR, Eslami et. al. 2016), including unsupervised learning, made possible by inductive biases present in the model structure. We use a moving multi-\textsc{mnist} dataset to show limitations of AIR in detecting overlapping or partially occluded objects, and show how \textsc{sqair} overcomes them by leveraging temporal consistency of objects. Finally, we also apply SQAIR to real-world pedestrian CCTV data, where it learns to reliably detect, track and generate walking pedestrians with no supervision. Adam R. Kosiorek, Hyunjik Kim, Yee Whye Teh, Ingmar Posner |
NeurIPS | 4 |
| 2017 | Find your own way: Weakly-supervised segmentation of path proposals for urban autonomyabstractWe present a weakly-supervised approach to segmenting proposed drivable paths in images with the goal of autonomous driving in complex urban environments. Using recorded routes from a data collection vehicle, our proposed method generates vast quantities of labelled images containing proposed paths and obstacles without requiring manual annotation, which we then use to train a deep semantic segmentation network. With the trained network we can segment proposed paths and obstacles at run-time using a vehicle equipped with only a monocular camera without relying on explicit modelling of road or lane markings. We evaluate our method on the large-scale KITTI and Oxford RobotCar datasets and demonstrate reliable path proposal and obstacle segmentation in a wide variety of environments under a range of lighting, weather and traffic conditions. We illustrate how the method can generalise to multiple path proposals at intersections and outline plans to incorporate the system into a framework for autonomous urban driving. Dan Barnes, Will Maddern, Ingmar Posner |
ICRA | 3 |
| 2017 | Vote3Deep: Fast object detection in 3D point clouds using efficient convolutional neural networksabstractThis paper proposes a computationally efficient approach to detecting objects natively in 3D point clouds using convolutional neural networks (CNNs). In particular, this is achieved by leveraging a feature-centric voting scheme to implement novel convolutional layers which explicitly exploit the sparsity encountered in the input. To this end, we examine the trade-off between accuracy and speed for different architectures and additionally propose to use an L1penalty on the filter activations to further encourage sparsity in the intermediate representations. To the best of our knowledge, this is the first work to propose sparse convolutional layers and L1regularisation for efficient large-scale processing of 3D data. We demonstrate the efficacy of our approach on the KITTI object detection benchmark and show that Vote3Deep models with as few as three layers outperform the previous state of the art in both laser and laser-vision based approaches by margins of up to 40% while remaining highly competitive in terms of processing time. Martin Engelcke, Dushyant Rao, Dominic Zeng Wang, Chi Hay Tong, Ingmar Posner |
ICRA | 5 |
| 2017 | Leveraging the urban soundscape: Auditory perception for smart vehiclesabstractUrban environments are characterised by the presence of distinctive audio signals which alert the drivers to events that require prompt action. The detection and interpretation of these signals would be highly beneficial for smart vehicle systems, as it would provide them with complementary information to navigate safely in the environment. In this paper, we present a framework that spots the presence of acoustic events, such as horns and sirens, using a two-stage approach. We first model the urban soundscape and use anomaly detection to identify the presence of an anomalous sound, and later determine the nature of this sound. As the audio samples are affected by copious non-stationary and unstructured noise, which can degrade classification performance, we propose a noise-removal technique to obtain a clean representation of the data we can use for classification and waveform reconstruction. The method is based on the idea of analysing the spectrograms of the incoming signals as images and applying spectrogram segmentation to isolate and extract the alerting signals from the background noise. We evaluate our framework on four hours of urban sounds collected driving around urban Oxford on different kinds of road and in different traffic conditions. When compared to traditional feature representations, such as Mel-frequency cepstrum coefficients, our framework shows an improvement of up to 31% in the classification rate. Letizia Marchegiani, Ingmar Posner |
ICRA | 2 |
| 2017 | What makes a place? Building bespoke place dependent object detectors for roboticsabstractThis paper is about enabling robots to improve their perceptual performance through repeated use in their operating environment, creating local expert detectors fitted to the places through which a robot moves. We leverage the concept of `experiences' in visual perception for robotics, accounting for bias in the data a robot sees by fitting object detector models to a particular `place'. The key question we seek to answer in this paper is simply: how do we define a place? We build bespoke pedestrian detector models for autonomous driving, highlighting the necessary trade off between generalisation and model capacity as we vary the extent of the `place' we fit to. We demonstrate a sizeable performance gain over a current state-of-the-art detector when using computationally lightweight bespoke place-fitted detector models. Jeffrey Hawke, Alex Bewley, Ingmar Posner |
IROS | 3 |
| 2017 | Addressing appearance change in outdoor robotics with adversarial domain adaptationabstractAppearance changes due to weather and seasonal conditions represent a strong impediment to the robust implementation of machine learning systems in outdoor robotics. While supervised learning optimises a model for the training domain, it will deliver degraded performance in application domains that underlie distributional shifts caused by these changes. Traditionally, this problem has been addressed via the collection of labelled data in multiple domains or by imposing priors on the type of shift between both domains. We frame the problem in the context of unsupervised domain adaptation and develop a framework for applying adversarial techniques to adapt popular, state-of-the-art network architectures with the additional objective to align features across domains. Moreover, as adversarial training is notoriously unstable, we first perform an extensive ablation study, adapting many techniques known to stabilise generative adversarial networks, and evaluate on a surrogate classification task with the same appearance change. The distilled insights are applied to the problem of free-space segmentation for motion planning in autonomous driving. Markus Wulfmeier, Alex Bewley, Ingmar Posner |
IROS | 3 |
| 2017 | Hierarchical Attentive Recurrent TrackingabstractClass-agnostic object tracking is particularly difficult in cluttered environments as target specific discriminative models cannot be learned a priori. Inspired by how the human visual cortex employs spatial attention and separate where'' andwhat'' processing pathways to actively suppress irrelevant visual features, this work develops a hierarchical attentive recurrent model for single object tracking in videos. The first layer of attention discards the majority of background by selecting a region containing the object of interest, while the subsequent layers tune in on visual features particular to the tracked object. This framework is fully differentiable and can be trained in a purely data driven fashion by gradient methods. To improve training convergence, we augment the loss function with terms for auxiliary tasks relevant for tracking. Evaluation of the proposed model is performed on two datasets: pedestrian tracking on the KTH activity recognition dataset and the more difficult KITTI object tracking dataset. Adam R. Kosiorek, Alex Bewley, Ingmar Posner |
NIPS | 3 |
| 2016 | Deep Tracking: Seeing Beyond Seeing Using Recurrent Neural NetworksabstractThis paper presents to the best of our knowledge the first end-to-end object tracking approach which directly maps from raw sensor input to object tracks in sensor space without requiring any feature engineering or system identification in the form of plant or sensor models. Specifically, our system accepts a stream of raw sensor data at one end and, in real-time, produces an estimate of the entire environment state at the output including even occluded objects. We achieve this by framing the problem as a deep learning task and exploit sequence models in the form of recurrent neural networks to learn a mapping from sensor measurements to object tracks. In particular, we propose a learning method based on a form of input dropout which allows learning in an unsupervised manner, only based on raw, occluded sensor data without access to ground-truth annotations. We demonstrate our approach using a synthetic dataset designed to mimic the task of tracking objects in 2D laser data — as commonly encountered in robotics applications — and show that it learns to track many dynamic objects despite occlusions and the presence of sensor noise. Peter Ondruska, Ingmar Posner |
AAAI | 2 |
| 2016 | Off the beaten track: Predicting localisation performance in visual teach and repeatabstractThis paper proposes an appearance-based approach to estimating localisation performance in the context of visual teach and repeat. Specifically, it aims to estimate the likely corridor around a taught trajectory within which a vision-based localisation system is still able to localise itself. In contrast to prior art, our system is able to predict this localisation envelope for trajectories in similar, yet geographically distant locations where no repeat runs have yet been performed. Thus, by characterising the localisation performance in one region, we are able to predict performance in another. To achieve this, we leverage a Gaussian Process regressor to estimate the likely number of feature matches for any keyframe in the teach run, based on a combination of trajectory properties such as curvature and an appearance model of the keyframe. Using data from real traversals, we demonstrate that our approach performs as well as prior art when it comes to interpolating localisation performance based on a number of repeat runs, while also performing well at generalising performance estimation to freshly taught trajectories. Julie Dequaire, Chi Hay Tong, Winston Churchill, Ingmar Posner |
ICRA | 4 |
| 2016 | Choosing a time and place for calibration of lidar-camera systemsabstractWe propose a calibration method that automatically estimates the extrinsic calibration between a sensor pose-graph from natural scenes. The sensor pose-graph represents a system of sensors comprising of lidars and cameras, without sensor co-visibility constraints. The method addresses the fact that each scene contributes differently to the calibration problem by introducing a diligent scene selection scheme. The algorithm searches over all scenes to extract a subset of exemplars, whose joint optimisation yields progressively better calibration estimates. This non-parametric method requires no knowledge of the physical world, and continuously finds scenes that better constrain the optimisation parameters. We explain the theory, implement the method, and provide detailed performance analyses with experiments on real-world data. Terry Scott 0002, Akshay A. Morye, Pedro Pinies, Lina María Paz, Ingmar Posner, Paul Newman 0001 |
ICRA | 5 |
| 2016 | Enabling intelligent energy management for robots using publicly available mapsabstractEnergy consumption represents one of the most basic constraints for mobile robot autonomy. We propose a new framework to predict energy consumption using information extracted from publicly available maps. This method avoids having to model internal robot configurations, which are often unavailable, while still providing invaluable predictions for both explored and unexplored trajectories. Our approach uses a heteroscedastic Gaussian Process to model the power consumption, which explicitly accounts for variance due to exogenous latent factors such as traffic and weather conditions. We evaluate our framework on 30km of data collected from a city centre environment with a mobile robot travelling on pedestrian walkways. Results across five different test routes show an average difference between predicted and measured power consumption of 3.3%, leading to an average error of 6.6% on predictions of energy consumption. The distinct advantage of our model is our ability to predict measurement variance. The variance predictions improved by 84.3% over a benchmark. Oliver Bartlett, Corina Gurau, Letizia Marchegiani, Ingmar Posner |
IROS | 4 |
| 2016 | Watch this: Scalable cost-function learning for path planning in urban environmentsabstractIn this work, we present an approach to learn cost maps for driving in complex urban environments from a large number of demonstrations of human driving behaviour. The learned cost maps are constructed directly from raw sensor measurements, bypassing the effort of manually designing cost maps as well as features. When deploying the cost maps, the trajectories generated not only replicate human-like driving behaviour but are also demonstrably robust against systematic errors in putative robot configuration. To achieve this we deploy a Maximum Entropy based, non-linear IRL framework which uses Fully Convolutional Neural Networks (FCNs) to represent the cost model underlying expert driving behaviour. Using a deep, parametric approach enables us to scale efficiently to large datasets and complex behaviours while being run-time independent of dataset extent during deployment. We demonstrate scalability and performance on an ambitious dataset collected over the course of one year including more than 25k demonstration trajectories extracted from over 120km of driving and 13 different drivers. We evaluate against a carefully designed cost map and, in addition, demonstrate robustness to systematic errors by learning precise cost-maps even in the presence of system calibration perturbations. Markus Wulfmeier, Dominic Zeng Wang, Ingmar Posner |
IROS | 3 |
| 2016 | Automated valet parking and charging for e-mobilityabstractAutomated valet parking services provide great potential to increase the attractiveness of electric vehicles by mitigating their two main current deficiencies: reduced driving ranges and prolonged refueling times. The European research project V-Charge aims at providing this service on designated parking lots using close-to-market sensors only. For this purpose the project developed a prototype capable of performing fully automated navigation in mixed traffic on designated parking lots and GPS-denied parking garages with cameras and ultrasonic sensors only. This paper summarizes the work of the project, comprising advances in network communication and parking space scheduling, multi-camera calibration, semantic mapping concepts, visual localization and motion planning. The project pushed visual localization, environment perception and automated parking to centimetre precision. The developed infrastructure-based camera calibration and semi-supervised semantic mapping concepts greatly reduce maintenance efforts. Results are presented from extensive month-long field tests. Ulrich Schwesinger, Mathias Bürki, Julian Timpner, Stephan Rottmann, Lars C. Wolf, Lina María Paz, Hugo Grimmett, Ingmar Posner, Paul Newman 0001, Christian Häne, Lionel Heng, Gim Hee Lee, Torsten Sattler, Marc Pollefeys, Marco Allodi, Francesco Valenti, Keiji Mimura, Bernd Goebelsmann, Wojciech Derendarz, Peter Mühlfellner, Stefan Wonneberger, Rene Waldmann, Sebastian Grysczyk, Carsten Last, Stefan Bruning, Sven Horstmann, Marc Bartholomaus, Clemens Brummer, Martin Stellmacher, Fabian Pucks, Marcel Nicklas, Roland Siegwart |
Intelligent Vehicles Symposium | 8 |
| 2015 | Learning to assess terrain from human demonstration using an introspective Gaussian-process classifierabstractThis paper presents an approach to learning robot terrain assessment from human demonstration. An operator drives a robot for a short period of time, supervising the gathering of traversable and untraversable terrain data. After this initial training period, the robot can then predict the traversability of new terrain based on its experiences. We improve on current methods in two ways: first, we maintain a richer (higher-dimensional) representation of the terrain that is better able to distinguish between different training examples. Second, we use a Gaussian-process classifier for terrain assessment due to its superior introspective abilities (leading to better uncertainty estimates) when compared to other classifier methods in the literature. Our method is tested on real data and shown to outperform current methods both in classification accuracy and uncertainty estimation. Laszlo-Peter Berczi, Ingmar Posner, Tim D. Barfoot |
ICRA | 2 |
| 2015 | Know your limits: Embedding localiser performance models in teach and repeat mapsabstractThis paper is about building maps which not only contain the traditional information useful for localising — such as point features — but also embeds a spatial model of expected localiser performance. This often overlooked second-order information provides vital context when it comes to map use and planning. Our motivation here is to improve the performance of the popular Teach and Repeat paradigm [1] which has been shown to enable truly large-scale field operation. When using the taught route for localisation, it is often assumed the robot is following exactly, or is sufficiently close to, the original path, enabling successful localisation. However, what happens if it is not possible, or not desirable to exactly follow the mapped path? How far off the beaten track can the robot travel before it gets lost? We present an approach for assessing this localisation area around a taught route, which we refer to as the localisation envelope. Using a combination of physical sampling and a Gaussian Process model, we are able to accurately predict the localisation performance at unseen points. Winston Churchill, Chi Hay Tong, Corina Gurau, Ingmar Posner, Paul Newman 0001 |
ICRA | 4 |
| 2015 | Integrating metric and semantic maps for vision-only automated parkingabstractWe present a framework for integrating two layers of map which are often required for fully automated operation: metric and semantic. Metric maps are likely to improve with subsequent visitations to the same place, while semantic maps can comprise both permanent and fluctuating features of the environment. However, it is not clear from the state of the art how to update the semantic layer as the metric map evolves. The strengths of our method are threefold: the framework allows for the unsupervised evolution of both maps as the environment is revisited by the robot; it uses vision-only sensors, making it appropriate for production cars; and the human labelling effort is minimised as far as possible while maintaining high fidelity. We evaluate this on two different car parks with a fully automated car, performing repeated automated parking manoeuvres to demonstrate the robustness of the system. Hugo Grimmett, Mathias Bürki, Lina María Paz, Pedro Pinies, Paul Timothy Furgale, Ingmar Posner, Paul Newman 0001 |
ICRA | 6 |
| 2015 | From dusk till dawn: Localisation at night using artificial light sourcesabstractThis paper is about localising at night in urban environments using vision. Despite it being dark exactly half of the time, surprisingly little attention has been given to this problem. A defining aspect of night-time urban scenes is the presence and effect of artificial lighting - be that in the form of street or interior lighting through windows. By building a model of the environment which includes a representation of the spatial location of every light source, localisation becomes possible using monocular cameras. One of the challenges we face is the gross change in light appearance as a function of distance due to flare, saturation and bleeding - city lights certainly do not appear as point features. To overcome this, we model the appearance of each light as a function of vehicle location, using this to inform our data-association decisions and to regularise the cost function which is used to infer vehicle pose. In this way we develop a place-dependent but stable sensor model which is customised for the particular environment in which we are operating. We demonstrate that our system is able to localise successfully at night over 12 km in situations where a traditional point feature based system fails. Peter Nelson, Winston Churchill, Ingmar Posner, Paul Newman 0001 |
ICRA | 3 |
| 2015 | Scheduled perception for energy-efficient path followingabstractThis paper explores the idea of reducing a robot's energy consumption while following a trajectory by turning off the main localisation subsystem and switching to a lower-powered, less accurate odometry source at appropriate times. This applies to scenarios where the robot is permitted to deviate from the original trajectory, which allows for energy savings. Sensor scheduling is formulated as a probabilistic belief planning problem. Two algorithms are presented which generate feasible perception schedules: the first is based upon a simple heuristic; the second leverages dynamic programming to obtain optimal plans. Both simulations and real-world experiments on a planetary rover prototype demonstrate over 50% savings in perception-related energy, which translates into a 12% reduction in total energy consumption. Peter Ondruska, Corina Gurau, Letizia Marchegiani, Chi Hay Tong, Ingmar Posner |
ICRA | 5 |
| 2015 | Exploiting known unknowns: Scene induced cross-calibration of lidar-stereo systemsabstractWe propose an automatic, targetless, data-driven, extrinsic calibration method to calibrate push-broom 2D lidars with a multi-camera system. The calibration problem is decoupled into alternating optimisers over two hierarchical levels, where both levels are linked with a penalty term. The lower-level optimises the six degrees-of-freedom (DoF) rigid-body transforms between the lidar and each camera of the multi-camera unit by minimising the Normalised Information Distance between intensity measurements obtained from both sensor modalities. The upper-level minimises a nonlinear least squares error between the lower-level solutions. We describe the theory, implement the method, and provide a detailed performance analysis with experiments on real-world data. Terry Scott 0002, Akshay A. Morye, Pedro Pinies, Lina María Paz, Ingmar Posner, Paul Newman 0001 |
IROS | 5 |
| 2015 | Exploiting 3D semantic scene priors for online traffic light interpretationabstractIn this paper we present a probabilistic framework for increasing online object detection performance when given a semantic 3D scene prior, which we apply to the task of traffic light detection for autonomous vehicles. Previous approaches to traffic light detection on autonomous vehicles have involved either precise knowledge of the relative 3D positions of the vehicle and the traffic light (requiring accurate and expensive mapping and localisation systems), or a classifier-based approach that searches for traffic lights in images (increasing the chance of false detections by searching all possible locations for traffic lights). We combine both approaches by explicitly incorporating both prior map and localisation uncertainty into a classifier-based object detection framework, generating a scale-space search region that only evaluates parts of the image likely to contain traffic lights, and weighting object detection scores by both the classifier score and the 3D occurrence prior distribution. We present results comparing a range of low- and high-cost localisation systems using over 30 km of data collected on an autonomous vehicle platform, demonstrating up to a 40% improvement in detection precision over no prior information and 15% improvement on unweighted detection scores. We demonstrate a 10x reduction in computation time compared to a naïve whole-image classification approach by considering only locations and scales in the image within a confidence bound of the predicted traffic light location. In addition to improvements in detection accuracy, our approach reduces computation time and enables the use of lower cost localisation sensors for reliable and cost-effective object detection. Dan Barnes, Will Maddern, Ingmar Posner |
Intelligent Vehicles Symposium | 3 |
| 2015 | Reading the Road: Road Marking Classification and InterpretationabstractRoad markings embody the rules of the road whilst capturing the upcoming road layout. These rules are diligently studied and applied to driving situations by human drivers who have read Highway Traffic driving manuals (road marking interpretation). An autonomous vehicle must however be taught to read the road, as a human might. This paper addresses the problem of automatically reading the rules encoded in road markings, by classifying them into seven distinct classes: single boundary, double boundary, separator, zig-zag, intersection, boxed junction and special lane. Our method employs a unique set of geometric feature functions within a probabilistic RUSBoost and Conditional Random Field (CRF) classification framework. This allows us to jointly classify extracted road markings. Furthermore, we infer the semantics of road scenes (pedestrian approaches and no drive regions) based on marking classification results. Finally, our algorithms are evaluated on a large real-life ground truth annotated dataset from our vehicle. Bonolo Mathibela, Paul Newman 0001, Ingmar Posner |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2014 | Probabilistic attainability maps: Efficiently predicting driver-specific electric vehicle rangeabstractThis paper concerns the efficient computation of a confidence level with which a particular driver will be able to reach a particular destination given the current state of charge of the battery of an electric vehicle. This probability of attainability is simultaneously computed for all destinations in a realistically sized map while taking into account the driver, the environment, on-board auxiliary systems and the vehicle battery system as potential sources of estimation noise. The model uses a feature-based linear regression framework which allows for a computationally efficient implementation capable of providing real-time updates of the resulting probabilistic attainability map. It was deployed on an all-electric Nissan Leaf and evaluated using data from over 140 miles of driving. The system proposed produces results of a quality commensurate with state-of-the-art approaches in terms of prediction accuracy. Peter Ondruska, Ingmar Posner |
Intelligent Vehicles Symposium | 2 |
| 2013 | Knowing when we don't know: Introspective classification for mission-critical decision makingabstractClassification precision and recall have been widely adopted by roboticists as canonical metrics to quantify the performance of learning algorithms. This paper advocates that for robotics applications, which often involve mission-critical decision making, good performance according to these standard metrics is desirable but insufficient to appropriately characterise system performance. We introduce and motivate the importance of a classifier's introspective capacity: the ability to mitigate potentially overconfident classifications by an appropriate assessment of how qualified the system is to make a judgement on the current test datum. We provide an intuition as to how this introspective capacity can be achieved and systematically investigate it in a selection of classification frameworks commonly used in robotics: support vector machines, LogitBoost classifiers and Gaussian Process classifiers (GPCs). Our experiments demonstrate that for common robotics tasks a framework such as a GPC exhibits a superior introspective capacity while maintaining commensurate classification performance to more popular, alternative approaches. Hugo Grimmett, Rohan Paul, Rudolph Triebel, Ingmar Posner |
ICRA | 4 |
| 2013 | A roadwork scene signature based on the opponent colour modelabstractThe presence of roadworks greatly affects the validity of prior maps used for navigation by autonomous vehicles. This paper addresses the problem of quickly and robustly assessing the gist of traffic scenes for whether roadworks might be present. Without explicitly modelling individual roadwork indicators such as traffic cones, construction barriers or traffic signs, our method instead only exploits the engineered visual saliency of such objects. We draw inspiration from opponent colour vision in humans to formulate a novel roadwork scene signature based on an opponent spatial prior combined with gradient information. Finally, we apply our roadwork scene signature to the task of roadwork scene recognition, within a classification framework based on soft assignment vec-torization and RUSBoost. We evaluate our roadwork signature on real life data from our autonomous vehicle. Bonolo Mathibela, Ingmar Posner, Paul Newman 0001 |
IROS | 2 |
| 2013 | Driven Learning for Driving: How Introspection Improves Semantic Mapping
Rudolph Triebel, Hugo Grimmett, Rohan Paul, Ingmar Posner |
ISRR | 4 |
| 2013 | A New Approach to Model-Free Tracking with 2D Lidar
Dominic Zeng Wang, Ingmar Posner, Paul Newman 0001 |
ISRR | 2 |
| 2013 | Toward automated driving in cities using close-to-market sensors: An overview of the V-Charge ProjectabstractFuture requirements for drastic reduction of CO2production and energy consumption will lead to significant changes in the way we see mobility in the years to come. However, the automotive industry has identified significant barriers to the adoption of electric vehicles, including reduced driving range and greatly increased refueling times. Automated cars have the potential to reduce the environmental impact of driving, and increase the safety of motor vehicle travel. The current state-of-the-art in vehicle automation requires a suite of expensive sensors. While the cost of these sensors is decreasing, integrating them into electric cars will increase the price and represent another barrier to adoption. The V-Charge Project, funded by the European Commission, seeks to address these problems simultaneously by developing an electric automated car, outfitted with close-to-market sensors, which is able to automate valet parking and recharging for integration into a future transportation system. The final goal is the demonstration of a fully operational system including automated navigation and parking. This paper presents an overview of the V-Charge system, from the platform setup to the mapping, perception, and planning sub-systems. Paul Timothy Furgale, Ulrich Schwesinger, Martin Rufli, Wojciech Derendarz, Hugo Grimmett, Peter Mühlfellner, Stefan Wonneberger, Julian Timpner, Stephan Rottmann, Bo Li 0018, Bastian Schmidt, Thien-Nghia Nguyen, Elena Cardarelli, Stefano Cattani, Stefan Bruning, Sven Horstmann, Martin Stellmacher, Holger Mielenz, Kevin Köser, Markus Beermann, Christian Häne, Lionel Heng, Gim Hee Lee, Friedrich Fraundorfer, René Iser, Rudolph Triebel, Ingmar Posner, Paul Newman 0001, Lars C. Wolf, Marc Pollefeys, Stefan Brosig, Jan Effertz, Cédric Pradalier, Roland Siegwart |
Intelligent Vehicles Symposium | 27 |
| 2012 | What could move? Finding cars, pedestrians and bicyclists in 3D laser dataabstractThis paper tackles the problem of segmenting things that could move from 3D laser scans of urban scenes. In particular, we wish to detect instances of classes of interest in autonomous driving applications - cars, pedestrians and bicyclists - amongst significant background clutter. Our aim is to provide the layout of an end-to-end pipeline which, when fed by a raw stream of 3D data, produces distinct groups of points which can be fed to downstream classifiers for categorisation. We postulate that, for the specific classes considered in this work, solving a binary classification task (i.e. separating the data into foreground and background first) outperforms approaches that tackle the multi-class problem directly. This is confirmed using custom and third-party datasets gathered of urban street scenes. While our system is agnostic to the specific clustering algorithm deployed we explore the use of a Euclidean Minimum Spanning Tree for an end-to-end segmentation pipeline and devise a RANSAC-based edge selection criterion. Dominic Zeng Wang, Ingmar Posner, Paul Newman 0001 |
ICRA | 2 |
| 2012 | Modelling Observation Correlations for Active Exploration and Robust Object DetectionabstractToday, mobile robots are expected to carry out increasingly complex tasks in multifarious, real-world environments. Often, the tasks require a certain semantic understanding of the workspace. Consider, for example, spoken instructions from a human collaborator referring to objects of interest; the robot must be able to accurately detect these objects to correctly understand the instructions. However, existing object detection, while competent, is not perfect. In particular, the performance of detection algorithms is commonly sensitive to the position of the sensor relative to the objects in the scene. This paper presents an online planning algorithm which learns an explicit model of the spatial dependence of object detection and generates plans which maximize the expected performance of the detection, and by extension the overall plan performance. Crucially, the learned sensor model incorporates spatial correlations between measurements, capturing the fact that successive measurements taken at the same or nearby locations are not independent. We show how this sensor model can be incorporated into an efficient forward search algorithm in the information space of detected objects, allowing the robot to generate motion plans efficiently. We investigate the performance of our approach by addressing the tasks of door and text detection in indoor environments and demonstrate significant improvement in detection performance during task execution over alternative methods in simulated and real robot experiments. Javier Vélez, Garrett Hemann, Albert S. Huang, Ingmar Posner, Nicholas Roy |
J. Artif. Intell. Res. | 4 |
| 2011 | Adaptive Data Compression for Robot Perception
Mike Smith 0002, Ingmar Posner, Paul Newman 0001 |
IJCAI | 2 |
| 2011 | Active Exploration for Robust Object Detection
Javier Vélez, Garrett Hemann, Albert S. Huang, Ingmar Posner, Nicholas Roy |
IJCAI | 4 |
| 2010 | Using text-spotting to query the worldabstractThe world we live in is labeled extensively for the benefit of humans. Yet, to date, robots have made little use of human readable text as a resource. In this paper we aim to draw attention to text as a readily available source of semantic information in robotics by implementing a system which allows robots to read visible text in natural scene images and to use this knowledge to interpret the content of a given scene. The reliable detection and parsing of text in natural scene images is an active area of research and remains a non-trivial problem. We extend a commonly adopted approach based on boosting for the detection and optical character recognition (OCR) for the parsing of text by a probabilistic error correction scheme incorporating a sensor-model for our pipeline. In order to interpret the scene content we introduce a generative model which explains spotted text in terms of arbitrary search terms. This allows the robot to estimate the relevance of a given scene with respect to arbitrary queries such as, for example, whether it is looking at a bank or a restaurant. We present results from images recorded by a robot in a busy cityscape. Ingmar Posner, Peter I. Corke, Paul Newman 0001 |
IROS | 1 |
| 2007 | Describing Composite Urban WorkspacesabstractIn this paper we present an appearance-based method for augmenting maps of outdoor urban environments with higher-order, semantic labels. Our motivation is to increase the value and utility of the typically low-level representations built by contemporary SLAM algorithms. A supervised learning scheme is employed to train a set of classifiers to respond to common scene attributes given a mixture of geometric and visual scene information. The union of classifier responses yields a composite description of the local workspace. We apply our method to three large data sets Ingmar Posner, Derik Schröter, Paul Newman 0001 |
ICRA | 1 |
| 2007 | Describing, Navigating and Recognising Urban Spaces - Building an End-to-End SLAM System
Paul Newman 0001, Manjari Chandran-Ramesh, Dave Cole, Mark Joseph Cummins, Alastair Harrison, Ingmar Posner, Derik Schröter |
ISRR | 6 |