Peter Ondruska

dblp:148/7155 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
3since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 3 since 2021Systems, architecture and hardware · 6 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2022 SafetyNet: Safe Planning for Real-World Self-Driving Vehicles Using Machine-Learned Policies
abstract
In this paper we present the first safe system for full control of self-driving vehicles trained from human demonstrations and deployed in challenging, real-world, urban environments. Current industry-standard solutions use rule-based systems for planning. Although they perform reasonably well in common scenarios, the engineering complexity renders this approach incompatible with human-level performance. On the other hand, the performance of machine-learned (ML) planning solutions can be improved by simply adding more exemplar data. However, ML methods cannot offer safety guarantees and sometimes behave unpredictably. To combat this, our approach uses a simple yet effective rule-based fallback layer that performs sanity checks on an ML planner's decisions (e.g. avoiding collision, assuring physical feasibility). This allows us to leverage ML to handle complex situations while still assuring the safety, reducing ML planner-only collisions by 95%. We train our ML planner on 300 hours of expert driving demonstrations using imitation learning and deploy it along with the fallback layer in downtown San Francisco, where it takes complete control of a real vehicle and navigates a wide variety of challenging urban driving scenarios.
Matt Vitelli, Yan Chang, Yawei Ye, Ana Sofia Rufino Ferreira, Maciej Wolczyk, Blazej Osinski, Moritz Niendorf, Hugo Grimmett, Qiangui Huang, Ashesh Jain, Peter Ondruska
ICRA11
2021 SimNet: Learning Reactive Self-driving Simulations from Real-world Observations
abstract
In this work we present a simple end-to-end trainable machine learning system capable of realistically simulating driving experiences. This can be used for verification of self-driving system performance without relying on expensive and time-consuming road testing. In particular, we frame the simulation problem as a Markov Process, leveraging deep neural networks to model both state distribution and transition function. These are trainable directly from the existing raw observations without the need of any handcrafting in the form of plant or kinematic models. All that is needed is a dataset of historical traffic episodes. Our formulation allows the system to construct never seen scenes that unfold realistically reacting to the self-driving car’s behaviour. We train our system directly from 1,000 hours of driving logs and measure both realism, reactivity of the simulation as the two key properties of the simulation. At the same time we apply the method to evaluate performance of a recently proposed state-of-the-art ML planning system [1] trained from human driving logs. We discover this planning system is prone to previously unreported causal confusion issues that are difficult to test by non-reactive simulation. To the best of our knowledge, this is the first work that directly merges highly realistic data-driven simulations with a closed loop evaluation for self-driving vehicles. We make the data, code, and pre-trained models publicly available to further stimulate simulation development.
Luca Bergamini, Yawei Ye, Oliver Scheel, Chih Hu, Luca Del Pero, Blazej Osinski, Hugo Grimmett, Peter Ondruska
ICRA9
2021 What data do we need for training an AV motion planner?
abstract
We investigate what grade of sensor data is required for training an imitation-learning-based AV planner on human expert demonstration. Machine-learned planners [1] are very hungry for training data, which is usually collected using vehicles equipped with the same sensors used for autonomous operation [1]. This is costly and non-scalable. If cheaper sensors could be used for collection instead, data availability would go up, which is crucial in a field where data volume requirements are large and availability is small. We present experiments using up to 1000 hours worth of expert demonstration and find that training with 10x lower-quality data outperforms 1x AV-grade data in terms of planner performance (see Fig. 1). The important implication of this is that cheaper sensors can indeed be used. This serves to improve data access and democratize the field of imitation-based motion planning. Alongside this, we perform a sensitivity analysis of planner performance as a function of perception range, field-of-view, accuracy, and data volume, and reason about why lower-quality data still provide good planning results.
Lukas Platinsky, Stefanie Speichert, Blazej Osinski, Oliver Scheel, Yawei Ye, Hugo Grimmett, Luca Del Pero, Peter Ondruska
ICRA9
2020 Collaborative Augmented Reality on Smartphones via Life-long City-scale Maps
abstract
In this paper we present the first published end-to-end production computer-vision system for powering city-scale shared augmented reality experiences on mobile devices. In doing so we propose a new formulation for an experience-based mapping framework as an effective solution to the key issues of city-scale SLAM scalability, robustness, map updates and all-time all-weather performance required by a production system. Furthermore, we propose an effective way of synchronising SLAM systems to deliver seamless real-time localisation of multiple edge devices at the same time. All this in the presence of network latency and bandwidth limitations. The resulting system is deployed and tested at scale in San Francisco where it delivers AR experiences in a mapped area of several hundred kilometers. To foster further development of this area we offer the data set to the public, constituting the largest of this kind to date.
Lukas Platinsky, Michal Szabados, Filip Hlasek, Ross Hemsley, Luca Del Pero, Andrej Pancik, Bryan Baum, Hugo Grimmett, Peter Ondruska
ISMAR9
2018 VALUE: Large Scale Voting-Based Automatic Labelling for Urban Environments
abstract
This paper presents a simple and robust method for the automatic localisation of static 3D objects in large-scale urban environments. By exploiting the potential to merge a large volume of noisy but accurately localised 2D image data, we achieve superior performance in terms of both robustness and accuracy of the recovered 3D information. The method is based on a simple distributed voting schema which can be fully distributed and parallelised to scale to large-scale scenarios. To evaluate the method we collected city-scale data sets from New York City and San Francisco consisting of almost 400k images spanning the area of 40 km2and used it to accurately recover the 3D positions of traffic lights. We demonstrate a robust performance and also show that the solution improves in quality over time as the amount of data increases.
Giacomo Dabisias, Emanuele Ruffaldi, Hugo Grimmett, Peter Ondruska
ICRA4
2018 Visual Vehicle Tracking Through Noise and Occlusions Using Crowd-Sourced Maps
abstract
We present a location-specific method to visually track the positions of observed vehicles based on large-scale crowd-sourced maps. We equipped a large fleet of cars that drive around cities with camera phones mounted on the dashboard, and performed city-scale structure-from-motion to accurately reconstruct the trajectories taken by the vehicles. We show that these data can be used to first create a system enabling high-accuracy localisation, and then to accurately predict the future motion of newly observed cars in the camera view. As a basis for the method we use a recently proposed system [1] for unsupervised motion prediction and extend it to a real-time visual tracking pipeline which can track vehicles through noise and extended occlusions using only a monocular camera. The system is tested using two large-scale datasets of San Francisco and New York City containing millions of frames. We demonstrate the performance of the system in a variety of traffic, time, and weather conditions. The presented system requires no manual annotation or knowledge of road infrastructure. To our knowledge, this is the first time a perception system based on a large-scale crowd-sourced maps has been evaluated at this scale.
M. S. Suraj, Hugo Grimmett, Lukas Platinsky, Peter Ondruska
IROS4
2018 Predicting trajectories of vehicles using large-scale motion priors
abstract
We present a simple yet effective paradigm to accurately predict the future trajectories of observed vehicles in dense city environments. We equipped a large fleet of cars with cameras and performed city-scale structure-from-motion to accurately reconstruct 10M positions of their trajectories spanning over 1000h of driving.We demonstrate that this information can be used as a powerful high-fidelity prior to predict future trajectories of newly observed vehicles in the area without the need for any knowledge of road infrastructure or vehicle motion models. By relating the current position of the observed car to a large dataset of the previously exhibited motion in the area we can directly perform prediction of its future position.We evaluate our method on two large-scale data sets from San Francisco and New York City and demonstrate an order of magnitude improvement compared to a linear-motion based method. We also demonstrate that the performance naturally improves with the amount of data and ultimately yields a system that can accurately predict vehicle motion in challenging situations across extremes in traffic, time, and weather conditions.
M. S. Suraj, Hugo Grimmett, Lukas Platinsky, Peter Ondruska
Intelligent Vehicles Symposium4
2016 Deep Tracking: Seeing Beyond Seeing Using Recurrent Neural Networks
abstract
This paper presents to the best of our knowledge the first end-to-end object tracking approach which directly maps from raw sensor input to object tracks in sensor space without requiring any feature engineering or system identification in the form of plant or sensor models. Specifically, our system accepts a stream of raw sensor data at one end and, in real-time, produces an estimate of the entire environment state at the output including even occluded objects. We achieve this by framing the problem as a deep learning task and exploit sequence models in the form of recurrent neural networks to learn a mapping from sensor measurements to object tracks. In particular, we propose a learning method based on a form of input dropout which allows learning in an unsupervised manner, only based on raw, occluded sensor data without access to ground-truth annotations. We demonstrate our approach using a synthetic dataset designed to mimic the task of tracking objects in 2D laser data — as commonly encountered in robotics applications — and show that it learns to track many dynamic objects despite occlusions and the presence of sensor noise.
Peter Ondruska, Ingmar Posner
AAAI1
2016 Ask Me Anything: Dynamic Memory Networks for Natural Language Processing
abstract
Most tasks in natural language processing can be cast into question answering (QA) problems over language input. We introduce the dynamic memory network (DMN), a neural network architecture which processes input sequences and questions, forms episodic memories, and generates relevant answers. Questions trigger an iterative attention process which allows the model to condition its attention on the inputs and the result of previous iterations. These results are then reasoned over in a hierarchical recurrent sequence model to generate answers. The DMN can be trained end-to-end and obtains state-of-the-art results on several types of tasks and datasets: question answering (Facebook’s bAbI dataset), text classification for sentiment analysis (Stanford Sentiment Treebank) and sequence modeling for part-of-speech tagging (WSJ-PTB). The training for these different tasks relies exclusively on trained word vector representations and input-question-answer triplets.
Ozan Irsoy, Peter Ondruska, Mohit Iyyer, James Bradbury 0002, Ishaan Gulrajani, Victor Zhong, Romain Paulus, Richard Socher
ICML3
2015 Scheduled perception for energy-efficient path following
abstract
This paper explores the idea of reducing a robot's energy consumption while following a trajectory by turning off the main localisation subsystem and switching to a lower-powered, less accurate odometry source at appropriate times. This applies to scenarios where the robot is permitted to deviate from the original trajectory, which allows for energy savings. Sensor scheduling is formulated as a probabilistic belief planning problem. Two algorithms are presented which generate feasible perception schedules: the first is based upon a simple heuristic; the second leverages dynamic programming to obtain optimal plans. Both simulations and real-world experiments on a planetary rover prototype demonstrate over 50% savings in perception-related energy, which translates into a 12% reduction in total energy consumption.
Peter Ondruska, Corina Gurau, Letizia Marchegiani, Chi Hay Tong, Ingmar Posner
ICRA1
2015 MobileFusion: Real-Time Volumetric Surface Reconstruction and Dense Tracking on Mobile Phones
abstract
We present the first pipeline for real-time volumetric surface reconstruction and dense 6DoF camera tracking running purely on standard, off-the-shelf mobile phones. Using only the embedded RGB camera, our system allows users to scan objects of varying shape, size, and appearance in seconds, with real-time feedback during the capture process. Unlike existing state of the art methods, which produce only point-based 3D models on the phone, or require cloud-based processing, our hybrid GPU/CPU pipeline is unique in that it creates a connected 3D surface model directly on the device at 25Hz. In each frame, we perform dense 6DoF tracking, which continuously registers the RGB input to the incrementally built 3D model, minimizing a noise aware photoconsistency error metric. This is followed by efficient key-frame selection, and dense per-frame stereo matching. These depth maps are fused volumetrically using a method akin to KinectFusion, producing compelling surface models. For each frame, the implicit surface is extracted for live user feedback and pose estimation. We demonstrate scans of a variety of objects, and compare to a Kinect-based baseline, showing on average ∼ 1.5cm error. We qualitatively compare to a state of the art point-based mobile phone method, demonstrating an order of magnitude faster scanning times, and fully connected surface models.
Peter Ondruska, Pushmeet Kohli, Shahram Izadi
IEEE Trans. Vis. Comput. Graph.1
2014 Probabilistic attainability maps: Efficiently predicting driver-specific electric vehicle range
abstract
This paper concerns the efficient computation of a confidence level with which a particular driver will be able to reach a particular destination given the current state of charge of the battery of an electric vehicle. This probability of attainability is simultaneously computed for all destinations in a realistically sized map while taking into account the driver, the environment, on-board auxiliary systems and the vehicle battery system as potential sources of estimation noise. The model uses a feature-based linear regression framework which allows for a computationally efficient implementation capable of providing real-time updates of the resulting probabilistic attainability map. It was deployed on an all-electric Nissan Leaf and evaluated using data from over 140 miles of driving. The system proposed produces results of a quality commensurate with state-of-the-art approaches in terms of prediction accuracy.
Peter Ondruska, Ingmar Posner
Intelligent Vehicles Symposium1