Przemyslaw Mazur

dblp:223/4490 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
0since 2021 · last 2020
0000-0003-0025-9410ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3Systems, architecture and hardware · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 42% Autonomous driving · 42% Probabilistic and Bayesian machine learning · 13%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Autonomous driving
end-to-end driving
0.412020
Urban Driving with Conditional Imitation Learning · ICRA 2020
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff
0.412019
Successor Uncertainties: Exploration and Uncertainty in Temporal Difference Learning · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › sampling
posterior sampling
0.412019
Successor Uncertainties: Exploration and Uncertainty in Temporal Difference Learning · NeurIPS 2019
Machine learning › Reinforcement learning › exploration
randomized value functions
0.412019
Successor Uncertainties: Exploration and Uncertainty in Temporal Difference Learning · NeurIPS 2019
Machine learning › Reinforcement learning
value function
0.412019
Successor Uncertainties: Exploration and Uncertainty in Temporal Difference Learning · NeurIPS 2019
Machine learning › Reinforcement learning
model-free reinforcement learning
0.112019
Learning to Drive in a Day · ICRA 2019

Methods — techniques the papers use, named apart from their topics

imitation learning · 0.4data balancing · 0.4convolutional neural network · 0.4temporal difference learning · 0.4neural network function approximation · 0.4deep reinforcement learning · 0.4continuous control · 0.4
YearPublicationVenuePosition
2020 Urban Driving with Conditional Imitation Learning
abstract
Hand-crafting generalised decision-making rules for real-world urban autonomous driving is hard. Alternatively, learning behaviour from easy-to-collect human driving demonstrations is appealing. Prior work has studied imitation learning (IL) for autonomous driving with a number of limitations. Examples include only performing lane-following rather than following a user-defined route, only using a single camera view or heavily cropped frames lacking state observability, only lateral (steering) control, but not longitudinal (speed) control and a lack of interaction with traffic. Importantly, the majority of such systems have been primarily evaluated in simulation - a simple domain, which lacks real-world complexities. Motivated by these challenges, we focus on learning representations of semantics, geometry and motion with computer vision for IL from human driving demonstrations. As our main contribution, we present an end-to-end conditional imitation learning approach, combining both lateral and longitudinal control on a real vehicle for following urban routes with simple traffic. We address inherent dataset bias by data balancing, training our final policy on approximately 30 hours of demonstrations gathered over six months. We evaluate our method on an autonomous vehicle by driving 35km of novel routes in European urban streets.
Jeffrey Hawke, Richard Shen, Corina Gurau, Daniele Reda, Nikolay Nikolov, Przemyslaw Mazur, Sean Micklethwaite, Nicolas Griffiths, Amar Shah 0001, Alex Kendall
ICRA7
2019 Learning to Drive in a Day
abstract
We demonstrate the first application of deep reinforcement learning to autonomous driving. From randomly initialised parameters, our model is able to learn a policy for lane following in a handful of training episodes using a single monocular image as input. We provide a general and easy to obtain reward: the distance travelled by the vehicle without the safety driver taking control. We use a continuous, model-free deep reinforcement learning algorithm, with all exploration and optimisation performed on-vehicle. This demonstrates a new framework for autonomous driving which moves away from reliance on defined logical rules, mapping, and direct supervision. We discuss the challenges and opportunities to scale this approach to a broader range of autonomous driving tasks.
Alex Kendall, Jeffrey Hawke, David Janz, Przemyslaw Mazur, Daniele Reda, John-Mark Allen, Vinh-Dieu Lam, Alex Bewley, Amar Shah 0001
ICRA4
2019 Successor Uncertainties: Exploration and Uncertainty in Temporal Difference Learning
abstract
Posterior sampling for reinforcement learning (PSRL) is an effective method for balancing exploration and exploitation in reinforcement learning. Randomised value functions (RVF) can be viewed as a promising approach to scaling PSRL. However, we show that most contemporary algorithms combining RVF with neural network function approximation do not possess the properties which make PSRL effective, and provably fail in sparse reward problems. Moreover, we find that propagation of uncertainty, a property of PSRL previously thought important for exploration, does not preclude this failure. We use these insights to design Successor Uncertainties (SU), a cheap and easy to implement RVF algorithm that retains key properties of PSRL. SU is highly effective on hard tabular exploration benchmarks. Furthermore, on the Atari 2600 domain, it surpasses human performance on 38 of 49 games tested (achieving a median human normalised score of 2.09), and outperforms its closest RVF competitor, Bootstrapped DQN, on 36 of those.
David Janz, Jiri Hron, Przemyslaw Mazur, Katja Hofmann, José Miguel Hernández-Lobato, Sebastian Tschiatschek
NeurIPS3