EDBT 2026 Demo / reviewers in the wild / expert
Przemyslaw Mazur
dblp:223/4490
· DBLP profile ↗
3ranked-venue papers
0as first author
0since 2021 · last 2020
0000-0003-0025-9410ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3Systems, architecture and hardware · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Reinforcement learning · 42% Autonomous driving · 42% Probabilistic and Bayesian machine learning · 13% |
Topics — the 6 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Autonomous driving
end-to-end driving |
0.4 | 1 | 2020 | Urban Driving with Conditional Imitation Learning · ICRA 2020 |
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff |
0.4 | 1 | 2019 | Successor Uncertainties: Exploration and Uncertainty in Temporal Difference Learning · NeurIPS 2019 |
Machine learning › Probabilistic and Bayesian machine learning › sampling
posterior sampling |
0.4 | 1 | 2019 | Successor Uncertainties: Exploration and Uncertainty in Temporal Difference Learning · NeurIPS 2019 |
Machine learning › Reinforcement learning › exploration
randomized value functions |
0.4 | 1 | 2019 | Successor Uncertainties: Exploration and Uncertainty in Temporal Difference Learning · NeurIPS 2019 |
Machine learning › Reinforcement learning
value function |
0.4 | 1 | 2019 | Successor Uncertainties: Exploration and Uncertainty in Temporal Difference Learning · NeurIPS 2019 |
Machine learning › Reinforcement learning
model-free reinforcement learning |
0.1 | 1 | 2019 | Learning to Drive in a Day · ICRA 2019 |
Methods — techniques the papers use, named apart from their topics
imitation learning · 0.4data balancing · 0.4convolutional neural network · 0.4temporal difference learning · 0.4neural network function approximation · 0.4deep reinforcement learning · 0.4continuous control · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Urban Driving with Conditional Imitation LearningabstractHand-crafting generalised decision-making rules for real-world urban autonomous driving is hard. Alternatively, learning behaviour from easy-to-collect human driving demonstrations is appealing. Prior work has studied imitation learning (IL) for autonomous driving with a number of limitations. Examples include only performing lane-following rather than following a user-defined route, only using a single camera view or heavily cropped frames lacking state observability, only lateral (steering) control, but not longitudinal (speed) control and a lack of interaction with traffic. Importantly, the majority of such systems have been primarily evaluated in simulation - a simple domain, which lacks real-world complexities. Motivated by these challenges, we focus on learning representations of semantics, geometry and motion with computer vision for IL from human driving demonstrations. As our main contribution, we present an end-to-end conditional imitation learning approach, combining both lateral and longitudinal control on a real vehicle for following urban routes with simple traffic. We address inherent dataset bias by data balancing, training our final policy on approximately 30 hours of demonstrations gathered over six months. We evaluate our method on an autonomous vehicle by driving 35km of novel routes in European urban streets. Jeffrey Hawke, Richard Shen, Corina Gurau, Daniele Reda, Nikolay Nikolov, Przemyslaw Mazur, Sean Micklethwaite, Nicolas Griffiths, Amar Shah 0001, Alex Kendall |
ICRA | 7 |
| 2019 | Learning to Drive in a DayabstractWe demonstrate the first application of deep reinforcement learning to autonomous driving. From randomly initialised parameters, our model is able to learn a policy for lane following in a handful of training episodes using a single monocular image as input. We provide a general and easy to obtain reward: the distance travelled by the vehicle without the safety driver taking control. We use a continuous, model-free deep reinforcement learning algorithm, with all exploration and optimisation performed on-vehicle. This demonstrates a new framework for autonomous driving which moves away from reliance on defined logical rules, mapping, and direct supervision. We discuss the challenges and opportunities to scale this approach to a broader range of autonomous driving tasks. Alex Kendall, Jeffrey Hawke, David Janz, Przemyslaw Mazur, Daniele Reda, John-Mark Allen, Vinh-Dieu Lam, Alex Bewley, Amar Shah 0001 |
ICRA | 4 |
| 2019 | Successor Uncertainties: Exploration and Uncertainty in Temporal Difference LearningabstractPosterior sampling for reinforcement learning (PSRL) is an effective method for balancing exploration and exploitation in reinforcement learning. Randomised value functions (RVF) can be viewed as a promising approach to scaling PSRL. However, we show that most contemporary algorithms combining RVF with neural network function approximation do not possess the properties which make PSRL effective, and provably fail in sparse reward problems. Moreover, we find that propagation of uncertainty, a property of PSRL previously thought important for exploration, does not preclude this failure. We use these insights to design Successor Uncertainties (SU), a cheap and easy to implement RVF algorithm that retains key properties of PSRL. SU is highly effective on hard tabular exploration benchmarks. Furthermore, on the Atari 2600 domain, it surpasses human performance on 38 of 49 games tested (achieving a median human normalised score of 2.09), and outperforms its closest RVF competitor, Bootstrapped DQN, on 36 of those. David Janz, Jiri Hron, Przemyslaw Mazur, Katja Hofmann, José Miguel Hernández-Lobato, Sebastian Tschiatschek |
NeurIPS | 3 |