Alexander Pritzel

dblp:168/8345 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Reinforcement learning · 54% Trustworthy machine learning · 10% Transfer learning and domain adaptation · 8%

Topics — the 23 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
deep reinforcement learning
0.622018
Fast deep reinforcement learning using online adjustments from the past · NeurIPS 2018
Neural Episodic Control · ICML 2017
Machine learning › Reinforcement learning › exploration
directed exploration
0.412020
Never Give Up: Learning Directed Exploration Strategies · ICLR 2020
Machine learning › Reinforcement learning › exploration
exploration strategies
0.412020
Never Give Up: Learning Directed Exploration Strategies · ICLR 2020
Machine learning › Reinforcement learning › exploration
intrinsic motivation
0.412020
Never Give Up: Learning Directed Exploration Strategies · ICLR 2020
Machine learning › Reinforcement learning › off-policy reinforcement learning
experience replay
0.312018
Fast deep reinforcement learning using online adjustments from the past · NeurIPS 2018
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning › state representation learning
latent dynamics model
0.312018
Generative Temporal Models with Spatial Memory for Partially Observed Environments · ICML 2018
Machine learning › Reinforcement learning
model-based reinforcement learning
0.312018
Generative Temporal Models with Spatial Memory for Partially Observed Environments · ICML 2018
Robotics › Motion planning and robot control › robot control › adaptive control
parameter adaptation
0.312018
Memory-based Parameter Adaptation · ICLR (Poster) 2018
Robotics › Robot navigation and mapping › spatial representation
spatial memory
0.312018
Generative Temporal Models with Spatial Memory for Partially Observed Environments · ICML 2018
Machine learning › Reinforcement learning
value function estimation
0.312018
Fast deep reinforcement learning using online adjustments from the past · NeurIPS 2018
Machine learning › Generative modeling
variational autoencoder
0.312018
Generative Temporal Models with Spatial Memory for Partially Observed Environments · ICML 2018
Computer vision › Video understanding and tracking
video prediction
0.312018
Generative Temporal Models with Spatial Memory for Partially Observed Environments · ICML 2018
Machine learning › Kernel, tree and ensemble methods › ensemble learning
deep ensembles
0.312017
Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles · NIPS 2017
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.312017
DARLA: Improving Zero-Shot Transfer in Reinforcement Learning · ICML 2017
Machine learning › Reinforcement learning › value-based reinforcement learning
episodic control
0.312017
Neural Episodic Control · ICML 2017
Machine learning › Trustworthy machine learning › uncertainty estimation
predictive uncertainty
0.312017
Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles · NIPS 2017
Machine learning › Trustworthy machine learning
uncertainty estimation
0.312017
Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles · NIPS 2017
Machine learning › Reinforcement learning
value-based reinforcement learning
0.312017
Neural Episodic Control · ICML 2017
Machine learning › Transfer learning and domain adaptation
zero-shot transfer
0.312017
DARLA: Improving Zero-Shot Transfer in Reinforcement Learning · ICML 2017
Machine learning › Reinforcement learning › deep reinforcement learning
deep q-network
0.212016
Deep Exploration via Bootstrapped DQN · NIPS 2016
Machine learning › Reinforcement learning
exploration
0.212016
Deep Exploration via Bootstrapped DQN · NIPS 2016
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection
0.112017
Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles · NIPS 2017
Machine learning › Trustworthy machine learning
robustness
0.112017
Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles · NIPS 2017

Methods — techniques the papers use, named apart from their topics

deep reinforcement learning · 0.4variational inference · 0.3state space model · 0.3prioritised sweeping · 0.3memory-based parameter adaptation · 0.3ephemeral value adjustments · 0.3disentangled representation learning · 0.3a3c · 0.3EC · 0.3DQN · 0.3
YearPublicationVenuePosition
2020 Never Give Up: Learning Directed Exploration Strategies
Adrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Guo, Bilal Piot, Steven Kapturowski, Olivier Tieleman, Martín Arjovsky, Alexander Pritzel, Andrew Bolt, Charles Blundell
ICLR9
2018 Memory-based Parameter Adaptation
Pablo Sprechmann, Siddhant M. Jayakumar, Jack W. Rae, Alexander Pritzel, Adrià Puigdomènech Badia, Benigno Uria, Oriol Vinyals, Demis Hassabis, Razvan Pascanu, Charles Blundell
ICLR (Poster)4
2018 Generative Temporal Models with Spatial Memory for Partially Observed Environments
abstract
In model-based reinforcement learning, generative and temporal models of environments can be leveraged to boost agent performance, either by tuning the agent’s representations during training or via use as part of an explicit planning mechanism. However, their application in practice has been limited to simplistic environments, due to the difficulty of training such models in larger, potentially partially-observed and 3D environments. In this work we introduce a novel action-conditioned generative model of such challenging environments. The model features a non-parametric spatial memory system in which we store learned, disentangled representations of the environment. Low-dimensional spatial updates are computed using a state-space model that makes use of knowledge on the prior dynamics of the moving agent, and high-dimensional visual observations are modelled with a Variational Auto-Encoder. The result is a scalable architecture capable of performing coherent predictions over hundreds of time steps across a range of partially observed 2D and 3D environments.
Marco Fraccaro, Danilo Jimenez Rezende, Yori Zwols, Alexander Pritzel, S. M. Ali Eslami, Fabio Viola
ICML4
2018 Fast deep reinforcement learning using online adjustments from the past
abstract
We propose Ephemeral Value Adjusments (EVA): a means of allowing deep reinforcement learning agents to rapidly adapt to experience in their replay buffer. EVA shifts the value predicted by a neural network with an estimate of the value function found by prioritised sweeping over experience tuples from the replay buffer near the current state. EVA combines a number of recent ideas around combining episodic memory-like structures into reinforcement learning agents: slot-based storage, content-based retrieval, and memory-based planning. We show that EVA is performant on a demonstration task and Atari games.
Steven Hansen 0001, Alexander Pritzel, Pablo Sprechmann, André Barreto 0001, Charles Blundell
NeurIPS2
2017 DARLA: Improving Zero-Shot Transfer in Reinforcement Learning
abstract
Domain adaptation is an important open problem in deep reinforcement learning (RL). In many scenarios of interest data is hard to obtain, so agents may learn a source policy in a setting where data is readily available, with the hope that it generalises well to the target domain. We propose a new multi-stage RL agent, DARLA (DisentAngled Representation Learning Agent), which learns to see before learning to act. DARLA’s vision is based on learning a disentangled representation of the observed environment. Once DARLA can see, it is able to acquire source policies that are robust to many domain shifts – even with no access to the target domain. DARLA significantly outperforms conventional baselines in zero-shot domain adaptation scenarios, an effect that holds across a variety of RL environments (Jaco arm, DeepMind Lab) and base RL algorithms (DQN, A3C and EC).
Irina Higgins, Arka Pal, Andrei A. Rusu, Loïc Matthey, Chris Burgess 0001, Alexander Pritzel, Matt M. Botvinick, Charles Blundell, Alexander Lerchner
ICML6
2017 Neural Episodic Control
abstract
Deep reinforcement learning methods attain super-human performance in a wide range of environments. Such methods are grossly inefficient, often taking orders of magnitudes more data than humans to achieve reasonable performance. We propose Neural Episodic Control: a deep reinforcement learning agent that is able to rapidly assimilate new experiences and act upon them. Our agent uses a semi-tabular representation of the value function: a buffer of past experience containing slowly changing state representations and rapidly updated estimates of the value function. We show across a wide range of environments that our agent learns significantly faster than other state-of-the-art, general purpose deep reinforcement learning agents.
Alexander Pritzel, Benigno Uria, Sriram Srinivasan 0005, Adrià Puigdomènech Badia, Oriol Vinyals, Demis Hassabis, Daan Wierstra, Charles Blundell
ICML1
2017 Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles
abstract
Deep neural networks (NNs) are powerful black box predictors that have recently achieved impressive performance on a wide spectrum of tasks. Quantifying predictive uncertainty in NNs is a challenging and yet unsolved problem. Bayesian NNs, which learn a distribution over weights, are currently the state-of-the-art for estimating predictive uncertainty; however these require significant modifications to the training procedure and are computationally expensive compared to standard (non-Bayesian) NNs. We propose an alternative to Bayesian NNs that is simple to implement, readily parallelizable, requires very little hyperparameter tuning, and yields high quality predictive uncertainty estimates. Through a series of experiments on classification and regression benchmarks, we demonstrate that our method produces well-calibrated uncertainty estimates which are as good or better than approximate Bayesian NNs. To assess robustness to dataset shift, we evaluate the predictive uncertainty on test examples from known and unknown distributions, and show that our method is able to express higher uncertainty on out-of-distribution examples. We demonstrate the scalability of our method by evaluating predictive uncertainty estimates on ImageNet.
Balaji Lakshminarayanan, Alexander Pritzel, Charles Blundell
NIPS2
2016 Deep Exploration via Bootstrapped DQN
abstract
Efficient exploration remains a major challenge for reinforcement learning (RL). Common dithering strategies for exploration, such as epsilon-greedy, do not carry out temporally-extended (or deep) exploration; this can lead to exponentially larger data requirements. However, most algorithms for statistically efficient RL are not computationally tractable in complex environments. Randomized value functions offer a promising approach to efficient exploration with generalization, but existing algorithms are not compatible with nonlinearly parameterized value functions. As a first step towards addressing such contexts we develop bootstrapped DQN. We demonstrate that bootstrapped DQN can combine deep exploration with deep neural networks for exponentially faster learning than any dithering strategy. In the Arcade Learning Environment bootstrapped DQN substantially improves learning speed and cumulative performance across most games.
Ian Osband, Charles Blundell, Alexander Pritzel, Benjamin Van Roy
NIPS3