VLDB 2026 Research / reviewers in the wild / expert
Matteo Hessel
dblp:162/3167
· DBLP profile ↗
24ranked-venue papers
4as first author
7since 2021 · last 2025
0009-0006-9946-4375ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
22 papers |
Reinforcement learning · 86% Transfer learning and domain adaptation · 6% Optimization for machine learning · 4% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 100% |
Topics — the 30 heaviest of 45, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
value-based reinforcement learning |
1.6 | 5 | 2020 | Discovering Reinforcement Learning Algorithms · NeurIPS 2020 Multi-Task Deep Reinforcement Learning with PopArt · AAAI 2019 Rainbow: Combining Improvements in Deep Reinforcement Learning · AAAI 2018 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
1.2 | 3 | 2021 | Self-Consistent Models and Values · NeurIPS 2021 When to use parametric models in reinforcement learning? · NeurIPS 2019 The Predictron: End-To-End Learning and Planning · ICML 2017 |
Machine learning › Reinforcement learning › off-policy reinforcement learning
experience replay |
1.1 | 3 | 2020 | Off-Policy Actor-Critic with Shared Experience Replay · ICML 2020 When to use parametric models in reinforcement learning? · NeurIPS 2019 Distributed Prioritized Experience Replay · ICLR (Poster) 2018 |
Machine learning › Reinforcement learning
temporal difference learning |
1.0 | 2 | 2021 | Emphatic Algorithms for Deep Reinforcement Learning · ICML 2021 Expected Eligibility Traces · AAAI 2021 |
Machine learning › Reinforcement learning
value function estimation |
0.9 | 2 | 2021 | Self-Consistent Models and Values · NeurIPS 2021 Discovery of Useful Questions as Auxiliary Tasks · NeurIPS 2019 |
Machine learning › Reinforcement learning
actor-critic methods |
0.9 | 2 | 2020 | A Self-Tuning Actor-Critic Algorithm · NeurIPS 2020 Off-Policy Actor-Critic with Shared Experience Replay · ICML 2020 |
Machine learning › Reinforcement learning › meta-reinforcement learning
meta-gradient reinforcement learning |
0.9 | 2 | 2020 | A Self-Tuning Actor-Critic Algorithm · NeurIPS 2020 Meta-Gradient Reinforcement Learning with an Objective Discovered Online · NeurIPS 2020 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
0.9 | 2 | 2020 | Meta-Gradient Reinforcement Learning with an Objective Discovered Online · NeurIPS 2020 Discovering Reinforcement Learning Algorithms · NeurIPS 2020 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.9 | 3 | 2021 | Multi-Task Deep Reinforcement Learning with PopArt · AAAI 2019 Rainbow: Combining Improvements in Deep Reinforcement Learning · AAAI 2018 Emphatic Algorithms for Deep Reinforcement Learning · ICML 2021 |
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent |
0.6 | 1 | 2022 | Learning by Directional Gradient Descent · ICLR 2022 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
credit assignment |
0.5 | 1 | 2021 | Expected Eligibility Traces · AAAI 2021 |
Machine learning › Reinforcement learning › reinforcement learning theory
deadly triad |
0.5 | 1 | 2021 | Emphatic Algorithms for Deep Reinforcement Learning · ICML 2021 |
Machine learning › Reinforcement learning › temporal difference learning
eligibility traces |
0.5 | 1 | 2021 | Expected Eligibility Traces · AAAI 2021 |
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.5 | 1 | 2021 | Discovery of Options via Meta-Learned Subgoals · NeurIPS 2021 |
Machine learning › Reinforcement learning
off-policy reinforcement learning |
0.5 | 1 | 2021 | Emphatic Algorithms for Deep Reinforcement Learning · ICML 2021 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option discovery |
0.5 | 1 | 2021 | Discovery of Options via Meta-Learned Subgoals · NeurIPS 2021 |
Machine learning › Reinforcement learning
policy evaluation |
0.5 | 1 | 2021 | Self-Consistent Models and Values · NeurIPS 2021 |
Machine learning › Reinforcement learning
policy optimization |
0.5 | 1 | 2021 | Muesli: Combining Improvements in Policy Optimization · ICML 2021 |
Natural language and speech › Language models and text generation
self-consistency |
0.5 | 1 | 2021 | Self-Consistent Models and Values · NeurIPS 2021 |
Machine learning › Optimization for machine learning
hyperparameter optimization |
0.4 | 1 | 2020 | A Self-Tuning Actor-Critic Algorithm · NeurIPS 2020 |
Machine learning › Reinforcement learning › exploration
intrinsic motivation |
0.4 | 1 | 2020 | What Can Learned Intrinsic Rewards Capture? · ICML 2020 |
Machine learning › Reinforcement learning
learned objective functions |
0.4 | 1 | 2020 | Meta-Gradient Reinforcement Learning with an Objective Discovered Online · NeurIPS 2020 |
Machine learning › Reinforcement learning › meta-reinforcement learning
learned update rules |
0.4 | 1 | 2020 | Discovering Reinforcement Learning Algorithms · NeurIPS 2020 |
Machine learning › Transfer learning and domain adaptation › meta-learning
meta-gradient |
0.4 | 1 | 2020 | What Can Learned Intrinsic Rewards Capture? · ICML 2020 |
Machine learning › Reinforcement learning › actor-critic methods
off-policy actor-critic |
0.4 | 1 | 2020 | Off-Policy Actor-Critic with Shared Experience Replay · ICML 2020 |
Machine learning › Reinforcement learning
reward learning |
0.4 | 1 | 2020 | What Can Learned Intrinsic Rewards Capture? · ICML 2020 |
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
auxiliary tasks |
0.4 | 1 | 2019 | Discovery of Useful Questions as Auxiliary Tasks · NeurIPS 2019 |
Machine learning › Reinforcement learning › value function estimation
general value functions |
0.4 | 1 | 2019 | Discovery of Useful Questions as Auxiliary Tasks · NeurIPS 2019 |
Machine learning › Reinforcement learning
multi-task reinforcement learning |
0.4 | 1 | 2019 | Multi-Task Deep Reinforcement Learning with PopArt · AAAI 2019 |
Machine learning › Reinforcement learning › large-scale reinforcement learning
distributed reinforcement learning |
0.3 | 1 | 2018 | Distributed Prioritized Experience Replay · ICLR (Poster) 2018 |
Methods — techniques the papers use, named apart from their topics
bootstrapping · 0.9meta-gradient · 0.9meta-gradient descent · 0.9actor-critic · 0.8experience replay · 0.7successor features · 0.5regularized policy optimization · 0.5muzero · 0.5model learning · 0.5emphatic temporal difference · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DataRater: Meta-Learned Dataset CurationabstractThe quality of foundation models depends heavily on their training data.
Consequently, great efforts have been put into dataset curation.
Yet most approaches rely on manual tuning of coarse-grained mixtures of large buckets of data, or filtering by hand-crafted heuristics.
An approach that is ultimately more scalable (let alone more satisfying) is to \emph{learn} which data is actually valuable for training.
This type of meta-learning could allow more sophisticated, fine-grained, and effective curation.
Our proposed \emph{DataRater} is an instance of this idea. It estimates the value of training on any particular data point. This is done by meta-learning using `meta-gradients', with the objective of improving training efficiency on held out data.
In extensive experiments across a range of model scales and datasets, we find that using our DataRater to filter data is highly effective, resulting in significantly improved compute efficiency. Dan Andrei Calian, Gregory Farquhar, Iurii Kemaev, Luisa M. Zintgraf, Matteo Hessel, Jeremy Shar, Junhyuk Oh, András György 0001, Tom Schaul, Jeffrey Dean, Hado van Hasselt, David Silver 0001 |
NeurIPS | 5 |
| 2022 | Learning by Directional Gradient Descent
David Silver 0001, Anirudh Goyal, Ivo Danihelka, Matteo Hessel, Hado van Hasselt |
ICLR | 4 |
| 2021 | Expected Eligibility TracesabstractThe question of how to determine which states and actions are responsible for a certain outcome is known as the credit assignment problem and remains a central research question in reinforcement learning and artificial intelligence. Eligibility traces enable efficient credit assignment to the recent sequence of states and actions experienced by the agent, but not to counterfactual sequences that could also have led to the current state. In this work, we introduce expected eligibility traces. Expected traces allow, with a single update, to update states and actions that could have preceded the current state, even if they did not do so on this occasion. We discuss when expected traces provide benefits over classic (instantaneous) traces in temporal-difference learning, and show that some- times substantial improvements can be attained. We provide a way to smoothly interpolate between instantaneous and expected traces by a mechanism similar to bootstrapping, which ensures that the resulting algorithm is a strict generalisation of TD(λ). Finally, we discuss possible extensions and connections to related ideas, such as successor features. Hado van Hasselt, Sephora Madjiheurem, Matteo Hessel, David Silver 0001, André Barreto 0001, Diana Borsa |
AAAI | 3 |
| 2021 | Muesli: Combining Improvements in Policy OptimizationabstractWe propose a novel policy update that combines regularized policy optimization with model learning as an auxiliary loss. The update (henceforth Muesli) matches MuZero’s state-of-the-art performance on Atari. Notably, Muesli does so without using deep search: it acts directly with a policy network and has computation speed comparable to model-free baselines. The Atari results are complemented by extensive ablations, and by additional results on continuous control and 9x9 Go. Matteo Hessel, Ivo Danihelka, Fabio Viola, Arthur Guez, Simon Schmitt, Laurent Sifre, Theophane Weber, David Silver 0001, Hado van Hasselt |
ICML | 1 |
| 2021 | Emphatic Algorithms for Deep Reinforcement LearningabstractOff-policy learning allows us to learn about possible policies of behavior from experience generated by a different behavior policy. Temporal difference (TD) learning algorithms can become unstable when combined with function approximation and off-policy sampling—this is known as the “deadly triad”. Emphatic temporal difference (ETD($\lambda$)) algorithm ensures convergence in the linear case by appropriately weighting the TD($\lambda$) updates. In this paper, we extend the use of emphatic methods to deep reinforcement learning agents. We show that naively adapting ETD($\lambda$) to popular deep reinforcement learning algorithms, which use forward view multi-step returns, results in poor performance. We then derive new emphatic algorithms for use in the context of such algorithms, and we demonstrate that they provide noticeable benefits in small problems designed to highlight the instability of TD methods. Finally, we observed improved performance when applying these algorithms at scale on classic Atari games from the Arcade Learning Environment. Ray Jiang, Tom Zahavy, Zhongwen Xu, Adam White 0001, Matteo Hessel, Charles Blundell, Hado van Hasselt |
ICML | 5 |
| 2021 | Self-Consistent Models and ValuesabstractLearned models of the environment provide reinforcement learning (RL) agents with flexible ways of making predictions about the environment.Models enable planning, i.e. using more computation to improve value functions or policies, without requiring additional environment interactions.In this work, we investigate a way of augmenting model-based RL, by additionally encouraging a learned model and value function to be jointly \emph{self-consistent}.This lies in contrast to classic planning methods like Dyna, which only update the value function to be consistent with the model.We propose a number of possible self-consistency updates, study them empirically in both the tabular and function approximation settings, and find that with appropriate choices self-consistency can be useful both for policy evaluation and control. Gregory Farquhar, Kate Baumli, Zita Marinho, Angelos Filos, Matteo Hessel, Hado van Hasselt, David Silver 0001 |
NeurIPS | 5 |
| 2021 | Discovery of Options via Meta-Learned SubgoalsabstractTemporal abstractions in the form of options have been shown to help reinforcement learning (RL) agents learn faster. However, despite prior work on this topic, the problem of discovering options through interaction with an environment remains a challenge. In this paper, we introduce a novel meta-gradient approach for discovering useful options in multi-task RL environments. Our approach is based on a manager-worker decomposition of the RL agent, in which a manager maximises rewards from the environment by learning a task-dependent policy over both a set of task-independent discovered-options and primitive actions. The option-reward and termination functions that define a subgoal for each option are parameterised as neural networks and trained via meta-gradients to maximise their usefulness. Empirical analysis on gridworld and DeepMind Lab tasks show that: (1) our approach can discover meaningful and diverse temporally-extended options in multi-task RL domains, (2) the discovered options are frequently used by the agent while learning to solve the training tasks, and (3) that the discovered options help a randomly initialised manager learn faster in completely new tasks. Vivek Veeriah, Tom Zahavy, Matteo Hessel, Zhongwen Xu, Junhyuk Oh, Iurii Kemaev, Hado van Hasselt, David Silver 0001, Satinder Singh 0001 |
NeurIPS | 3 |
| 2020 | Behaviour Suite for Reinforcement Learning
Ian Osband, Yotam Doron, Matteo Hessel, John Aslanides, Eren Sezener, Andre Saraiva 0001, Katrina McKinney, Tor Lattimore, Csaba Szepesvári, Satinder Singh 0001, Benjamin Van Roy, Richard S. Sutton, David Silver 0001, Hado van Hasselt |
ICLR | 3 |
| 2020 | Off-Policy Actor-Critic with Shared Experience ReplayabstractWe investigate the combination of actor-critic reinforcement learning algorithms with a uniform large-scale experience replay and propose solutions for two ensuing challenges: (a) efficient actor-critic learning with experience replay (b) the stability of off-policy learning where agents learn from other agents behaviour. To this end we analyze the bias-variance tradeoffs in V-trace, a form of importance sampling for actor-critic methods. Based on our analysis, we then argue for mixing experience sampled from replay with on-policy experience, and propose a new trust region scheme that scales effectively to data distributions where V-trace becomes unstable. We provide extensive empirical validation of the proposed solutions on DMLab-30 and further show the benefits of this setup in two training regimes for Atari: (1) a single agent is trained up until 200M environment frames per game (2) a population of agents is trained up until 200M environment frames each and may share experience. We demonstrate state-of-the-art data efficiency among model-free agents in both regimes. Simon Schmitt, Matteo Hessel, Karen Simonyan |
ICML | 2 |
| 2020 | What Can Learned Intrinsic Rewards Capture?abstractThe objective of a reinforcement learning agent is to behave so as to maximise the sum of a suitable scalar function of state: the reward. These rewards are typically given and immutable. In this paper, we instead consider the proposition that the reward function itself can be a good locus of learned knowledge. To investigate this, we propose a scalable meta-gradient framework for learning useful intrinsic reward functions across multiple lifetimes of experience. Through several proof-of-concept experiments, we show that it is feasible to learn and capture knowledge about long-term exploration and exploitation into a reward function. Furthermore, we show that unlike policy transfer methods that capture “how” the agent should behave, the learned reward functions can generalise to other kinds of agents and to changes in the dynamics of the environment by capturing “what” the agent should strive to do. Junhyuk Oh, Matteo Hessel, Zhongwen Xu, Manuel Kroiss, Hado van Hasselt, David Silver 0001, Satinder Singh 0001 |
ICML | 3 |
| 2020 | Discovering Reinforcement Learning AlgorithmsabstractReinforcement learning (RL) algorithms update an agent’s parameters according to one of several possible rules, discovered manually through years of research. Automating the discovery of update rules from data could lead to more efficient algorithms, or algorithms that are better adapted to specific environments. Although there have been prior attempts at addressing this significant scientific challenge, it remains an open question whether it is feasible to discover alternatives to fundamental concepts of RL such as value functions and temporal-difference learning. This paper introduces a new meta-learning approach that discovers an entire update rule which includes both what to predict' (e.g. value functions) andhow to learn from it' (e.g. bootstrapping) by interacting with a set of environments. The output of this method is an RL algorithm that we call Learned Policy Gradient (LPG). Empirical results show that our method discovers its own alternative to the concept of value functions. Furthermore it discovers a bootstrapping mechanism to maintain and use its predictions. Surprisingly, when trained solely on toy environments, LPG generalises effectively to complex Atari games and achieves non-trivial performance. This shows the potential to discover general RL algorithms from data. Junhyuk Oh, Matteo Hessel, Wojciech Czarnecki 0001, Zhongwen Xu, Hado van Hasselt, Satinder Singh 0001, David Silver 0001 |
NeurIPS | 2 |
| 2020 | Meta-Gradient Reinforcement Learning with an Objective Discovered OnlineabstractDeep reinforcement learning includes a broad family of algorithms that parameterise an internal representation, such as a value function or policy, by a deep neural network. Each algorithm optimises its parameters with respect to an objective, such as Q-learning or policy gradient, that defines its semantics. In this work, we propose an algorithm based on meta-gradient descent that discovers its own objective, flexibly parameterised by a deep neural network, solely from interactive experience with its environment. Over time, this allows the agent to learn how to learn increasingly effectively. Furthermore, because the objective is discovered online, it can adapt to changes over time. We demonstrate that the algorithm discovers how to address several important issues in RL, such as bootstrapping, non-stationarity, and off-policy learning. On the Atari Learning Environment, the meta-gradient algorithm adapts over time to learn with greater efficiency, eventually outperforming the median score of a strong actor-critic baseline. Zhongwen Xu, Hado van Hasselt, Matteo Hessel, Junhyuk Oh, Satinder Singh 0001, David Silver 0001 |
NeurIPS | 3 |
| 2020 | A Self-Tuning Actor-Critic AlgorithmabstractReinforcement learning algorithms are highly sensitive to the choice of hyperparameters, typically requiring significant manual effort to identify hyperparameters that perform well on a new domain. In this paper, we take a step towards addressing this issue by using metagradients to automatically adapt hyperparameters online by meta-gradient descent (Xu et al., 2018). We apply our algorithm, Self-Tuning Actor-Critic (STAC), to self-tune all the differentiable hyperparameters of an actor-critic loss function, to discover auxiliary tasks, and to improve off-policy learning using a novel leaky V-trace operator. STAC is simple to use, sample efficient and does not require a significant increase in compute. Ablative studies show that the overall performance of STAC improved as we adapt more hyperparameters. When applied to the Arcade Learning Environment (Bellemare et al. 2012), STAC improved the median human normalized score in 200M steps from 243% to 364%. When applied to the DM Control suite (Tassa et al., 2018), STAC improved the mean score in 30M steps from 217 to 389 when learning with features, from 108 to 202 when learning from pixels, and from 195 to 295 in the Real-World Reinforcement Learning Challenge (Dulac-Arnold et al., 2020). Tom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel, Junhyuk Oh, Hado van Hasselt, David Silver 0001, Satinder Singh 0001 |
NeurIPS | 4 |
| 2019 | Multi-Task Deep Reinforcement Learning with PopArtabstractThe reinforcement learning (RL) community has made great strides in designing algorithms capable of exceeding human performance on specific tasks. These algorithms are mostly trained one task at the time, each new task requiring to train a brand new agent instance. This means the learning algorithm is general, but each solution is not; each agent can only solve the one task it was trained on. In this work, we study the problem of learning to master not one but multiple sequentialdecision tasks at once. A general issue in multi-task learning is that a balance must be found between the needs of multiple tasks competing for the limited resources of a single learning system. Many learning algorithms can get distracted by certain tasks in the set of tasks to solve. Such tasks appear more salient to the learning process, for instance because of the density or magnitude of the in-task rewards. This causes the algorithm to focus on those salient tasks at the expense of generality. We propose to automatically adapt the contribution of each task to the agent’s updates, so that all tasks have a similar impact on the learning dynamics. This resulted in state of the art performance on learning to play all games in a set of 57 diverse Atari games. Excitingly, our method learned a single trained policy - with a single set of weights - that exceeds median human performance. To our knowledge, this was the first time a single agent surpassed human-level performance on this multi-task domain. The same approach also demonstrated state of the art performance on a set of 30 tasks in the 3D reinforcement learning platform DeepMind Lab. Matteo Hessel, Hubert Soyer, Lasse Espeholt, Wojciech Czarnecki 0001, Simon Schmitt, Hado van Hasselt |
AAAI | 1 |
| 2019 | When to use parametric models in reinforcement learning?abstractWe examine the question of when and how parametric models are most useful in reinforcement learning. In particular, we look at commonalities and differences between parametric models and experience replay. Replay-based learning algorithms share important traits with model-based approaches, including the ability to plan: to use more computation without additional data to improve predictions and behaviour. We discuss when to expect benefits from either approach, and interpret prior work in this context. We hypothesise that, under suitable conditions, replay-based algorithms should be competitive to or better than model-based algorithms if the model is used only to generate fictional transitions from observed states for an update rule that is otherwise model-free. We validated this hypothesis on Atari 2600 video games. The replay-based algorithm attained state-of-the-art data efficiency, improving over prior results with parametric models. Additionally, we discuss different ways to use models. We show that it can be better to plan backward than to plan forward when using models to perform credit assignment (e.g., to directly learn a value or policy), even though the latter seems more common. Finally, we argue and demonstrate that it can be beneficial to plan forward for immediate behaviour, rather than for credit assignment. Hado van Hasselt, Matteo Hessel, John Aslanides |
NeurIPS | 2 |
| 2019 | Discovery of Useful Questions as Auxiliary TasksabstractArguably, intelligent agents ought to be able to discover their own questions so that in learning answers for them they learn unanticipated useful knowledge and skills; this departs from the focus in much of machine learning on agents learning answers to externally defined questions. We present a novel method for a reinforcement learning (RL) agent to discover questions formulated as general value functions or GVFs, a fairly rich form of knowledge representation. Specifically, our method uses non-myopic meta-gradients to learn GVF-questions such that learning answers to them, as an auxiliary task, induces useful representations for the main task faced by the RL agent. We demonstrate that auxiliary tasks based on the discovered GVFs are sufficient, on their own, to build representations that support main task learning, and that they do so better than popular hand-designed auxiliary tasks from the literature. Furthermore, we show, in the context of Atari2600 videogames, how such auxiliary tasks, meta-learned alongside the main task, can improve the data efficiency of an actor-critic agent. Vivek Veeriah, Matteo Hessel, Zhongwen Xu, Janarthanan Rajendran, Richard L. Lewis, Junhyuk Oh, Hado van Hasselt, David Silver 0001, Satinder Singh 0001 |
NeurIPS | 2 |
| 2018 | Rainbow: Combining Improvements in Deep Reinforcement LearningabstractThe deep reinforcement learning community has made several independent improvements to the DQN algorithm. However, it is unclear which of these extensions are complementary and can be fruitfully combined. This paper examines six extensions to the DQN algorithm and empirically studies their combination. Our experiments show that the combination provides state-of-the-art performance on the Atari 2600 benchmark, both in terms of data efficiency and final performance. We also provide results from a detailed ablation study that shows the contribution of each component to overall performance. Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Daniel Horgan, Bilal Piot, Mohammad Gheshlaghi Azar, David Silver 0001 |
AAAI | 1 |
| 2018 | Noisy Networks For Exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Matteo Hessel, Ian Osband, Alex Graves, Volodymyr Mnih, Rémi Munos, Demis Hassabis, Olivier Pietquin, Charles Blundell, Shane Legg |
ICLR (Poster) | 5 |
| 2018 | Distributed Prioritized Experience Replay
Daniel Horgan, John Quan, David Budden, Gabriel Barth-Maron, Matteo Hessel, Hado van Hasselt, David Silver 0001 |
ICLR (Poster) | 5 |
| 2018 | Transfer in Deep Reinforcement Learning Using Successor Features and Generalised Policy ImprovementabstractThe ability to transfer skills across tasks has the potential to scale up reinforcement learning (RL) agents to environments currently out of reach. Recently, a framework based on two ideas, successor features (SFs) and generalised policy improvement (GPI), has been introduced as a principled way of transferring skills. In this paper we extend the SF&GPI framework in two ways. One of the basic assumptions underlying the original formulation of SF&GPI is that rewards for all tasks of interest can be computed as linear combinations of a fixed set of features. We relax this constraint and show that the theoretical guarantees supporting the framework can be extended to any set of tasks that only differ in the reward function. Our second contribution is to show that one can use the reward functions themselves as features for future tasks, without any loss of expressiveness, thus removing the need to specify a set of features beforehand. This makes it possible to combine SF&GPI with deep learning in a more stable way. We empirically verify this claim on a complex 3D environment where observations are images from a first-person perspective. We show that the transfer promoted by SF&GPI leads to very good policies on unseen tasks almost instantaneously. We also describe how to learn policies specialised to the new tasks in a way that allows them to be added to the agent’s set of skills, and thus be reused in the future. André Barreto 0001, Diana Borsa, John Quan, Tom Schaul, David Silver 0001, Matteo Hessel, Daniel J. Mankowitz, Augustin Zídek, Rémi Munos |
ICML | 6 |
| 2017 | The Predictron: End-To-End Learning and PlanningabstractOne of the key challenges of artificial intelligence is to learn models that are effective in the context of planning. In this document we introduce the predictron architecture. The predictron consists of a fully abstract model, represented by a Markov reward process, that can be rolled forward multiple “imagined” planning steps. Each forward pass of the predictron accumulates internal rewards and values over multiple planning depths. The predictron is trained end-to-end so as to make these accumulated values accurately approximate the true value function. We applied the predictron to procedurally generated random mazes and a simulator for the game of pool. The predictron yielded significantly more accurate predictions than conventional deep neural network architectures. David Silver 0001, Hado van Hasselt, Matteo Hessel, Tom Schaul, Arthur Guez, Tim Harley, Gabriel Dulac-Arnold, David P. Reichert, Neil C. Rabinowitz, André Barreto 0001, Thomas Degris |
ICML | 3 |
| 2016 | Dueling Network Architectures for Deep Reinforcement LearningabstractIn recent years there have been many successes of using deep representations in reinforcement learning. Still, many of these applications use conventional architectures, such as convolutional networks, LSTMs, or auto-encoders. In this paper, we present a new neural network architecture for model-free reinforcement learning. Our dueling network represents two separate estimators: one for the state value function and one for the state-dependent action advantage function. The main benefit of this factoring is to generalize learning across actions without imposing any change to the underlying reinforcement learning algorithm. Our results show that this architecture leads to better policy evaluation in the presence of many similar-valued actions. Moreover, the dueling architecture enables our RL agent to outperform the state-of-the-art on the Atari 2600 domain. Ziyu Wang 0001, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot, Nando de Freitas |
ICML | 3 |
| 2016 | Learning values across many orders of magnitudeabstractMost learning algorithms are not invariant to the scale of the signal that is being approximated. We propose to adaptively normalize the targets used in the learning updates. This is important in value-based reinforcement learning, where the magnitude of appropriate value approximations can change over time when we update the policy of behavior. Our main motivation is prior work on learning to play Atari games, where the rewards were clipped to a predetermined range. This clipping facilitates learning across many different games with a single learning algorithm, but a clipped reward function can result in qualitatively different behavior. Using adaptive normalization we can remove this domain-specific heuristic without diminishing overall performance. Hado van Hasselt, Arthur Guez, Matteo Hessel, Volodymyr Mnih, David Silver 0001 |
NIPS | 3 |
| 2014 | A novel approach to model design and tuning through automatic parameter screening and optimization theory and application to a helicopter flight simulator case-study
Matteo Hessel, Francesco Borgatelli, Fabio Ortalli |
SIMULTECH | 1 |