Adrià Puigdomènech Badia

dblp:176/5597 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
4since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Reinforcement learning · 78% Representation and self-supervised learning · 5% Language models and text generation · 5%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 28 heaviest of 32, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
deep reinforcement learning
1.642023
Human-level Atari 200x faster · ICLR 2023
Agent57: Outperforming the Atari Human Benchmark · ICML 2020
Neural Episodic Control · ICML 2017
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning
0.712023
Human-level Atari 200x faster · ICLR 2023
Natural language and speech › Language models and text generation › large language model training › language model pretraining
BERT-style pretraining
0.612022
CoBERL: Contrastive BERT for Reinforcement Learning · ICLR 2022
Machine learning › Representation and self-supervised learning
contrastive learning
0.612022
CoBERL: Contrastive BERT for Reinforcement Learning · ICLR 2022
Machine learning › Reinforcement learning
offline reinforcement learning
0.612022
Retrieval-Augmented Reinforcement Learning · ICML 2022
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning
0.612022
CoBERL: Contrastive BERT for Reinforcement Learning · ICLR 2022
Machine learning › Reinforcement learning › partially observable reinforcement learning › memory-based reinforcement learning
retrieval-augmented reinforcement learning
0.612022
Retrieval-Augmented Reinforcement Learning · ICML 2022
Information retrieval › evaluation
benchmark
0.612022
The CLRS Algorithmic Reasoning Benchmark · ICML 2022
Machine learning › Reinforcement learning › exploration
directed exploration
0.412020
Never Give Up: Learning Directed Exploration Strategies · ICLR 2020
Machine learning › Reinforcement learning › memory architectures
episodic memory
0.412020
MEMO: A Deep Network for Flexible Combination of Episodic Memories · ICLR 2020
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff
0.412020
Agent57: Outperforming the Atari Human Benchmark · ICML 2020
Machine learning › Reinforcement learning › exploration
exploration strategies
0.412020
Never Give Up: Learning Directed Exploration Strategies · ICLR 2020
Machine learning › Reinforcement learning › exploration
intrinsic motivation
0.412020
Never Give Up: Learning Directed Exploration Strategies · ICLR 2020
Machine learning › Deep learning architectures and training
memory-augmented neural networks
0.412020
MEMO: A Deep Network for Flexible Combination of Episodic Memories · ICLR 2020
Machine learning › Reinforcement learning
generalization in reinforcement learning
0.412019
Generalization of Reinforcement Learners with Working and Episodic Memory · NeurIPS 2019
Machine learning › Reinforcement learning › partially observable reinforcement learning
memory-based reinforcement learning
0.412019
Generalization of Reinforcement Learners with Working and Episodic Memory · NeurIPS 2019
Robotics › Motion planning and robot control › robot control › adaptive control
parameter adaptation
0.312018
Memory-based Parameter Adaptation · ICLR (Poster) 2018
Machine learning › Reinforcement learning › value-based reinforcement learning
episodic control
0.312017
Neural Episodic Control · ICML 2017
Machine learning › Reinforcement learning
model-based reinforcement learning
0.312017
Imagination-Augmented Agents for Deep Reinforcement Learning · NIPS 2017
Machine learning › Reinforcement learning
model-free reinforcement learning
0.312017
Imagination-Augmented Agents for Deep Reinforcement Learning · NIPS 2017
Machine learning › Reinforcement learning
policy learning
0.312017
Imagination-Augmented Agents for Deep Reinforcement Learning · NIPS 2017
Machine learning › Reinforcement learning
value-based reinforcement learning
0.312017
Neural Episodic Control · ICML 2017
Machine learning › Reinforcement learning
actor-critic methods
0.212016
Asynchronous Methods for Deep Reinforcement Learning · ICML 2016
Machine learning › Efficient and distributed learning › distributed training
asynchronous training
0.212016
Asynchronous Methods for Deep Reinforcement Learning · ICML 2016
Knowledge, reasoning and agents › Multi-agent systems
agent architecture
0.112019
Generalization of Reinforcement Learners with Working and Episodic Memory · NeurIPS 2019
Machine learning › Reinforcement learning › model-based reinforcement learning › model-based planning
planning with learned models
0.112017
Imagination-Augmented Agents for Deep Reinforcement Learning · NIPS 2017
Machine learning › Reinforcement learning
continuous control
0.112016
Asynchronous Methods for Deep Reinforcement Learning · ICML 2016
Robotics › Robot navigation and mapping
visual navigation
0.112016
Asynchronous Methods for Deep Reinforcement Learning · ICML 2016

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.3graph neural network · 1.1episodic memory · 0.7retrieval network · 0.6r2d2 · 0.6contrastive learning · 0.6DQN · 0.6BERT · 0.6neural network policy parameterization · 0.4deep reinforcement learning · 0.4adaptive policy selection · 0.4
YearPublicationVenuePosition
2023 Human-level Atari 200x faster
Steven Kapturowski, Victor Campos 0001, Ray Jiang, Nemanja Rakicevic, Hado van Hasselt, Charles Blundell, Adrià Puigdomènech Badia
ICLR7
2022 CoBERL: Contrastive BERT for Reinforcement Learning
Andrea Banino, Adrià Puigdomènech Badia, Jacob C. Walker, Tim Scholtes, Jovana Mitrovic, Charles Blundell
ICLR2
2022 Retrieval-Augmented Reinforcement Learning
abstract
Most deep reinforcement learning (RL) algorithms distill experience into parametric behavior policies or value functions via gradient updates. While effective, this approach has several disadvantages: (1) it is computationally expensive, (2) it can take many updates to integrate experiences into the parametric model, (3) experiences that are not fully integrated do not appropriately influence the agent’s behavior, and (4) behavior is limited by the capacity of the model. In this paper we explore an alternative paradigm in which we train a network to map a dataset of past experiences to optimal behavior. Specifically, we augment an RL agent with a retrieval process (parameterized as a neural network) that has direct access to a dataset of experiences. This dataset can come from the agent’s past experiences, expert demonstrations, or any other relevant source. The retrieval process is trained to retrieve information from the dataset that may be useful in the current context, to help the agent achieve its goal faster and more efficiently. The proposed method facilitates learning agents that at test time can condition their behavior on the entire dataset and not only the current state, or current trajectory. We integrate our method into two different RL agents: an offline DQN agent and an online R2D2 agent. In offline multi-task problems, we show that the retrieval-augmented DQN agent avoids task interference and learns faster than the baseline DQN agent. On Atari, we show that retrieval-augmented R2D2 learns significantly faster than the baseline R2D2 agent and achieves higher scores. We run extensive ablations to measure the contributions of the components of our proposed method.
Anirudh Goyal, Abram L. Friesen, Andrea Banino, Theophane Weber, Nan Rosemary Ke, Adrià Puigdomènech Badia, Arthur Guez, Mehdi Mirza, Peter Conway Humphreys, Ksenia Konyushkova, Michal Valko, Simon Osindero, Timothy P. Lillicrap, Nicolas Heess, Charles Blundell
ICML6
2022 The CLRS Algorithmic Reasoning Benchmark
abstract
Learning representations of algorithms is an emerging area of machine learning, seeking to bridge concepts from neural networks with classical algorithms. Several important works have investigated whether neural networks can effectively reason like algorithms, typically by learning to execute them. The common trend in the area, however, is to generate targeted kinds of algorithmic data to evaluate specific hypotheses, making results hard to transfer across publications, and increasing the barrier of entry. To consolidate progress and work towards unified evaluation, we propose the CLRS Algorithmic Reasoning Benchmark, covering classical algorithms from the Introduction to Algorithms textbook. Our benchmark spans a variety of algorithmic reasoning procedures, including sorting, searching, dynamic programming, graph algorithms, string algorithms and geometric algorithms. We perform extensive experiments to demonstrate how several popular algorithmic reasoning baselines perform on these tasks, and consequently, highlight links to several open challenges. Our library is readily available at https://github.com/deepmind/clrs.
Petar Velickovic, Adrià Puigdomènech Badia, David Budden, Razvan Pascanu, Andrea Banino, Misha Dashevskiy, Raia Hadsell, Charles Blundell
ICML2
2020 Never Give Up: Learning Directed Exploration Strategies
Adrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Guo, Bilal Piot, Steven Kapturowski, Olivier Tieleman, Martín Arjovsky, Alexander Pritzel, Andrew Bolt, Charles Blundell
ICLR1
2020 MEMO: A Deep Network for Flexible Combination of Episodic Memories
Andrea Banino, Adrià Puigdomènech Badia, Raphael Koster, Martin J. Chadwick, Vinícius Flores Zambaldi, Demis Hassabis, Caswell Barry, Matt M. Botvinick, Dharshan Kumaran, Charles Blundell
ICLR2
2020 Agent57: Outperforming the Atari Human Benchmark
abstract
Atari games have been a long-standing benchmark in the reinforcement learning (RL) community for the past decade. This benchmark was proposed to test general competency of RL algorithms. Previous work has achieved good average performance by doing outstandingly well on many games of the set, but very poorly in several of the most challenging games. We propose Agent57, the first deep RL agent that outperforms the standard human benchmark on all 57 Atari games. To achieve this result, we train a neural network which parameterizes a family of policies ranging from very exploratory to purely exploitative. We propose an adaptive mechanism to choose which policy to prioritize throughout the training process. Additionally, we utilize a novel parameterization of the architecture that allows for more consistent and stable learning.
Adrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Guo, Charles Blundell
ICML1
2019 Generalization of Reinforcement Learners with Working and Episodic Memory
abstract
Memory is an important aspect of intelligence and plays a role in many deep reinforcement learning models. However, little progress has been made in understanding when specific memory systems help more than others and how well they generalize. The field also has yet to see a prevalent consistent and rigorous approach for evaluating agent performance on holdout data. In this paper, we aim to develop a comprehensive methodology to test different kinds of memory in an agent and assess how well the agent can apply what it learns in training to a holdout set that differs from the training set along dimensions that we suggest are relevant for evaluating memory-specific generalization. To that end, we first construct a diverse set of memory tasks that allow us to evaluate test-time generalization across multiple dimensions. Second, we develop and perform multiple ablations on an agent architecture that combines multiple memory systems, observe its baseline models, and investigate its performance against the task suite.
Meire Fortunato, Melissa Tan, Ryan Faulkner 0001, Steven Hansen 0001, Adrià Puigdomènech Badia, Gavin Buttimore, Charlie Deck, Joel Z. Leibo, Charles Blundell
NeurIPS5
2018 Memory-based Parameter Adaptation
Pablo Sprechmann, Siddhant M. Jayakumar, Jack W. Rae, Alexander Pritzel, Adrià Puigdomènech Badia, Benigno Uria, Oriol Vinyals, Demis Hassabis, Razvan Pascanu, Charles Blundell
ICLR (Poster)5
2017 Neural Episodic Control
abstract
Deep reinforcement learning methods attain super-human performance in a wide range of environments. Such methods are grossly inefficient, often taking orders of magnitudes more data than humans to achieve reasonable performance. We propose Neural Episodic Control: a deep reinforcement learning agent that is able to rapidly assimilate new experiences and act upon them. Our agent uses a semi-tabular representation of the value function: a buffer of past experience containing slowly changing state representations and rapidly updated estimates of the value function. We show across a wide range of environments that our agent learns significantly faster than other state-of-the-art, general purpose deep reinforcement learning agents.
Alexander Pritzel, Benigno Uria, Sriram Srinivasan 0005, Adrià Puigdomènech Badia, Oriol Vinyals, Demis Hassabis, Daan Wierstra, Charles Blundell
ICML4
2017 Imagination-Augmented Agents for Deep Reinforcement Learning
abstract
We introduce Imagination-Augmented Agents (I2As), a novel architecture for deep reinforcement learning combining model-free and model-based aspects. In contrast to most existing model-based reinforcement learning and planning methods, which prescribe how a model should be used to arrive at a policy, I2As learn to interpret predictions from a trained environment model to construct implicit plans in arbitrary ways, by using the predictions as additional context in deep policy networks. I2As show improved data efficiency, performance, and robustness to model misspecification compared to several strong baselines.
Sébastien Racanière, Theophane Weber, David P. Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adrià Puigdomènech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li 0001, Razvan Pascanu, Peter W. Battaglia, Demis Hassabis, David Silver 0001, Daan Wierstra
NIPS7
2016 Asynchronous Methods for Deep Reinforcement Learning
abstract
We propose a conceptually simple and lightweight framework for deep reinforcement learning that uses asynchronous gradient descent for optimization of deep neural network controllers. We present asynchronous variants of four standard reinforcement learning algorithms and show that parallel actor-learners have a stabilizing effect on training allowing all four methods to successfully train neural network controllers. The best performing method, an asynchronous variant of actor-critic, surpasses the current state-of-the-art on the Atari domain while training for half the time on a single multi-core CPU instead of a GPU. Furthermore, we show that asynchronous actor-critic succeeds on a wide variety of continuous motor control problems as well as on a new task of navigating random 3D mazes using a visual input.
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver 0001, Koray Kavukcuoglu
ICML2