EDBT 2026 Demo / reviewers in the wild / expert
Steven Kapturowski
dblp:245/4747
· DBLP profile ↗
10ranked-venue papers
2as first author
5since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Reinforcement learning · 81% Deep learning architectures and training · 11% Optimization for machine learning · 3% | |
| Human-computer interaction and pervasive computing
1 paper |
Games and playful interaction · 100% |
Topics — the 26 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
exploration |
1.2 | 2 | 2024 | Unlocking the Power of Representations in Long-term Novelty-based Exploration · ICLR 2024 Making Efficient Use of Demonstrations to Solve Hard Exploration Problems · ICLR 2020 |
Machine learning › Reinforcement learning
deep reinforcement learning |
1.1 | 2 | 2023 | Human-level Atari 200x faster · ICLR 2023 Agent57: Outperforming the Atari Human Benchmark · ICML 2020 |
Machine learning › Reinforcement learning
actor-critic methods |
0.8 | 1 | 2024 | Offline Actor-Critic Reinforcement Learning Scales to Large Models · ICML 2024 |
Machine learning › Reinforcement learning › exploration › novelty-based exploration
count-based exploration |
0.8 | 1 | 2024 | Unlocking the Power of Representations in Long-term Novelty-based Exploration · ICLR 2024 |
Machine learning › Reinforcement learning
multi-task reinforcement learning |
0.8 | 1 | 2024 | Offline Actor-Critic Reinforcement Learning Scales to Large Models · ICML 2024 |
Machine learning › Reinforcement learning › exploration
novelty-based exploration |
0.8 | 1 | 2024 | Unlocking the Power of Representations in Long-term Novelty-based Exploration · ICLR 2024 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.8 | 1 | 2024 | Offline Actor-Critic Reinforcement Learning Scales to Large Models · ICML 2024 |
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | Transformers need glasses! Information over-squashing in language tasks · NeurIPS 2024 |
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning |
0.7 | 1 | 2023 | Human-level Atari 200x faster · ICLR 2023 |
Machine learning › Optimization for machine learning
convergence analysis |
0.5 | 1 | 2021 | Revisiting Peng's Q(λ) for Modern Reinforcement Learning · ICML 2021 |
Machine learning › Reinforcement learning
temporal difference learning |
0.5 | 1 | 2021 | Revisiting Peng's Q(λ) for Modern Reinforcement Learning · ICML 2021 |
Machine learning › Reinforcement learning › exploration
directed exploration |
0.4 | 1 | 2020 | Never Give Up: Learning Directed Exploration Strategies · ICLR 2020 |
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff |
0.4 | 1 | 2020 | Agent57: Outperforming the Atari Human Benchmark · ICML 2020 |
Machine learning › Reinforcement learning › exploration
exploration strategies |
0.4 | 1 | 2020 | Never Give Up: Learning Directed Exploration Strategies · ICLR 2020 |
Machine learning › Reinforcement learning › exploration › exploration in markov decision processes
hard exploration |
0.4 | 1 | 2020 | Making Efficient Use of Demonstrations to Solve Hard Exploration Problems · ICLR 2020 |
Machine learning › Reinforcement learning › exploration
intrinsic motivation |
0.4 | 1 | 2020 | Never Give Up: Learning Directed Exploration Strategies · ICLR 2020 |
Robotics › Robot manipulation
learning from demonstration |
0.4 | 1 | 2020 | Making Efficient Use of Demonstrations to Solve Hard Exploration Problems · ICLR 2020 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.4 | 1 | 2020 | Value-driven Hindsight Modelling · NeurIPS 2020 |
Machine learning › Reinforcement learning
value-based reinforcement learning |
0.4 | 1 | 2020 | Value-driven Hindsight Modelling · NeurIPS 2020 |
Machine learning › Reinforcement learning
value function estimation |
0.4 | 1 | 2020 | Value-driven Hindsight Modelling · NeurIPS 2020 |
Machine learning › Reinforcement learning › large-scale reinforcement learning
distributed reinforcement learning |
0.4 | 1 | 2019 | Recurrent Experience Replay in Distributed Reinforcement Learning · ICLR (Poster) 2019 |
Machine learning › Reinforcement learning › off-policy reinforcement learning
experience replay |
0.4 | 1 | 2019 | Recurrent Experience Replay in Distributed Reinforcement Learning · ICLR (Poster) 2019 |
Natural language and speech › Language models and text generation › pre-trained language model
decoder-only language model |
0.2 | 1 | 2024 | Transformers need glasses! Information over-squashing in language tasks · NeurIPS 2024 |
Machine learning › Reinforcement learning › deep reinforcement learning
scaling laws for reinforcement learning |
0.2 | 1 | 2024 | Offline Actor-Critic Reinforcement Learning Scales to Large Models · ICML 2024 |
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
state representation |
0.2 | 1 | 2024 | Unlocking the Power of Representations in Long-term Novelty-based Exploration · ICLR 2024 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.1 | 1 | 2019 | Recurrent Experience Replay in Distributed Reinforcement Learning · ICLR (Poster) 2019 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.3transformer · 0.8signal propagation analysis · 0.8perceiver · 0.8masked transformer · 0.8low-precision floating-point analysis · 0.8inverse dynamics loss · 0.8clustering-based online density estimation · 0.8behavioral cloning · 0.8conservative policy iteration · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Unlocking the Power of Representations in Long-term Novelty-based ExplorationabstractWe introduce Robust Exploration via Clustering-based Online Density Estimation (RECODE), a non-parametric method for novelty-based exploration that estimates visitation counts for clusters of states based on their similarity in a chosen embedding space. By adapting classical clustering to the nonstationary setting of Deep RL, RECODE can efficiently track state visitation counts over thousands of episodes. We further propose a novel generalization of the inverse dynamics loss, which leverages masked transformer architectures for multi-step prediction; which in conjunction with \DETOCS achieves a new state-of-the-art in a suite of challenging 3D-exploration tasks in DM-Hard-8. RECODE also sets new state-of-the-art in hard exploration Atari games, and is the first agent to reach the end screen in "Pitfall!" Alaa Saade, Steven Kapturowski, Daniele Calandriello, Charles Blundell, Pablo Sprechmann, Leopoldo Sarra, Oliver Groth, Michal Valko, Bilal Piot |
ICLR | 2 |
| 2024 | Offline Actor-Critic Reinforcement Learning Scales to Large ModelsabstractWe show that offline actor-critic reinforcement learning can scale to large models - such as transformers - and follows similar scaling laws as supervised learning. We find that offline actor-critic algorithms can outperform strong, supervised, behavioral cloning baselines for multi-task training on a large dataset; containing both sub-optimal and expert behavior on 132 continuous control tasks. We introduce a Perceiver-based actor-critic model and elucidate the key features needed to make offline RL work with self- and cross-attention modules. Overall, we find that: i) simple offline actor critic algorithms are a natural choice for gradually moving away from the currently predominant paradigm of behavioral cloning, and ii) via offline RL it is possible to learn multi-task policies that master many domains simultaneously, including real robotics tasks, from sub-optimal demonstrations or self-generated data. Jost Tobias Springenberg, Abbas Abdolmaleki, Jingwei Zhang 0001, Oliver Groth, Michael Bloesch, Thomas Lampe, Philemon Brakel, Sarah Bechtle, Steven Kapturowski, Roland Hafner, Nicolas Heess, Martin A. Riedmiller |
ICML | 9 |
| 2024 | Transformers need glasses! Information over-squashing in language tasksabstractWe study how information propagates in decoder-only Transformers, which are the architectural foundation of most existing frontier large language models (LLMs). We rely on a theoretical signal propagation analysis---specifically, we analyse the representations of the last token in the final layer of the Transformer, as this is the representation used for next-token prediction. Our analysis reveals a representational collapse phenomenon: we prove that certain distinct pairs of inputs to the Transformer can yield arbitrarily close representations in the final token. This effect is exacerbated by the low-precision floating-point formats frequently used in modern LLMs. As a result, the model is provably unable to respond to these sequences in different ways---leading to errors in, e.g., tasks involving counting or copying. Further, we show that decoder-only Transformer language models can lose sensitivity to specific tokens in the input, which relates to the well-known phenomenon of over-squashing in graph neural networks. We provide empirical evidence supporting our claims on contemporary LLMs. Our theory points to simple solutions towards ameliorating these issues. Federico Barbero, Andrea Banino, Steven Kapturowski, Dharshan Kumaran, João G. M. Araújo, Alex Vitvitskyi, Razvan Pascanu, Petar Velickovic |
NeurIPS | 3 |
| 2023 | Human-level Atari 200x faster
Steven Kapturowski, Victor Campos 0001, Ray Jiang, Nemanja Rakicevic, Hado van Hasselt, Charles Blundell, Adrià Puigdomènech Badia |
ICLR | 1 |
| 2021 | Revisiting Peng's Q(λ) for Modern Reinforcement LearningabstractOff-policy multi-step reinforcement learning algorithms consist of conservative and non-conservative algorithms: the former actively cut traces, whereas the latter do not. Recently, Munos et al. (2016) proved the convergence of conservative algorithms to an optimal Q-function. In contrast, non-conservative algorithms are thought to be unsafe and have a limited or no theoretical guarantee. Nonetheless, recent studies have shown that non-conservative algorithms empirically outperform conservative ones. Motivated by the empirical results and the lack of theory, we carry out theoretical analyses of Peng’s Q($\lambda$), a representative example of non-conservative algorithms. We prove that \emph{it also converges to an optimal policy} provided that the behavior policy slowly tracks a greedy policy in a way similar to conservative policy iteration. Such a result has been conjectured to be true but has not been proven. We also experiment with Peng’s Q($\lambda$) in complex continuous control tasks, confirming that Peng’s Q($\lambda$) often outperforms conservative algorithms despite its simplicity. These results indicate that Peng’s Q($\lambda$), which was thought to be unsafe, is a theoretically-sound and practically effective algorithm. Tadashi Kozuno, Yunhao Tang, Mark Rowland 0001, Rémi Munos, Steven Kapturowski, Will Dabney, Michal Valko, David Abel |
ICML | 5 |
| 2020 | Never Give Up: Learning Directed Exploration Strategies
Adrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Guo, Bilal Piot, Steven Kapturowski, Olivier Tieleman, Martín Arjovsky, Alexander Pritzel, Andrew Bolt, Charles Blundell |
ICLR | 6 |
| 2020 | Making Efficient Use of Demonstrations to Solve Hard Exploration Problems
Caglar Gulcehre, Tom Le Paine, Bobak Shahriari, Misha Denil, Matt Hoffman 0001, Hubert Soyer, Richard Tanburn, Steven Kapturowski, Neil C. Rabinowitz, Duncan Williams, Gabriel Barth-Maron, Ziyu Wang 0001, Nando de Freitas |
ICLR | 8 |
| 2020 | Agent57: Outperforming the Atari Human BenchmarkabstractAtari games have been a long-standing benchmark in the reinforcement learning (RL) community for the past decade. This benchmark was proposed to test general competency of RL algorithms. Previous work has achieved good average performance by doing outstandingly well on many games of the set, but very poorly in several of the most challenging games. We propose Agent57, the first deep RL agent that outperforms the standard human benchmark on all 57 Atari games. To achieve this result, we train a neural network which parameterizes a family of policies ranging from very exploratory to purely exploitative. We propose an adaptive mechanism to choose which policy to prioritize throughout the training process. Additionally, we utilize a novel parameterization of the architecture that allows for more consistent and stable learning. Adrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Guo, Charles Blundell |
ICML | 3 |
| 2020 | Value-driven Hindsight ModellingabstractValue estimation is a critical component of the reinforcement learning (RL) paradigm. The question of how to effectively learn value predictors from data is one of the major problems studied by the RL community, and different approaches exploit structure in the problem domain in different ways. Model learning can make use of the rich transition structure present in sequences of observations, but this approach is usually not sensitive to the reward function. In contrast, model-free methods directly leverage the quantity of interest from the future, but receive a potentially weak scalar signal (an estimate of the return). We develop an approach for representation learning in RL that sits in between these two extremes: we propose to learn what to model in a way that can directly help value prediction. To this end, we determine which features of the future trajectory provide useful information to predict the associated return. This provides tractable prediction targets that are directly relevant for a task, and can thus accelerate learning the value function. The idea can be understood as reasoning, in hindsight, about which aspects of the future observations could help past value prediction. We show how this can help dramatically even in simple policy evaluation settings. We then test our approach at scale in challenging domains, including on 57 Atari 2600 games. Arthur Guez, Fabio Viola, Theophane Weber, Lars Buesing, Steven Kapturowski, Doina Precup, David Silver 0001, Nicolas Heess |
NeurIPS | 5 |
| 2019 | Recurrent Experience Replay in Distributed Reinforcement Learning
Steven Kapturowski, Georg Ostrovski, John Quan, Rémi Munos, Will Dabney |
ICLR (Poster) | 1 |