VLDB 2026 Research / reviewers in the wild / expert
Steven Hansen 0001
dblp:80/9799 · also Steven Stenberg Hansen
· DBLP profile ↗
13ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Reinforcement learning · 99% Multi-agent systems · 1% | |
| Theoretical computer science
1 paper |
Information theory · 100% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
exploration |
1.1 | 2 | 2022 | Learning more skills through optimistic exploration · ICLR 2022 Entropic Desired Dynamics for Intrinsic Control · NeurIPS 2021 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
skill learning |
1.1 | 2 | 2022 | Learning more skills through optimistic exploration · ICLR 2022 Relative Variational Intrinsic Control · AAAI 2021 |
Machine learning › Reinforcement learning › exploration › intrinsically motivated reinforcement learning
empowerment |
0.9 | 1 | 2025 | Plasticity as the Mirror of Empowerment · NeurIPS 2025 |
Information theory › information measures › multiterminal information measures
directed information |
0.9 | 1 | 2025 | Plasticity as the Mirror of Empowerment · NeurIPS 2025 |
Machine learning › Reinforcement learning › meta-reinforcement learning
in-context reinforcement learning |
0.7 | 1 | 2023 | In-context Reinforcement Learning with Algorithm Distillation · ICLR 2023 |
Machine learning › Reinforcement learning › exploration
optimistic exploration |
0.6 | 1 | 2022 | Learning more skills through optimistic exploration · ICLR 2022 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning › skill learning
skill discovery |
0.6 | 1 | 2022 | Learning more skills through optimistic exploration · ICLR 2022 |
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.5 | 1 | 2021 | Relative Variational Intrinsic Control · AAAI 2021 |
Machine learning › Reinforcement learning
task inference |
0.4 | 1 | 2020 | Fast Task Inference with Variational Intrinsic Successor Features · ICLR 2020 |
Machine learning › Reinforcement learning
generalization in reinforcement learning |
0.4 | 1 | 2019 | Generalization of Reinforcement Learners with Working and Episodic Memory · NeurIPS 2019 |
Machine learning › Reinforcement learning › partially observable reinforcement learning
memory-based reinforcement learning |
0.4 | 1 | 2019 | Generalization of Reinforcement Learners with Working and Episodic Memory · NeurIPS 2019 |
Machine learning › Reinforcement learning
reward learning |
0.4 | 1 | 2019 | Unsupervised Control Through Non-Parametric Discriminative Rewards · ICLR (Poster) 2019 |
Machine learning › Reinforcement learning
unsupervised reinforcement learning |
0.4 | 1 | 2019 | Unsupervised Control Through Non-Parametric Discriminative Rewards · ICLR (Poster) 2019 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.3 | 1 | 2018 | Fast deep reinforcement learning using online adjustments from the past · NeurIPS 2018 |
Machine learning › Reinforcement learning › off-policy reinforcement learning
experience replay |
0.3 | 1 | 2018 | Fast deep reinforcement learning using online adjustments from the past · NeurIPS 2018 |
Machine learning › Reinforcement learning
value function estimation |
0.3 | 1 | 2018 | Fast deep reinforcement learning using online adjustments from the past · NeurIPS 2018 |
Knowledge, reasoning and agents › Multi-agent systems
agent architecture |
0.1 | 1 | 2019 | Generalization of Reinforcement Learners with Working and Episodic Memory · NeurIPS 2019 |
Methods — techniques the papers use, named apart from their topics
information-theoretic measures · 0.9information-theoretic measure · 0.9optimistic exploration · 0.6mutual information · 0.5latent dynamics · 0.5entropic desired dynamics · 0.5variational inference · 0.4successor features · 0.4non-parametric discriminative rewards · 0.4episodic memory · 0.4ablation study · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Plasticity as the Mirror of EmpowermentabstractAgents are minimally entities that are influenced by their past observations and act to influence future observations. This latter capacity is captured by empowerment, which has served as a vital framing concept across artificial intelligence and cognitive science. This former capacity, however, is equally foundational: In what ways, and to what extent, can an agent be influenced by what it observes? In this paper, we ground this concept in a universal agent-centric measure that we refer to as plasticity, and reveal a fundamental connection to empowerment. Following a set of desiderata on a suitable definition, we define plasticity using a new information-theoretic quantity we call the generalized directed information. We show that this new quantity strictly generalizes the directed information introduced by Massey (1990) while preserving all of its desirable properties. Under this definition, we find that plasticity is well thought of as the mirror of empowerment: The two concepts are defined using the same measure, with only the direction of influence reversed. Our main result establishes a tension between the plasticity and empowerment of an agent, suggesting that agent design needs to be mindful of both characteristics. We explore the implications of these findings, and suggest that plasticity, empowerment, and their relationship are essential to understanding agency. David Abel, Michael H. Bowling, André Barreto 0001, Will Dabney, Steven Hansen 0001, Anna Harutyunyan, Khimya Khetarpal, Clare Lyle, Razvan Pascanu, Georgios Piliouras, Doina Precup, Jonathan Richens, Mark Rowland 0001, Tom Schaul, Satinder Singh 0001 |
NeurIPS | 6 |
| 2023 | In-context Reinforcement Learning with Algorithm Distillation
Michael Laskin, Luyu Wang, Junhyuk Oh, Emilio Parisotto, Stephen Spencer, Richie Steigerwald, DJ Strouse, Steven Hansen 0001, Angelos Filos, Ethan Brooks, Maxime Gazeau, Himanshu Sahni, Satinder Singh 0001, Volodymyr Mnih |
ICLR | 8 |
| 2022 | Learning more skills through optimistic exploration
DJ Strouse, Kate Baumli, David Warde-Farley, Volodymyr Mnih, Steven Hansen 0001 |
ICLR | 5 |
| 2021 | Relative Variational Intrinsic ControlabstractIn the absence of external rewards, agents can still learn useful behaviors by identifying and mastering a set of diverse skills within their environment. Existing skill learning methods use mutual information objectives to incentivize each skill to be diverse and distinguishable from the rest. However, if care is not taken to constrain the ways in which the skills are diverse, trivially diverse skill sets can arise. To ensure useful skill diversity, we propose a novel skill learning objective, Relative Variational Intrinsic Control (RVIC), which incentivizes learning skills that are distinguishable in how they change the agent's relationship to its environment. The resulting set of skills tiles the space of affordances available to the agent. We qualitatively analyze skill behaviors on multiple environments and show how RVIC skills are more useful than skills discovered by existing methods in hierarchical reinforcement learning. Kate Baumli, David Warde-Farley, Steven Hansen 0001, Volodymyr Mnih |
AAAI | 3 |
| 2021 | Entropic Desired Dynamics for Intrinsic ControlabstractAn agent might be said, informally, to have mastery of its environment when it has maximised the effective number of states it can reliably reach. In practice, this often means maximizing the number of latent codes that can be discriminated from future states under some short time horizon (e.g. \cite{eysenbach2018diversity}). By situating these latent codes in a globally consistent coordinate system, we show that agents can reliably reach more states in the long term while still optimizing a local objective. A simple instantiation of this idea, \textbf{E}ntropic \textbf{D}esired \textbf{D}ynamics for \textbf{I}ntrinsic \textbf{C}on\textbf{T}rol (EDDICT), assumes fixed additive latent dynamics, which results in tractable learning and an interpretable latent space. Compared to prior methods, EDDICT's globally consistent codes allow it to be far more exploratory, as demonstrated by improved state coverage and increased unsupervised performance on hard exploration games such as Montezuma's Revenge. Steven Hansen 0001, Guillaume Desjardins, Kate Baumli, David Warde-Farley, Nicolas Heess, Simon Osindero, Volodymyr Mnih |
NeurIPS | 1 |
| 2020 | Fast Task Inference with Variational Intrinsic Successor Features
Steven Hansen 0001, Will Dabney, André Barreto 0001, David Warde-Farley, Tom Van de Wiele, Volodymyr Mnih |
ICLR | 1 |
| 2019 | Unsupervised Control Through Non-Parametric Discriminative Rewards
David Warde-Farley, Tom Van de Wiele, Tejas D. Kulkarni, Catalin Ionescu, Steven Hansen 0001, Volodymyr Mnih |
ICLR (Poster) | 5 |
| 2019 | Generalization of Reinforcement Learners with Working and Episodic MemoryabstractMemory is an important aspect of intelligence and plays a role in many deep reinforcement learning models. However, little progress has been made in understanding when specific memory systems help more than others and how well they generalize. The field also has yet to see a prevalent consistent and rigorous approach for evaluating agent performance on holdout data. In this paper, we aim to develop a comprehensive methodology to test different kinds of memory in an agent and assess how well the agent can apply what it learns in training to a holdout set that differs from the training set along dimensions that we suggest are relevant for evaluating memory-specific generalization. To that end, we first construct a diverse set of memory tasks that allow us to evaluate test-time generalization across multiple dimensions. Second, we develop and perform multiple ablations on an agent architecture that combines multiple memory systems, observe its baseline models, and investigate its performance against the task suite. Meire Fortunato, Melissa Tan, Ryan Faulkner 0001, Steven Hansen 0001, Adrià Puigdomènech Badia, Gavin Buttimore, Charlie Deck, Joel Z. Leibo, Charles Blundell |
NeurIPS | 4 |
| 2018 | Fast deep reinforcement learning using online adjustments from the pastabstractWe propose Ephemeral Value Adjusments (EVA): a means of allowing deep reinforcement learning agents to rapidly adapt to experience in their replay buffer. EVA shifts the value predicted by a neural network with an estimate of the value function found by prioritised sweeping over experience tuples from the replay buffer near the current state. EVA combines a number of recent ideas around combining episodic memory-like structures into reinforcement learning agents: slot-based storage, content-based retrieval, and memory-based planning. We show that EVA is performant on a demonstration task and Atari games. Steven Hansen 0001, Alexander Pritzel, Pablo Sprechmann, André Barreto 0001, Charles Blundell |
NeurIPS | 1 |
| 2017 | Neural responses decrease while performance increases with practice: A neural network model
Milena Rabovsky, Steven Hansen 0001, James L. McClelland |
CogSci | 2 |
| 2016 | Tutorial Workshop on Contemporary Deep Neural Network Models
James L. McClelland, Steven Hansen 0001, Andrew M. Saxe |
CogSci | 2 |
| 2016 | N400 amplitudes reflect change in a probabilistic representation of meaning: Evidence from a connectionist model
Milena Rabovsky, Steven Hansen 0001, James L. McClelland |
CogSci | 2 |
| 2014 | Two Plus Three Is Five: Discovering Efficient Addition Strategies without Metacognition
Steven Hansen 0001, Cameron R. L. McKenzie, James L. McClelland |
CogSci | 1 |