Steven Hansen 0001

dblp:80/9799 · also Steven Stenberg Hansen · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Reinforcement learning · 99% Multi-agent systems · 1%
Theoretical computer science
1 paper
Information theory · 100%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
exploration
1.122022
Learning more skills through optimistic exploration · ICLR 2022
Entropic Desired Dynamics for Intrinsic Control · NeurIPS 2021
Machine learning › Reinforcement learning › hierarchical reinforcement learning
skill learning
1.122022
Learning more skills through optimistic exploration · ICLR 2022
Relative Variational Intrinsic Control · AAAI 2021
Machine learning › Reinforcement learning › exploration › intrinsically motivated reinforcement learning
empowerment
0.912025
Plasticity as the Mirror of Empowerment · NeurIPS 2025
Information theory › information measures › multiterminal information measures
directed information
0.912025
Plasticity as the Mirror of Empowerment · NeurIPS 2025
Machine learning › Reinforcement learning › meta-reinforcement learning
in-context reinforcement learning
0.712023
In-context Reinforcement Learning with Algorithm Distillation · ICLR 2023
Machine learning › Reinforcement learning › exploration
optimistic exploration
0.612022
Learning more skills through optimistic exploration · ICLR 2022
Machine learning › Reinforcement learning › hierarchical reinforcement learning › skill learning
skill discovery
0.612022
Learning more skills through optimistic exploration · ICLR 2022
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.512021
Relative Variational Intrinsic Control · AAAI 2021
Machine learning › Reinforcement learning
task inference
0.412020
Fast Task Inference with Variational Intrinsic Successor Features · ICLR 2020
Machine learning › Reinforcement learning
generalization in reinforcement learning
0.412019
Generalization of Reinforcement Learners with Working and Episodic Memory · NeurIPS 2019
Machine learning › Reinforcement learning › partially observable reinforcement learning
memory-based reinforcement learning
0.412019
Generalization of Reinforcement Learners with Working and Episodic Memory · NeurIPS 2019
Machine learning › Reinforcement learning
reward learning
0.412019
Unsupervised Control Through Non-Parametric Discriminative Rewards · ICLR (Poster) 2019
Machine learning › Reinforcement learning
unsupervised reinforcement learning
0.412019
Unsupervised Control Through Non-Parametric Discriminative Rewards · ICLR (Poster) 2019
Machine learning › Reinforcement learning
deep reinforcement learning
0.312018
Fast deep reinforcement learning using online adjustments from the past · NeurIPS 2018
Machine learning › Reinforcement learning › off-policy reinforcement learning
experience replay
0.312018
Fast deep reinforcement learning using online adjustments from the past · NeurIPS 2018
Machine learning › Reinforcement learning
value function estimation
0.312018
Fast deep reinforcement learning using online adjustments from the past · NeurIPS 2018
Knowledge, reasoning and agents › Multi-agent systems
agent architecture
0.112019
Generalization of Reinforcement Learners with Working and Episodic Memory · NeurIPS 2019

Methods — techniques the papers use, named apart from their topics

information-theoretic measures · 0.9information-theoretic measure · 0.9optimistic exploration · 0.6mutual information · 0.5latent dynamics · 0.5entropic desired dynamics · 0.5variational inference · 0.4successor features · 0.4non-parametric discriminative rewards · 0.4episodic memory · 0.4ablation study · 0.4
YearPublicationVenuePosition
2025 Plasticity as the Mirror of Empowerment
abstract
Agents are minimally entities that are influenced by their past observations and act to influence future observations. This latter capacity is captured by empowerment, which has served as a vital framing concept across artificial intelligence and cognitive science. This former capacity, however, is equally foundational: In what ways, and to what extent, can an agent be influenced by what it observes? In this paper, we ground this concept in a universal agent-centric measure that we refer to as plasticity, and reveal a fundamental connection to empowerment. Following a set of desiderata on a suitable definition, we define plasticity using a new information-theoretic quantity we call the generalized directed information. We show that this new quantity strictly generalizes the directed information introduced by Massey (1990) while preserving all of its desirable properties. Under this definition, we find that plasticity is well thought of as the mirror of empowerment: The two concepts are defined using the same measure, with only the direction of influence reversed. Our main result establishes a tension between the plasticity and empowerment of an agent, suggesting that agent design needs to be mindful of both characteristics. We explore the implications of these findings, and suggest that plasticity, empowerment, and their relationship are essential to understanding agency.
David Abel, Michael H. Bowling, André Barreto 0001, Will Dabney, Steven Hansen 0001, Anna Harutyunyan, Khimya Khetarpal, Clare Lyle, Razvan Pascanu, Georgios Piliouras, Doina Precup, Jonathan Richens, Mark Rowland 0001, Tom Schaul, Satinder Singh 0001
NeurIPS6
2023 In-context Reinforcement Learning with Algorithm Distillation
Michael Laskin, Luyu Wang, Junhyuk Oh, Emilio Parisotto, Stephen Spencer, Richie Steigerwald, DJ Strouse, Steven Hansen 0001, Angelos Filos, Ethan Brooks, Maxime Gazeau, Himanshu Sahni, Satinder Singh 0001, Volodymyr Mnih
ICLR8
2022 Learning more skills through optimistic exploration
DJ Strouse, Kate Baumli, David Warde-Farley, Volodymyr Mnih, Steven Hansen 0001
ICLR5
2021 Relative Variational Intrinsic Control
abstract
In the absence of external rewards, agents can still learn useful behaviors by identifying and mastering a set of diverse skills within their environment. Existing skill learning methods use mutual information objectives to incentivize each skill to be diverse and distinguishable from the rest. However, if care is not taken to constrain the ways in which the skills are diverse, trivially diverse skill sets can arise. To ensure useful skill diversity, we propose a novel skill learning objective, Relative Variational Intrinsic Control (RVIC), which incentivizes learning skills that are distinguishable in how they change the agent's relationship to its environment. The resulting set of skills tiles the space of affordances available to the agent. We qualitatively analyze skill behaviors on multiple environments and show how RVIC skills are more useful than skills discovered by existing methods in hierarchical reinforcement learning.
Kate Baumli, David Warde-Farley, Steven Hansen 0001, Volodymyr Mnih
AAAI3
2021 Entropic Desired Dynamics for Intrinsic Control
abstract
An agent might be said, informally, to have mastery of its environment when it has maximised the effective number of states it can reliably reach. In practice, this often means maximizing the number of latent codes that can be discriminated from future states under some short time horizon (e.g. \cite{eysenbach2018diversity}). By situating these latent codes in a globally consistent coordinate system, we show that agents can reliably reach more states in the long term while still optimizing a local objective. A simple instantiation of this idea, \textbf{E}ntropic \textbf{D}esired \textbf{D}ynamics for \textbf{I}ntrinsic \textbf{C}on\textbf{T}rol (EDDICT), assumes fixed additive latent dynamics, which results in tractable learning and an interpretable latent space. Compared to prior methods, EDDICT's globally consistent codes allow it to be far more exploratory, as demonstrated by improved state coverage and increased unsupervised performance on hard exploration games such as Montezuma's Revenge.
Steven Hansen 0001, Guillaume Desjardins, Kate Baumli, David Warde-Farley, Nicolas Heess, Simon Osindero, Volodymyr Mnih
NeurIPS1
2020 Fast Task Inference with Variational Intrinsic Successor Features
Steven Hansen 0001, Will Dabney, André Barreto 0001, David Warde-Farley, Tom Van de Wiele, Volodymyr Mnih
ICLR1
2019 Unsupervised Control Through Non-Parametric Discriminative Rewards
David Warde-Farley, Tom Van de Wiele, Tejas D. Kulkarni, Catalin Ionescu, Steven Hansen 0001, Volodymyr Mnih
ICLR (Poster)5
2019 Generalization of Reinforcement Learners with Working and Episodic Memory
abstract
Memory is an important aspect of intelligence and plays a role in many deep reinforcement learning models. However, little progress has been made in understanding when specific memory systems help more than others and how well they generalize. The field also has yet to see a prevalent consistent and rigorous approach for evaluating agent performance on holdout data. In this paper, we aim to develop a comprehensive methodology to test different kinds of memory in an agent and assess how well the agent can apply what it learns in training to a holdout set that differs from the training set along dimensions that we suggest are relevant for evaluating memory-specific generalization. To that end, we first construct a diverse set of memory tasks that allow us to evaluate test-time generalization across multiple dimensions. Second, we develop and perform multiple ablations on an agent architecture that combines multiple memory systems, observe its baseline models, and investigate its performance against the task suite.
Meire Fortunato, Melissa Tan, Ryan Faulkner 0001, Steven Hansen 0001, Adrià Puigdomènech Badia, Gavin Buttimore, Charlie Deck, Joel Z. Leibo, Charles Blundell
NeurIPS4
2018 Fast deep reinforcement learning using online adjustments from the past
abstract
We propose Ephemeral Value Adjusments (EVA): a means of allowing deep reinforcement learning agents to rapidly adapt to experience in their replay buffer. EVA shifts the value predicted by a neural network with an estimate of the value function found by prioritised sweeping over experience tuples from the replay buffer near the current state. EVA combines a number of recent ideas around combining episodic memory-like structures into reinforcement learning agents: slot-based storage, content-based retrieval, and memory-based planning. We show that EVA is performant on a demonstration task and Atari games.
Steven Hansen 0001, Alexander Pritzel, Pablo Sprechmann, André Barreto 0001, Charles Blundell
NeurIPS1
2017 Neural responses decrease while performance increases with practice: A neural network model
Milena Rabovsky, Steven Hansen 0001, James L. McClelland
CogSci2
2016 Tutorial Workshop on Contemporary Deep Neural Network Models
James L. McClelland, Steven Hansen 0001, Andrew M. Saxe
CogSci2
2016 N400 amplitudes reflect change in a probabilistic representation of meaning: Evidence from a connectionist model
Milena Rabovsky, Steven Hansen 0001, James L. McClelland
CogSci2
2014 Two Plus Three Is Five: Discovering Efficient Addition Strategies without Metacognition
Steven Hansen 0001, Cameron R. L. McKenzie, James L. McClelland
CogSci1