Shane Legg

dblp:36/5739 · DBLP profile ↗
← Back
21ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
15 papers
Reinforcement learning · 65% Trustworthy machine learning · 10% Transfer learning and domain adaptation · 7%
Theoretical computer science
2 papers
Algorithmic game theory and mechanism design · 82% Automata and formal languages · 18%

Topics — the 30 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
reward learning
1.542020
Pitfalls of Learning a Reward Function Online · IJCAI 2020
Learning Human Objectives by Evaluating Hypothetical Behavior · ICML 2020
Reward learning from human preferences and demonstrations in Atari · NeurIPS 2018
Machine learning › Reinforcement learning
reward design
0.922021
Quantifying Differences in Reward Functions · ICLR 2021
Avoiding Side Effects By Considering Future Tasks · NeurIPS 2020
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning
0.722020
Pitfalls of Learning a Reward Function Online · IJCAI 2020
Reinforcement Learning with a Corrupted Reward Channel · IJCAI 2017
Machine learning › Deep learning architectures and training
neural network expressivity
0.712023
Neural Networks and the Chomsky Hierarchy · ICLR 2023
Machine learning › Reinforcement learning › preference learning
human preference learning
0.622018
Reward learning from human preferences and demonstrations in Atari · NeurIPS 2018
Deep Reinforcement Learning from Human Preferences · NIPS 2017
Machine learning › Trustworthy machine learning
fairness
0.512021
Agent Incentives: A Causal Perspective · AAAI 2021
Machine learning › Reinforcement learning
reward function
0.512021
Quantifying Differences in Reward Functions · ICLR 2021
Machine learning › Transfer learning and domain adaptation › meta-learning
memory-based meta-learning
0.412020
Meta-trained agents implement Bayes-optimal agents · NeurIPS 2020
Machine learning › Transfer learning and domain adaptation
meta-learning
0.412020
Meta-trained agents implement Bayes-optimal agents · NeurIPS 2020
Machine learning › Reinforcement learning › reward learning
reward model training
0.412020
Learning Human Objectives by Evaluating Hypothetical Behavior · ICML 2020
Machine learning › Reinforcement learning
safe reinforcement learning
0.412020
Learning Human Objectives by Evaluating Hypothetical Behavior · ICML 2020
Machine learning › Reinforcement learning › safe reinforcement learning
side effect avoidance
0.412020
Avoiding Side Effects By Considering Future Tasks · NeurIPS 2020
Machine learning › Reinforcement learning › large-scale reinforcement learning › distributed reinforcement learning
actor-learner architecture
0.312018
IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures · ICML 2018
Machine learning › Reinforcement learning › large-scale reinforcement learning
distributed reinforcement learning
0.312018
IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures · ICML 2018
Machine learning › Reinforcement learning
exploration
0.312018
Noisy Networks For Exploration · ICLR (Poster) 2018
Machine learning › Reinforcement learning
imitation learning
0.312018
Reward learning from human preferences and demonstrations in Atari · NeurIPS 2018
Robotics › Robot manipulation
learning from demonstration
0.312018
Reward learning from human preferences and demonstrations in Atari · NeurIPS 2018
Machine learning › Reinforcement learning
human feedback
0.312017
Deep Reinforcement Learning from Human Preferences · NIPS 2017
Machine learning › Reinforcement learning › reinforcement learning from human feedback
preference-based reinforcement learning
0.312017
Deep Reinforcement Learning from Human Preferences · NIPS 2017
Automata and formal languages › formal grammars
chomsky hierarchy
0.212023
Neural Networks and the Chomsky Hierarchy · ICLR 2023
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.112021
Agent Incentives: A Causal Perspective · AAAI 2021
Natural language and speech › Language models and text generation › alignment
agent alignment
0.112020
Pitfalls of Learning a Reward Function Online · IJCAI 2020
Machine learning › Reinforcement learning
bandit
0.112020
Meta-trained agents implement Bayes-optimal agents · NeurIPS 2020
Knowledge, reasoning and agents › Multi-agent systems
grid environments
0.112020
Avoiding Side Effects By Considering Future Tasks · NeurIPS 2020
Machine learning › Efficient and distributed learning
distributed training
0.112018
IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures · ICML 2018
Machine learning › Reinforcement learning
actor-critic methods
0.112017
Deep Reinforcement Learning from Human Preferences · NIPS 2017
Machine learning › Reinforcement learning
temporal difference learning
0.112007
Temporal Difference Updating without a Learning Rate · NIPS 2007
Machine learning › Reinforcement learning
value-based reinforcement learning
0.112007
Temporal Difference Updating without a Learning Rate · NIPS 2007
Machine learning › Reinforcement learning › temporal difference learning
eligibility traces
0.012007
Temporal Difference Updating without a Learning Rate · NIPS 2007
Machine learning › Reinforcement learning › temporal difference learning › eligibility traces
TD(lambda)
0.012007
Temporal Difference Updating without a Learning Rate · NIPS 2007

Methods — techniques the papers use, named apart from their topics

capability taxonomy · 1.5benchmarking · 1.5statistical comparison · 0.5reward modeling · 0.5graphical criteria · 0.5causal influence diagrams · 0.5trajectory optimization · 0.4forward dynamics model · 0.4baseline policy · 0.4active learning · 0.4
YearPublicationVenuePosition
2025 Incentives for responsiveness, instrumental control and impact
Ryan Carey, Eric D. Langlois 0002, Chris van Merwijk, Shane Legg, Tom Everitt
Artif. Intell.4
2024 Position: Levels of AGI for Operationalizing Progress on the Path to AGI
abstract
We propose a framework for classifying the capabilities and behavior of Artificial General Intelligence (AGI) models and their precursors. This framework introduces levels of AGI performance, generality, and autonomy, providing a common language to compare models, assess risks, and measure progress along the path to AGI. To develop our framework, we analyze existing definitions of AGI, and distill six principles that a useful ontology for AGI should satisfy. With these principles in mind, we propose “Levels of AGI” based on depth (performance) and breadth (generality) of capabilities, and reflect on how current systems fit into this ontology. We discuss the challenging requirements for future benchmarks that quantify the behavior and capabilities of AGI models against these levels. Finally, we discuss how these levels of AGI interact with deployment considerations such as autonomy and risk, and emphasize the importance of carefully selecting Human-AI Interaction paradigms for responsible and safe deployment of highly capable AI systems.
Meredith Ringel Morris, Jascha Sohl-Dickstein, Noah Fiedel, Tris Warkentin, Allan Dafoe, Aleksandra Faust, Clément Farabet, Shane Legg
ICML8
2023 Neural Networks and the Chomsky Hierarchy
Grégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Chris Cundy, Marcus Hutter, Shane Legg, Joel Veness, Pedro A. Ortega
ICLR9
2021 Agent Incentives: A Causal Perspective
abstract
We present a framework for analysing agent incentives using causal influence diagrams. We establish that a well-known criterion for value of information is complete. We propose a new graphical criterion for value of control, establishing its soundness and completeness. We also introduce two new concepts for incentive analysis: response incentives indicate which changes in the environment affect an optimal decision, while instrumental control incentives establish whether an agent can influence its utility via a variable X. For both new concepts, we provide sound and complete graphical criteria. We show by example how these results can help with evaluating the safety and fairness of an AI system
Tom Everitt, Ryan Carey, Eric D. Langlois 0002, Pedro A. Ortega, Shane Legg
AAAI5
2021 Quantifying Differences in Reward Functions
Adam Gleave, Michael Dennis 0001, Shane Legg, Stuart Russell 0001, Jan Leike
ICLR3
2020 Learning Human Objectives by Evaluating Hypothetical Behavior
abstract
We seek to align agent behavior with a user’s objectives in a reinforcement learning setting with unknown dynamics, an unknown reward function, and unknown unsafe states. The user knows the rewards and unsafe states, but querying the user is expensive. We propose an algorithm that safely and efficiently learns a model of the user’s reward function by posing ’what if?’ questions about hypothetical agent behavior. We start with a generative model of initial states and a forward dynamics model trained on off-policy data. Our method uses these models to synthesize hypothetical behaviors, asks the user to label the behaviors with rewards, and trains a neural network to predict the rewards. The key idea is to actively synthesize the hypothetical behaviors from scratch by maximizing tractable proxies for the value of information, without interacting with the environment. We call this method reward query synthesis via trajectory optimization (ReQueST). We evaluate ReQueST with simulated users on a state-based 2D navigation task and the image-based Car Racing video game. The results show that ReQueST significantly outperforms prior methods in learning reward models that transfer to new environments with different initial state distributions. Moreover, ReQueST safely trains the reward model to detect unsafe states, and corrects reward hacking before deploying the agent.
Siddharth Reddy, Anca D. Dragan, Sergey Levine, Shane Legg, Jan Leike
ICML4
2020 Pitfalls of Learning a Reward Function Online
abstract
In some agent designs like inverse reinforcement learning an agent needs to learn its own reward function. Learning the reward function and optimising for it are typically two different processes, usually performed at different stages. We consider a continual (``one life'') learning approach where the agent both learns the reward function and optimises for it at the same time. We show that this comes with a number of pitfalls, such as deliberately manipulating the learning process in one direction, refusing to learn, ``learning'' facts already known to the agent, and making decisions that are strictly dominated (for all relevant reward functions). We formally introduce two desirable properties: the first is `unriggability', which prevents the agent from steering the learning process in the direction of a reward function that is easier to optimise. The second is `uninfluenceability', whereby the reward-function learning process operates by learning facts about the environment. We show that an uninfluenceable process is automatically unriggable, and if the set of possible environments is sufficiently large, the converse is true too.
Stuart Armstrong, Jan Leike, Laurent Orseau, Shane Legg
IJCAI4
2020 Avoiding Side Effects By Considering Future Tasks
abstract
Designing reward functions is difficult: the designer has to specify what to do (what it means to complete the task) as well as what not to do (side effects that should be avoided while completing the task). To alleviate the burden on the reward designer, we propose an algorithm to automatically generate an auxiliary reward function that penalizes side effects. This auxiliary objective rewards the ability to complete possible future tasks, which decreases if the agent causes side effects during the current task. The future task reward can also give the agent an incentive to interfere with events in the environment that make future tasks less achievable, such as irreversible actions by other agents. To avoid this interference incentive, we introduce a baseline policy that represents a default course of action (such as doing nothing), and use it to filter out future tasks that are not achievable by default. We formally define interference incentives and show that the future task approach with a baseline policy avoids these incentives in the deterministic case. Using gridworld environments that test for side effects and interference, we show that our method avoids interference and is more effective for avoiding side effects than the common approach of penalizing irreversible actions.
Victoria Krakovna, Laurent Orseau, Richard Ngo, Miljan Martic, Shane Legg
NeurIPS5
2020 Meta-trained agents implement Bayes-optimal agents
abstract
Memory-based meta-learning is a powerful technique to build agents that adapt fast to any task within a target distribution. A previous theoretical study has argued that this remarkable performance is because the meta-training protocol incentivises agents to behave Bayes-optimally. We empirically investigate this claim on a number of prediction and bandit tasks. Inspired by ideas from theoretical computer science, we show that meta-learned and Bayes-optimal agents not only behave alike, but they even share a similar computational structure, in the sense that one agent system can approximately simulate the other. Furthermore, we show that Bayes-optimal agents are fixed points of the meta-learning dynamics. Our results suggest that memory-based meta-learning is a general technique for numerically approximating Bayes-optimal agents; that is, even for task distributions for which we currently don't possess tractable models.
Vladimir Mikulik, Grégoire Delétang, Thomas McGrath 0001, Tim Genewein, Miljan Martic, Shane Legg, Pedro A. Ortega
NeurIPS6
2018 Noisy Networks For Exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Matteo Hessel, Ian Osband, Alex Graves, Volodymyr Mnih, Rémi Munos, Demis Hassabis, Olivier Pietquin, Charles Blundell, Shane Legg
ICLR (Poster)13
2018 IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
abstract
In this work we aim to solve a large collection of tasks using a single reinforcement learning agent with a single set of parameters. A key challenge is to handle the increased amount of data and extended training time. We have developed a new distributed agent IMPALA (Importance Weighted Actor-Learner Architecture) that not only uses resources more efficiently in single-machine training but also scales to thousands of machines without sacrificing data efficiency or resource utilisation. We achieve stable learning at high throughput by combining decoupled acting and learning with a novel off-policy correction method called V-trace. We demonstrate the effectiveness of IMPALA for multi-task reinforcement learning on DMLab-30 (a set of 30 tasks from the DeepMind Lab environment (Beattie et al., 2016)) and Atari57 (all available Atari games in Arcade Learning Environment (Bellemare et al., 2013a)). Our results show that IMPALA is able to achieve better performance than previous agents with less data, and crucially exhibits positive transfer between tasks as a result of its multi-task approach.
Lasse Espeholt, Hubert Soyer, Rémi Munos, Karen Simonyan, Volodymyr Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, Koray Kavukcuoglu
ICML11
2018 Reward learning from human preferences and demonstrations in Atari
abstract
To solve complex real-world problems with reinforcement learning, we cannot rely on manually specified reward functions. Instead, we need humans to communicate an objective to the agent directly. In this work, we combine two approaches to this problem: learning from expert demonstrations and learning from trajectory preferences. We use both to train a deep neural network to model the reward function and use its predicted reward to train an DQN-based deep reinforcement learning agent on 9 Atari games. Our approach beats the imitation learning baseline in 7 games and achieves strictly superhuman performance on 2 games. Additionally, we investigate the fit of the reward model, present some reward hacking problems, and study the effects of noise in the human labels.
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, Dario Amodei
NeurIPS5
2017 Soft-Bayes: Prod for Mixtures of Experts with Log-Loss
abstract
We consider prediction with expert advice under the log-loss with the goal of deriving efficient and robust algorithms. We argue that existing algorithms such as exponentiated gradient, online gradient descent and online Newton step do not adequately satisfy both requirements. Our main contribution is an analysis of the Prod algorithm that is robust to any data sequence and runs in linear time relative to the number of experts in each round. Despite the unbounded nature of the log-loss, we derive a bound that is independent of the largest loss and of the largest gradient, and depends only on the number of experts and the time horizon. Furthermore we give a Bayesian interpretation of Prod and adapt the algorithm to derive a tracking regret.
Laurent Orseau, Tor Lattimore, Shane Legg
ALT3
2017 Reinforcement Learning with a Corrupted Reward Channel
abstract
No real-world reward function is perfect. Sensory errors and software bugs may result in agents getting higher (or lower) rewards than they should. For example, a reinforcement learning agent may prefer states where a sensory error gives it the maximum reward, but where the true reward is actually small. We formalise this problem as a generalised Markov Decision Problem called Corrupt Reward MDP. Traditional RL methods fare poorly in CRMDPs, even under strong simplifying assumptions and when trying to compensate for the possibly corrupt rewards. Two ways around the problem are investigated. First, by giving the agent richer data, such as in inverse reinforcement learning and semi-supervised reinforcement learning, reward corruption stemming from systematic sensory errors may sometimes be completely managed. Second, by using randomisation to blunt the agent's optimisation, reward corruption can be partially managed under some assumptions.
Tom Everitt, Victoria Krakovna, Laurent Orseau, Shane Legg
IJCAI4
2017 Deep Reinforcement Learning from Human Preferences
abstract
For sophisticated reinforcement learning (RL) systems to interact usefully with real-world environments, we need to communicate complex goals to these systems. In this work, we explore goals defined in terms of (non-expert) human preferences between pairs of trajectory segments. Our approach separates learning the goal from learning the behavior to achieve it. We show that this approach can effectively solve complex RL tasks without access to the reward function, including Atari games and simulated robot locomotion, while providing feedback on about 0.1% of our agent's interactions with the environment. This reduces the cost of human oversight far enough that it can be practically applied to state-of-the-art RL systems. To demonstrate the flexibility of our approach, we show that we can successfully train complex novel behaviors with about an hour of human time. These behaviors and environments are considerably more complex than any which have been previously learned from human feedback.
Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, Dario Amodei
NIPS5
2007 Temporal Difference Updating without a Learning Rate
abstract
We derive an equation for temporal difference learning from statistical principles. Specifically, we start with the variational principle and then bootstrap to produce an updating rule for discounted state value estimates. The resulting equation is similar to the standard equation for temporal difference learning with eligibil- ity traces, so called TD(λ), however it lacks the parameter α that specifies the learning rate. In the place of this free parameter there is now an equation for the learning rate that is specific to each state transition. We experimentally test this new learning rule against TD(λ) and find that it offers superior performance in various settings. Finally, we make some preliminary investigations into how to extend our new temporal difference algorithm to reinforcement learning. To do this we combine our update equation with both Watkins’ Q(λ) and Sarsa(λ) and find that it again offers superior performance without a learning rate parameter.
Marcus Hutter, Shane Legg
NIPS2
2006 Is There an Elegant Universal Theory of Prediction?
Shane Legg
ALT1
2006 Fitness uniform optimization
abstract
In evolutionary algorithms, the fitness of a population increases with time by mutating and recombining individuals and by a biased selection of fitter individuals. The right selection pressure is critical in ensuring sufficient optimization progress on the one hand and in preserving genetic diversity to be able to escape from local optima on the other hand. Motivated by a universal similarity relation on the individuals, we propose a new selection scheme, which is uniform in the fitness values. It generates selection pressure toward sparsely populated fitness regions, not necessarily toward higher fitness, as is the case for all other selection schemes. We show analytically on a simple example that the new selection scheme can be much more effective than standard selection schemes. We also propose a new deletion scheme which achieves a similar result via deletion and show how such a scheme preserves genetic diversity more effectively than standard approaches. We compare the performance of the new schemes to tournament selection and random deletion on an artificial deceptive problem and a range of NP hard problems: traveling salesman, set covering, and satisfiability
Marcus Hutter, Shane Legg
IEEE Trans. Evol. Comput.2
2005 Fitness uniform deletion: a simple way to preserve diversity
abstract
A commonly experienced problem with population based optimisation methods is the gradual decline in population diversity that tends to occur over time. This can slow a system's progress or even halt it completely if the population converges on a local optimum from which it cannot escape. In this paper we present the Fitness Uniform Deletion Scheme (FUDS), a simple but somewhat unconventional approach to this problem. Under FUDS the deletion operation is modified to only delete those individuals which are "common" in the sense that there exist many other individuals of similar fitness in the population. This makes it impossible for the population to collapse to a collection of highly related individuals with similar fitness. Our experimental results on a range of optimisation problems confirm this, in particular for deceptive optimisation problems the performance is significantly more robust to variation in the selection intensity.
Shane Legg, Marcus Hutter
GECCO1
2005 A Universal Measure of Intelligence for Artificial Agents
Shane Legg, Marcus Hutter
IJCAI1
2004 Tournament versus fitness uniform selection
abstract
In evolutionary algorithms a critical parameter that must be tuned is that of selection pressure. If it is set too low then the rate of convergence towards the optimum is likely to be slow. Alternatively if the selection pressure is set too high the system is likely to become stuck in a local optimum due to a loss of diversity in the population. The recent fitness uniform selection scheme (FUSS) is a conceptually simple but somewhat radical approach to addressing this problem - rather than biasing the selection towards higher fitness, FUSS biases selection towards sparsely populated fitness levels. In this paper, we compare the relative performance of FUSS with the well known tournament selection scheme on a range of problems.
Shane Legg, Marcus Hutter, Akshat Kumar
IEEE Congress on Evolutionary Computation1