EDBT 2026 Demo / reviewers in the wild / expert
Tom Everitt
dblp:151/4259
· DBLP profile ↗
21ranked-venue papers
4as first author
16since 2021 · last 2025
0000-0003-1210-9866ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 4 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 7 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Limits of Predicting Agents from BehaviourabstractAs the complexity of AI systems and their interactions with the world increases, generating explanations for their behaviour is important for safely deploying AI. For agents, the most natural abstractions for predicting behaviour attribute beliefs, intentions and goals to the system. If an agent behaves as if it has a certain goal or belief, then we can make reasonable predictions about how it will behave in novel situations, including those where comprehensive safety evaluations are untenable. How well can we infer an agent’s beliefs from their behaviour, and how reliably can these inferred beliefs predict the agent’s behaviour in novel situations? We provide a precise answer to this question under the assumption that the agent’s behaviour is guided by a world model. Our contribution is the derivation of novel bounds on the agent's behaviour in new (unseen) deployment environments, which represent a theoretical limit for predicting intentional agents from behavioural data alone. We discuss the implications of these results for several research areas including fairness and safety. Alexis Bellot, Jonathan Richens, Tom Everitt |
ICML | 3 |
| 2025 | General agents need world modelsabstractAre world models a necessary ingredient for flexible, goal-directed behaviour, or is model-free learning sufficient?
We provide a formal answer to this question, showing that any agent capable of generalizing to multi-step goal-directed tasks must have learned a predictive model of its environment.
We show that this model can be extracted from the agent's policy, and that increasing the agents performance or the complexity of the goals it can achieve requires learning increasingly accurate world models.
This has a number of consequences: from developing safe and general agents, to bounding agent capabilities in complex environments, and providing new algorithms for eliciting world models from agents. Jonathan Richens, Tom Everitt, David Abel |
ICML | 2 |
| 2025 | Incentives for responsiveness, instrumental control and impact
Ryan Carey, Eric D. Langlois 0002, Chris van Merwijk, Shane Legg, Tom Everitt |
Artif. Intell. | 5 |
| 2024 | Reasoning about Causality in Games (Abstract Reprint)abstractCausal reasoning and game-theoretic reasoning are fundamental topics in artificial intelligence, among many other disciplines: this paper is concerned with their intersection. Despite their importance, a formal framework that supports both these forms of reasoning has, until now, been lacking. We offer a solution in the form of (structural) causal games, which can be seen as extending Pearl's causal hierarchy to the game-theoretic domain, or as extending Koller and Milch's multi-agent influence diagrams to the causal domain. We then consider three key questions: i) How can the (causal) dependencies in games – either between variables, or between strategies – be modelled in a uniform, principled manner? ii) How may causal queries be computed in causal games, and what assumptions does this require? iii) How do causal games compare to existing formalisms? To address question i), we introduce mechanised games, which encode dependencies between agents' decision rules and the distributions governing the game. In response to question ii), we present definitions of predictions, interventions, and counterfactuals, and discuss the assumptions required for each. Regarding question iii), we describe correspondences between causal games and other formalisms, and explain how causal games can be used to answer queries that other causal or game-theoretic models do not support. Finally, we highlight possible applications of causal games, aided by an extensive open-source Python library. Lewis Hammond, James Fox, Tom Everitt, Ryan Carey, Alessandro Abate, Michael J. Wooldridge |
AAAI | 3 |
| 2024 | Discovering Agents (Abstract Reprint)abstractCausal models of agents have been used to analyse the safety aspects of machine learning systems. But identifying agents is non-trivial – often the causal model is just assumed by the modeller without much justification – and modelling failures can lead to mistakes in the safety analysis. This paper proposes the first formal causal definition of agents – roughly that agents are systems that would adapt their policy if their actions influenced the world in a different way. From this we derive the first causal discovery algorithm for discovering the presence of agents from empirical data, given a set of variables and under certain assumptions. We also provide algorithms for translating between causal models and game-theoretic influence diagrams. We demonstrate our approach by resolving some previous confusions caused by incorrect causal modelling of agents. Zachary Kenton, Ramana Kumar, Sebastian Farquhar, Jonathan Richens, Matt MacDermott, Tom Everitt |
AAAI | 6 |
| 2024 | Robust agents learn causal world modelsabstractIt has long been hypothesised that causal reasoning plays a fundamental role in robust and general intelligence. However, it is not known if agents must learn causal models in order to generalise to new domains, or if other inductive biases are sufficient. We answer this question, showing that any agent capable of satisfying a regret bound for a large set of distributional shifts must have learned an approximate causal model of the data generating process, which converges to the true causal model for optimal agents. We discuss the implications of this result for several research areas including transfer learning and causal inference. Jonathan Richens, Tom Everitt |
ICLR | 2 |
| 2024 | Measuring Goal-DirectednessabstractWe define maximum entropy goal-directedness (MEG), a formal measure of goal-
directedness in causal models and Markov decision processes, and give algorithms
for computing it. Measuring goal-directedness is important, as it is a critical
element of many concerns about harm from AI. It is also of philosophical interest,
as goal-directedness is a key aspect of agency. MEG is based on an adaptation of
the maximum causal entropy framework used in inverse reinforcement learning. It
can measure goal-directedness with respect to a known utility function, a hypothesis
class of utility functions, or a set of random variables. We prove that MEG satisfies
several desiderata and demonstrate our algorithms with small-scale experiments. Matt MacDermott, James Fox, Francesco Belardinelli, Tom Everitt |
NeurIPS | 4 |
| 2023 | Honesty Is the Best Policy: Defining and Mitigating AI DeceptionabstractDeceptive agents are a challenge for the safety, trustworthiness, and cooperation of AI systems. We focus on the problem that agents might deceive in order to achieve their goals (for instance, in our experiments with language models, the goal of being evaluated as truthful).
There are a number of existing definitions of deception in the literature on game theory and symbolic AI, but there is no overarching theory of deception for learning agents in games.
We introduce a formal
definition of deception in structural causal games, grounded in the philosophy
literature, and applicable to real-world machine learning systems.
Several examples and results illustrate that our formal definition aligns with the philosophical and commonsense meaning of deception.
Our main technical result is to provide graphical criteria for deception.
We show, experimentally, that these results can be used to mitigate deception in reinforcement learning agents and language models. Francis Rhys Ward, Francesca Toni, Francesco Belardinelli, Tom Everitt |
NeurIPS | 4 |
| 2023 | Human Control: Definitions and AlgorithmsabstractHow can humans stay in control of advanced artificial intelligence systems? One proposal is corrigibility, which requires the agent to follow the instructions of a human overseer, without inappropriately influencing them. In this paper, we formally define a variant of corrigibility called shutdown instructability, and show that it implies appropriate shutdown behavior, retention of human autonomy, and avoidance of user harm. We also analyse the related concepts of non-obstruction and shutdown alignment, three previously proposed algorithms for human control, and one new algorithm. Ryan Carey, Tom Everitt |
UAI | 2 |
| 2023 | Reasoning about causality in gamesabstractCausal reasoning and game-theoretic reasoning are fundamental topics in artificial intelligence, among many other disciplines: this paper is concerned with their intersection. Despite their importance, a formal framework that supports both these forms of reasoning has, until now, been lacking. We offer a solution in the form of (structural) causal games, which can be seen as extending Pearl's causal hierarchy to the game-theoretic domain, or as extending Koller and Milch's multi-agent influence diagrams to the causal domain. We then consider three key questions: How can the (causal) dependencies in games – either between variables, or between strategies – be modelled in a uniform, principled manner? How may causal queries be computed in causal games, and what assumptions does this require? How do causal games compare to existing formalisms? Lewis Hammond, James Fox, Tom Everitt, Ryan Carey, Alessandro Abate, Michael J. Wooldridge |
Artif. Intell. | 3 |
| 2023 | Discovering agentsabstractCausal models of agents have been used to analyse the safety aspects of machine learning systems. But identifying agents is non-trivial – often the causal model is just assumed by the modeller without much justification – and modelling failures can lead to mistakes in the safety analysis. This paper proposes the first formal causal definition of agents – roughly that agents are systems that would adapt their policy if their actions influenced the world in a different way. From this we derive the first causal discovery algorithm for discovering the presence of agents from empirical data, given a set of variables and under certain assumptions. We also provide algorithms for translating between causal models and game-theoretic influence diagrams. We demonstrate our approach by resolving some previous confusions caused by incorrect causal modelling of agents. Zachary Kenton, Ramana Kumar, Sebastian Farquhar, Jonathan Richens, Matt MacDermott, Tom Everitt |
Artif. Intell. | 6 |
| 2022 | Why Fair Labels Can Yield Unfair Predictions: Graphical Conditions for Introduced UnfairnessabstractIn addition to reproducing discriminatory relationships in the training data, machine learning (ML) systems can also introduce or amplify discriminatory effects. We refer to this as introduced unfairness, and investigate the conditions under which it may arise. To this end, we propose introduced total variation as a measure of introduced unfairness, and establish graphical conditions under which it may be incentivised to occur. These criteria imply that adding the sensitive attribute as a feature removes the incentive for introduced variation under well-behaved loss functions. Additionally, taking a causal perspective, introduced path-specific effects shed light on the issue of when specific paths should be considered fair. Carolyn Ashurst, Ryan Carey, Silvia Chiappa, Tom Everitt |
AAAI | 4 |
| 2022 | Path-Specific Objectives for Safer Agent IncentivesabstractWe present a general framework for training safe agents whose naive incentives are unsafe. As an example, manipulative or deceptive behaviour can improve rewards but should be avoided. Most approaches fail here: agents maximize expected return by any means necessary. We formally describe settings with `delicate' parts of the state which should not be used as a means to an end. We then train agents to maximize the causal effect of actions on the expected return which is not mediated by the delicate parts of state, using Causal Influence Diagram analysis. The resulting agents have no incentive to control the delicate state. We further show how our framework unifies and generalizes existing proposals. Sebastian Farquhar, Ryan Carey, Tom Everitt |
AAAI | 3 |
| 2022 | A Complete Criterion for Value of Information in Soluble Influence DiagramsabstractInfluence diagrams have recently been used to analyse the safety and fairness properties of AI systems. A key building block for this analysis is a graphical criterion for value of information (VoI). This paper establishes the first complete graphical criterion for VoI in influence diagrams with multiple decisions. Along the way, we establish two techniques for proving properties of multi-decision influence diagrams: ID homomorphisms are structure-preserving transformations of influence diagrams, while a Tree of Systems is a collection of paths that captures how information and control can flow in an influence diagram. Chris van Merwijk, Ryan Carey, Tom Everitt |
AAAI | 3 |
| 2021 | Agent Incentives: A Causal PerspectiveabstractWe present a framework for analysing agent incentives using causal influence diagrams. We establish that a well-known criterion for value of information is complete. We propose a new graphical criterion for value of control, establishing its soundness and completeness. We also introduce two new concepts for incentive analysis: response incentives indicate which changes in the environment affect an optimal decision, while instrumental control incentives establish whether an agent can influence its utility via a variable X. For both new concepts, we provide sound and complete graphical criteria. We show by example how these results can help with evaluating the safety and fairness of an AI system Tom Everitt, Ryan Carey, Eric D. Langlois 0002, Pedro A. Ortega, Shane Legg |
AAAI | 1 |
| 2021 | How RL Agents Behave When Their Actions Are ModifiedabstractReinforcement learning in complex environments may require supervision to prevent the agent from attempting dangerous actions. As a result of supervisor intervention, the executed action may differ from the action specified by the policy. How does this affect learning? We present the Modified-Action Markov Decision Process, an extension of the MDP model that allows actions to differ from the policy. We analyze the asymptotic behaviours of common reinforcement learning algorithms in this setting and show that they adapt in different ways: some completely ignore modifications while others go to various lengths in trying to avoid action modifications that decrease reward. By choosing the right algorithm, developers can prevent their agents from learning to circumvent interruptions or constraints, and better control agent responses to other kinds of action modification, like self-damage. Eric D. Langlois 0002, Tom Everitt |
AAAI | 2 |
| 2018 | AGI Safety Literature ReviewabstractThe development of Artificial General Intelligence (AGI) promises to be a major event. Along with its many potential benefits, it also raises serious safety concerns. The intention of this paper is to provide an easily accessible and up-to-date collection of references for the emerging field of AGI safety. A significant number of safety problems for AGI have been identified. We list these, and survey recent research on solving them. We also cover works on how best to think of AGI from the limited knowledge we have today, predictions for when AGI will first be created, and what will happen after its creation. Finally, we review the current public policy on AGI. Tom Everitt, Gary Lea, Marcus Hutter |
IJCAI | 1 |
| 2017 | Reinforcement Learning with a Corrupted Reward ChannelabstractNo real-world reward function is perfect. Sensory errors and software bugs may result in agents getting higher (or lower) rewards than they should. For example, a reinforcement learning agent may prefer states where a sensory error gives it the maximum reward, but where the true reward is actually small. We formalise this problem as a generalised Markov Decision Problem called Corrupt Reward MDP. Traditional RL methods fare poorly in CRMDPs, even under strong simplifying assumptions and when trying to compensate for the possibly corrupt rewards. Two ways around the problem are investigated. First, by giving the agent richer data, such as in inverse reinforcement learning and semi-supervised reinforcement learning, reward corruption stemming from systematic sensory errors may sometimes be completely managed. Second, by using randomisation to blunt the agent's optimisation, reward corruption can be partially managed under some assumptions. Tom Everitt, Victoria Krakovna, Laurent Orseau, Shane Legg |
IJCAI | 1 |
| 2017 | Count-Based Exploration in Feature Space for Reinforcement LearningabstractWe introduce a new count-based optimistic exploration algorithm for Reinforcement Learning (RL) that is feasible in environments with high-dimensional state-action spaces. The success of RL algorithms in these domains depends crucially on generalisation from limited training experience. Function approximation techniques enable RL agents to generalise in order to estimate the value of unvisited states, but at present few methods enable generalisation regarding uncertainty. This has prevented the combination of scalable RL algorithms with efficient exploration strategies that drive the agent to reduce its uncertainty. We present a new method for computing a generalised state visit-count, which allows the agent to estimate the uncertainty associated with any state. Our \phi-pseudocount achieves generalisation by exploiting same feature representation of the state space that is used for value function approximation. States that have less frequently observed features are deemed more uncertain. The \phi-Exploration-Bonus algorithm rewards the agent for exploring in feature space rather than in the untransformed state space. The method is simpler and less computationally expensive than some previous proposals, and achieves near state-of-the-art results on high-dimensional RL benchmarks. Jarryd Martin, Suraj Narayanan Sasikumar, Tom Everitt, Marcus Hutter |
IJCAI | 3 |
| 2014 | Free Lunch for optimisation under the universal distributionabstractFunction optimisation is a major challenge in computer science. The No Free Lunch theorems state that if all functions with the same histogram are assumed to be equally probable then no algorithm outperforms any other in expectation. We argue against the uniform assumption and suggest a universal prior exists for which there is a free lunch, but where no particular class of functions is favoured over another. We also prove upper and lower bounds on the size of the free lunch. Tom Everitt, Tor Lattimore, Marcus Hutter |
IEEE Congress on Evolutionary Computation | 1 |
| 2014 | Can we measure the difficulty of an optimization problem?abstractCan we measure the difficulty of an optimization problem? Although optimization plays a crucial role in modern science and technology, a formal framework that puts problems and solution algorithms into a broader context has not been established. This paper presents a conceptual approach which gives a positive answer to the question for a broad class of optimization problems. Adopting an information and computational perspective, the proposed framework builds upon Shannon and algorithmic information theories. As a starting point, a concrete model and definition of optimization problems is provided. Then, a formal definition of optimization difficulty is introduced which builds upon algorithmic information theory. Following an initial analysis, lower and upper bounds on optimization difficulty are established. One of the upper-bounds is closely related to Shannon information theory and black-box optimization. Finally, various computational issues and future research directions are discussed. Tansu Alpcan, Tom Everitt, Marcus Hutter |
ITW | 2 |