Edward Hughes 0001

dblp:217/2003 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-2434-2334ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
15 papers
Reinforcement learning · 49% Multi-agent systems · 23% Knowledge representation and reasoning · 7%
Theoretical computer science
3 papers
Algorithmic game theory and mechanism design · 100%

Topics — the 30 heaviest of 38, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-agent reinforcement learning
2.972021
Collaborating with Humans without Human Data · NeurIPS 2021
The Hanabi challenge: A new frontier for AI research · Artif. Intell. 2020
Learning to Incentivize Other Learning Agents · NeurIPS 2020
Machine learning › Reinforcement learning
meta-reinforcement learning
1.422024
Towards a Pretrained Model for Restless Bandits via Multi-arm Generalization · IJCAI 2024
Human-Timescale Adaptation in an Open-Ended Task Space · ICML 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning
open-world learning
1.022024
Position: Open-Endedness is Essential for Artificial Superhuman Intelligence · ICML 2024
Artificial Generational Intelligence: Cultural Accumulation in Reinforcement Learning · NeurIPS 2024
Knowledge, reasoning and agents › Multi-agent systems
multi-agent learning
0.922020
A Generalized Training Approach for Multiagent Learning · ICLR 2020
Smooth markets: A basic mechanism for organizing gradient-based learners · ICLR 2020
Machine learning › Deep learning architectures and training
foundation model
0.812024
Position: Open-Endedness is Essential for Artificial Superhuman Intelligence · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › sequential latent variable model
latent action models
0.812024
Genie: Generative Interactive Environments · ICML 2024
Machine learning › Reinforcement learning › multi-armed bandit
restless bandits
0.812024
Towards a Pretrained Model for Restless Bandits via Multi-arm Generalization · IJCAI 2024
Knowledge, reasoning and agents › Multi-agent systems › multi-agent learning
social learning
0.812024
Artificial Generational Intelligence: Cultural Accumulation in Reinforcement Learning · NeurIPS 2024
Machine learning › Reinforcement learning › model-based reinforcement learning
world model
0.812024
Genie: Generative Interactive Environments · ICML 2024
Machine learning › Reinforcement learning
generalization in reinforcement learning
0.712023
Human-Timescale Adaptation in an Open-Ended Task Space · ICML 2023
Machine learning › Reinforcement learning › meta-reinforcement learning
in-context reinforcement learning
0.712023
Human-Timescale Adaptation in an Open-Ended Task Space · ICML 2023
Machine learning › Transfer learning and domain adaptation
zero-shot transfer
0.712023
Human-Timescale Adaptation in an Open-Ended Task Space · ICML 2023
Knowledge, reasoning and agents › Multi-agent systems › multi-agent collaboration
ad hoc teamwork
0.512021
Collaborating with Humans without Human Data · NeurIPS 2021
Machine learning › Reinforcement learning › multi-agent reinforcement learning
human-AI collaboration
0.512021
Collaborating with Humans without Human Data · NeurIPS 2021
Knowledge, reasoning and agents › Multi-agent systems
automated negotiation
0.412020
Negotiating team formation using deep reinforcement learning · Artif. Intell. 2020
Knowledge, reasoning and agents › Multi-agent systems › game theory
cooperative game
0.412020
The Hanabi challenge: A new frontier for AI research · Artif. Intell. 2020
Machine learning › Reinforcement learning
deep reinforcement learning
0.412020
Negotiating team formation using deep reinforcement learning · Artif. Intell. 2020
Machine learning › Efficient and distributed learning › federated learning
incentive mechanism
0.412020
Learning to Incentivize Other Learning Agents · NeurIPS 2020
Knowledge, reasoning and agents › Multi-agent systems
multi-agent collaboration
0.412020
The Hanabi challenge: A new frontier for AI research · Artif. Intell. 2020
Machine learning › Reinforcement learning › multi-agent reinforcement learning
opponent shaping
0.412020
Learning to Incentivize Other Learning Agents · NeurIPS 2020
Knowledge, reasoning and agents › Multi-agent systems
team formation
0.412020
Negotiating team formation using deep reinforcement learning · Artif. Intell. 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning
theory of mind
0.412020
The Hanabi challenge: A new frontier for AI research · Artif. Intell. 2020
Algorithmic game theory and mechanism design › market design
market mechanism
0.412020
Smooth markets: A basic mechanism for organizing gradient-based learners · ICLR 2020
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.412019
Bayesian Action Decoder for Deep Multi-Agent Reinforcement Learning · ICML 2019
Knowledge, reasoning and agents › Multi-agent systems
emergent communication
0.412019
Social Influence as Intrinsic Motivation for Multi-Agent Deep Reinforcement Learning · ICML 2019
Machine learning › Reinforcement learning › exploration
intrinsic motivation
0.412019
Social Influence as Intrinsic Motivation for Multi-Agent Deep Reinforcement Learning · ICML 2019
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
partially observable markov decision process
0.412019
Bayesian Action Decoder for Deep Multi-Agent Reinforcement Learning · ICML 2019
Machine learning › Reinforcement learning › reward learning
reward modeling
0.412019
Learning to Understand Goal Specifications by Modelling Reward · ICLR (Poster) 2019
Machine learning › Reinforcement learning › multi-agent reinforcement learning
cooperative multi-agent reinforcement learning
0.312018
Inequity aversion improves cooperation in intertemporal social dilemmas · NeurIPS 2018
Knowledge, reasoning and agents › Multi-agent systems › game theory
social dilemmas
0.312018
Inequity aversion improves cooperation in intertemporal social dilemmas · NeurIPS 2018

Methods — techniques the papers use, named apart from their topics

unsupervised learning · 0.8spatiotemporal tokenizer · 0.8novelty and learnability · 0.8in-weights learning · 0.8in-context learning · 0.8imitation · 0.8formal definition of open-endedness · 0.8autoregressive dynamics models · 0.8automated curriculum · 0.7attention-based memory · 0.7incentive function learning · 0.4gradient-based learning · 0.4game theory · 0.4actor-critic · 0.4multi-agent reinforcement learning · 0.3inequity aversion modeling · 0.3
YearPublicationVenuePosition
2025 Cultural Evolution of Cooperation among LLM Agents
Aron Vallinder, Edward Hughes 0001
AAMAS2
2024 Position: Open-Endedness is Essential for Artificial Superhuman Intelligence
abstract
In recent years there has been a tremendous surge in the general capabilities of AI systems, mainly fuelled by training foundation models on internet-scale data. Nevertheless, the creation of open-ended, ever self-improving AI remains elusive. **In this position paper, we argue that the ingredients are now in place to achieve *open-endedness* in AI systems with respect to a human observer. Furthermore, we claim that such open-endedness is an essential property of any artificial superhuman intelligence (ASI).** We begin by providing a concrete formal definition of open-endedness through the lens of novelty and learnability. We then illustrate a path towards ASI via open-ended systems built on top of foundation models, capable of making novel, human-relevant discoveries. We conclude by examining the safety implications of generally-capable open-ended AI. We expect that open-ended foundation models will prove to be an increasingly fertile and safety-critical area of research in the near future.
Edward Hughes 0001, Michael Dennis 0001, Jack Parker-Holder, Feryal M. P. Behbahani, Aditi Mavalankar, Yuge Shi, Tom Schaul, Tim Rocktäschel
ICML1
2024 Genie: Generative Interactive Environments
abstract
We introduce Genie, the first *generative interactive environment* trained in an unsupervised manner from unlabelled Internet videos. The model can be prompted to generate an endless variety of action-controllable virtual worlds described through text, synthetic images, photographs, and even sketches. At 11B parameters, Genie can be considered a *foundation world model*. It is comprised of a spatiotemporal video tokenizer, an autoregressive dynamics model, and a simple and scalable latent action model. Genie enables users to act in the generated environments on a frame-by-frame basis *despite training without any ground-truth action labels* or other domain specific requirements typically found in the world model literature. Further the resulting learned latent action space facilitates training agents to imitate behaviors from unseen videos, opening the path for training generalist agents of the future.
Jake Bruce, Michael Dennis 0001, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes 0001, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, Yusuf Aytar, Sarah Bechtle, Feryal M. P. Behbahani, Stephanie C. Y. Chan, Nicolas Heess, Lucy Gonzalez, Simon Osindero, Sherjil Ozair, Scott E. Reed, Jingwei Zhang 0001, Konrad Zolna, Jeff Clune, Nando de Freitas, Satinder Singh 0001, Tim Rocktäschel
ICML6
2024 Towards a Pretrained Model for Restless Bandits via Multi-arm Generalization
Yunfan Zhao, Nikhil Behari, Edward Hughes 0001, Edwin Zhang, Dheeraj Nagaraj, Karl Tuyls, Aparna Taneja, Milind Tambe
IJCAI3
2024 Artificial Generational Intelligence: Cultural Accumulation in Reinforcement Learning
abstract
Cultural accumulation drives the open-ended and diverse progress in capabilities spanning human history. It builds an expanding body of knowledge and skills by combining individual exploration with inter-generational information transmission. Despite its widespread success among humans, the capacity for artificial learning agents to accumulate culture remains under-explored. In particular, approaches to reinforcement learning typically strive for improvements over only a single lifetime. Generational algorithms that do exist fail to capture the open-ended, emergent nature of cultural accumulation, which allows individuals to trade-off innovation and imitation. Building on the previously demonstrated ability for reinforcement learning agents to perform social learning, we find that training setups which balance this with independent learning give rise to cultural accumulation. These accumulating agents outperform those trained for a single lifetime with the same cumulative experience. We explore this accumulation by constructing two models under two distinct notions of a generation: episodic generations, in which accumulation occurs via in-context learning and train-time generations, in which accumulation occurs via in-weights learning. In-context and in-weights cultural accumulation can be interpreted as analogous to knowledge and skill accumulation, respectively. To the best of our knowledge, this work is the first to present general models that achieve emergent cultural accumulation in reinforcement learning, opening up new avenues towards more open-ended learning systems, as well as presenting new opportunities for modelling human culture.
Jonathan Cook 0004, Chris Lu 0001, Edward Hughes 0001, Joel Z. Leibo, Jakob N. Foerster
NeurIPS3
2023 Human-Timescale Adaptation in an Open-Ended Task Space
abstract
Foundation models have shown impressive adaptation and scalability in supervised and self-supervised learning problems, but so far these successes have not fully translated to reinforcement learning (RL). In this work, we demonstrate that training an RL agent at scale leads to a general in-context learning algorithm that can adapt to open-ended novel embodied 3D problems as quickly as humans. In a vast space of held-out environment dynamics, our adaptive agent (AdA) displays on-the-fly hypothesis-driven exploration, efficient exploitation of acquired knowledge, and can successfully be prompted with first-person demonstrations. Adaptation emerges from three ingredients: (1) meta-reinforcement learning across a vast, smooth and diverse task distribution, (2) a policy parameterised as a large-scale attention-based memory architecture, and (3) an effective automated curriculum that prioritises tasks at the frontier of an agent’s capabilities. We demonstrate characteristic scaling laws with respect to network size, memory length, and richness of the training task distribution. We believe our results lay the foundation for increasingly general and adaptive RL agents that perform well across ever-larger open-ended domains.
Jakob Bauer, Kate Baumli, Feryal M. P. Behbahani, Avishkar Bhoopchand, Nathalie Bradley-Schmieg, Natalie Clay, Adrian Collister, Vibhavari Dasagi, Lucy Gonzalez, Karol Gregor, Edward Hughes 0001, Sheleem Kashem, Maria Loks-Thompson, Hannah Openshaw, Jack Parker-Holder, Shreya Pathak, Nicolas Perez Nieves, Nemanja Rakicevic, Tim Rocktäschel, Yannick Schroecker, Satinder Singh 0001, Jakub Sygnowski, Karl Tuyls, Sarah York, Alexander Zacherl, Lei M. Zhang
ICML12
2021 Collaborating with Humans without Human Data
abstract
Collaborating with humans requires rapidly adapting to their individual strengths, weaknesses, and preferences. Unfortunately, most standard multi-agent reinforcement learning techniques, such as self-play (SP) or population play (PP), produce agents that overfit to their training partners and do not generalize well to humans. Alternatively, researchers can collect human data, train a human model using behavioral cloning, and then use that model to train "human-aware" agents ("behavioral cloning play", or BCP). While such an approach can improve the generalization of agents to new human co-players, it involves the onerous and expensive step of collecting large amounts of human data first. Here, we study the problem of how to train agents that collaborate well with human partners without using human data. We argue that the crux of the problem is to produce a diverse set of training partners. Drawing inspiration from successful multi-agent approaches in competitive domains, we find that a surprisingly simple approach is highly effective. We train our agent partner as the best response to a population of self-play agents and their past checkpoints taken throughout training, a method we call Fictitious Co-Play (FCP). Our experiments focus on a two-player collaborative cooking simulator that has recently been proposed as a challenge problem for coordination with humans. We find that FCP agents score significantly higher than SP, PP, and BCP when paired with novel agent and human partners. Furthermore, humans also report a strong subjective preference to partnering with FCP agents over all baselines.
DJ Strouse, Kevin R. McKee, Matt M. Botvinick, Edward Hughes 0001, Richard Everett 0001
NeurIPS4
2020 Smooth markets: A basic mechanism for organizing gradient-based learners
David Balduzzi, Wojciech Czarnecki 0001, Thomas W. Anthony 0001, Ian Gemp, Edward Hughes 0001, Joel Z. Leibo, Georgios Piliouras, Thore Graepel
ICLR5
2020 A Generalized Training Approach for Multiagent Learning
Paul Muller, Shayegan Omidshafiei, Mark Rowland 0001, Karl Tuyls, Julien Pérolat, Siqi Liu 0002, Daniel Hennes, Luke Marris, Marc Lanctot, Edward Hughes 0001, Zhe Wang 0055, Guy Lever, Nicolas Heess, Thore Graepel, Rémi Munos
ICLR10
2020 Learning to Incentivize Other Learning Agents
abstract
The challenge of developing powerful and general Reinforcement Learning (RL) agents has received increasing attention in recent years. Much of this effort has focused on the single-agent setting, in which an agent maximizes a predefined extrinsic reward function. However, a long-term question inevitably arises: how will such independent agents cooperate when they are continually learning and acting in a shared multi-agent environment? Observing that humans often provide incentives to influence others' behavior, we propose to equip each RL agent in a multi-agent environment with the ability to give rewards directly to other agents, using a learned incentive function. Each agent learns its own incentive function by explicitly accounting for its impact on the learning of recipients and, through them, the impact on its own extrinsic objective. We demonstrate in experiments that such agents significantly outperform standard RL and opponent-shaping agents in challenging general-sum Markov games, often by finding a near-optimal division of labor. Our work points toward more opportunities and challenges along the path to ensure the common good in a multi-agent future.
Mehrdad Farajtabar, Peter Sunehag, Edward Hughes 0001, Hongyuan Zha
NeurIPS5
2020 Bounds and dynamics for empirical game theoretic analysis
abstract
Abstract This paper provides several theoretical results for empirical game theory. Specifically, we introduce bounds for empirical game theoretical analysis of complex multi-agent interactions. In doing so we provide insights in the empirical meta game showing that a Nash equilibrium of the estimated meta-game is an approximate Nash equilibrium of the true underlying meta-game. We investigate and show how many data samples are required to obtain a close enough approximation of the underlying game. Additionally, we extend the evolutionary dynamics analysis of meta-games using heuristic payoff tables (HPTs) to asymmetric games. The state-of-the-art has only considered evolutionary dynamics of symmetric HPTs in which agents have access to the same strategy sets and the payoff structure is symmetric, implying that agents are interchangeable. Finally, we carry out an empirical illustration of the generalised method in several domains, illustrating the theory and evolutionary dynamics of several versions of theAlphaGoalgorithm (symmetric), the dynamics of the Colonel Blotto game played by human players on Facebook (symmetric), the dynamics of several teams of players in the capture the flag game (symmetric), and an example of a meta-game in Leduc Poker (asymmetric), generated by the policy-space response oracle multi-agent learning algorithm.
Karl Tuyls, Julien Pérolat, Marc Lanctot, Edward Hughes 0001, Richard Everett 0001, Joel Z. Leibo, Csaba Szepesvári, Thore Graepel
Auton. Agents Multi Agent Syst.4
2020 Negotiating team formation using deep reinforcement learning
Yoram Bachrach, Richard Everett 0001, Edward Hughes 0001, Angeliki Lazaridou, Joel Z. Leibo, Marc Lanctot, Michael Johanson, Wojciech Czarnecki 0001, Thore Graepel
Artif. Intell.3
2020 The Hanabi challenge: A new frontier for AI research
abstract
From the early days of computing, games have been important testbeds for studying how well machines can do sophisticated decision making. In recent years, machine learning has made dramatic advances with artificial agents reaching superhuman performance in challenge domains like Go, Atari, and some variants of poker. As with their predecessors of chess, checkers, and backgammon, these game domains have driven research by providing sophisticated yet well-defined challenges for artificial intelligence practitioners. We continue this tradition by proposing the game of Hanabi as a new challenge domain with novel problems that arise from its combination of purely cooperative gameplay with two to five players and imperfect information. In particular, we argue that Hanabi elevates reasoning about the beliefs and intentions of other agents to the foreground. We believe developing novel techniques for such theory of mind reasoning will not only be crucial for success in Hanabi, but also in broader collaborative efforts, especially those with human partners. To facilitate future research, we introduce the open-source Hanabi Learning Environment, propose an experimental framework for the research community to evaluate algorithmic advances, and assess the performance of current state-of-the-art techniques.
Nolan Bard, Jakob N. Foerster, Sarath Chandar, Neil Burch, Marc Lanctot, H. Francis Song, Emilio Parisotto, Vincent Dumoulin, Subhodeep Moitra, Edward Hughes 0001, Iain Dunning, Shibl Mourad, Hugo Larochelle, Marc G. Bellemare, Michael H. Bowling
Artif. Intell.10
2019 Learning to Understand Goal Specifications by Modelling Reward
Dzmitry Bahdanau, Felix Hill, Jan Leike, Edward Hughes 0001, Seyed Arian Hosseini, Pushmeet Kohli, Edward Grefenstette
ICLR (Poster)4
2019 Bayesian Action Decoder for Deep Multi-Agent Reinforcement Learning
abstract
When observing the actions of others, humans make inferences about why they acted as they did, and what this implies about the world; humans also use the fact that their actions will be interpreted in this manner, allowing them to act informatively and thereby communicate efficiently with others. Although learning algorithms have recently achieved superhuman performance in a number of two-player, zero-sum games, scalable multi-agent reinforcement learning algorithms that can discover effective strategies and conventions in complex, partially observable settings have proven elusive. We present the Bayesian action decoder (BAD), a new multi-agent learning method that uses an approximate Bayesian update to obtain a public belief that conditions on the actions taken by all agents in the environment. BAD introduces a new Markov decision process, the public belief MDP, in which the action space consists of all deterministic partial policies, and exploits the fact that an agent acting only on this public belief state can still learn to use its private information if the action space is augmented to be over all partial policies mapping private information into environment actions. The Bayesian update is closely related to the theory of mind reasoning that humans carry out when observing others’ actions. We first validate BAD on a proof-of-principle two-step matrix game, where it outperforms policy gradient methods; we then evaluate BAD on the challenging, cooperative partial-information card game Hanabi, where, in the two-player setting, it surpasses all previously published learning and hand-coded approaches, establishing a new state of the art.
Jakob N. Foerster, H. Francis Song, Edward Hughes 0001, Neil Burch, Iain Dunning, Shimon Whiteson, Matt M. Botvinick, Michael H. Bowling
ICML3
2019 Social Influence as Intrinsic Motivation for Multi-Agent Deep Reinforcement Learning
abstract
We propose a unified mechanism for achieving coordination and communication in Multi-Agent Reinforcement Learning (MARL), through rewarding agents for having causal influence over other agents’ actions. Causal influence is assessed using counterfactual reasoning. At each timestep, an agent simulates alternate actions that it could have taken, and computes their effect on the behavior of other agents. Actions that lead to bigger changes in other agents’ behavior are considered influential and are rewarded. We show that this is equivalent to rewarding agents for having high mutual information between their actions. Empirical results demonstrate that influence leads to enhanced coordination and communication in challenging social dilemma environments, dramatically increasing the learning curves of the deep RL agents, and leading to more meaningful learned communication protocols. The influence rewards for all agents can be computed in a decentralized way by enabling agents to learn a model of other agents using deep neural networks. In contrast, key previous works on emergent communication in the MARL setting were unable to learn diverse policies in a decentralized manner and had to resort to centralized training. Consequently, the influence reward opens up a window of new opportunities for research in this area.
Natasha Jaques, Angeliki Lazaridou, Edward Hughes 0001, Caglar Gulcehre, Pedro A. Ortega, DJ Strouse, Joel Z. Leibo, Nando de Freitas
ICML3
2018 Inequity aversion improves cooperation in intertemporal social dilemmas
abstract
Groups of humans are often able to find ways to cooperate with one another in complex, temporally extended social dilemmas. Models based on behavioral economics are only able to explain this phenomenon for unrealistic stateless matrix games. Recently, multi-agent reinforcement learning has been applied to generalize social dilemma problems to temporally and spatially extended Markov games. However, this has not yet generated an agent that learns to cooperate in social dilemmas as humans do. A key insight is that many, but not all, human individuals have inequity averse social preferences. This promotes a particular resolution of the matrix game social dilemma wherein inequity-averse individuals are personally pro-social and punish defectors. Here we extend this idea to Markov games and show that it promotes cooperation in several types of sequential social dilemma, via a profitable interaction with policy learnability. In particular, we find that inequity aversion improves temporal credit assignment for the important class of intertemporal social dilemmas. These results help explain how large-scale cooperation may emerge and persist.
Edward Hughes 0001, Joel Z. Leibo, Matthew Phillips, Karl Tuyls, Edgar A. Duéñez-Guzmán, Antonio García Castañeda, Iain Dunning, Tina Zhu, Kevin R. McKee, Raphael Koster, Heather Roff, Thore Graepel
NeurIPS1