David J. Wu 0002

dblp:32/10400-2 · also David Jian Wu · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2024
0000-0002-5834-4936ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 8 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Reinforcement learning · 61% Multi-agent systems · 20% Planning, search and constraint satisfaction · 12%
Theoretical computer science
2 papers
Algorithmic game theory and mechanism design · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-agent reinforcement learning
2.342023
Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning · ICLR 2023
Self-Explaining Deviations for Coordination · NeurIPS 2022
Modeling Strong and Human-Like Gameplay with KL-Regularized Search · ICML 2022
Machine learning › Reinforcement learning › model-based reinforcement learning › model-based planning
decision-time planning
1.422024
The Update-Equivalence Framework for Decision-Time Planning · ICLR 2024
Abstracting Imperfect Information Away from Two-Player Zero-Sum Games · ICML 2023
Knowledge, reasoning and agents › Multi-agent systems
imperfect information games
1.422024
The Update-Equivalence Framework for Decision-Time Planning · ICLR 2024
Abstracting Imperfect Information Away from Two-Player Zero-Sum Games · ICML 2023
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
game playing
0.712023
Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning · ICLR 2023
Algorithmic game theory and mechanism design
game solving
0.712023
Abstracting Imperfect Information Away from Two-Player Zero-Sum Games · ICML 2023
Algorithmic game theory and mechanism design
zero-sum game
0.712023
Abstracting Imperfect Information Away from Two-Player Zero-Sum Games · ICML 2023
Natural language and speech › Language models and text generation
large language model fine-tuning
0.612022
A Fine-Tuning Approach to Belief State Modeling · ICLR 2022
Knowledge, reasoning and agents › Multi-agent systems
multi-agent coordination
0.612022
Self-Explaining Deviations for Coordination · NeurIPS 2022
Machine learning › Reinforcement learning › multi-agent reinforcement learning
opponent modeling
0.612022
Modeling Strong and Human-Like Gameplay with KL-Regularized Search · ICML 2022
Machine learning › Reinforcement learning
regret minimization
0.612022
Modeling Strong and Human-Like Gameplay with KL-Regularized Search · ICML 2022
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
search-based planning
0.612022
Modeling Strong and Human-Like Gameplay with KL-Regularized Search · ICML 2022
Algorithmic game theory and mechanism design › equilibrium computation
double oracle algorithm
0.512021
No-Press Diplomacy from Scratch · NeurIPS 2021
Knowledge, reasoning and agents › Knowledge representation and reasoning
theory of mind
0.212022
Self-Explaining Deviations for Coordination · NeurIPS 2022
Machine learning › Reinforcement learning › multi-agent reinforcement learning
self-play
0.112021
No-Press Diplomacy from Scratch · NeurIPS 2021

Methods — techniques the papers use, named apart from their topics

regularized equilibrium computation · 1.3update equivalence · 0.8mirror descent · 0.8magnetic mirror descent · 0.8reinforcement learning · 0.7planning · 0.7diversity optimization · 0.7adversarial policy generation · 0.7fine-tuning · 0.6alphazero-style search · 0.6value iteration · 0.5policy proposal network · 0.5equilibrium search · 0.5double oracle · 0.5
YearPublicationVenuePosition
2024 The Update-Equivalence Framework for Decision-Time Planning
abstract
The process of revising (or constructing) a policy at execution time---known as decision-time planning---has been key to achieving superhuman performance in perfect-information games like chess and Go. A recent line of work has extended decision-time planning to imperfect-information games, leading to superhuman performance in poker. However, these methods involve solving subgames whose sizes grow quickly in the amount of non-public information, making them unhelpful when the amount of non-public information is large. Motivated by this issue, we introduce an alternative framework for decision-time planning that is not based on solving subgames, but rather on update equivalence. In this update-equivalence framework, decision-time planning algorithms replicate the updates of last-iterate algorithms, which need not rely on public information. This facilitates scalability to games with large amounts of non-public information. Using this framework, we derive a provably sound search algorithm for fully cooperative games based on mirror descent and a search algorithm for adversarial games based on magnetic mirror descent. We validate the performance of these algorithms in cooperative and adversarial domains, notably in Hanabi, the standard benchmark for search in fully cooperative imperfect-information games. Here, our mirror descent approach exceeds or matches the performance of public information-based search while using two orders of magnitude less search time. This is the first instance of a non-public-information-based algorithm outperforming public-information-based approaches in a domain they have historically dominated.
Samuel Sokota, Gabriele Farina, David J. Wu 0002, Hengyuan Hu, Kevin A. Wang, J. Zico Kolter, Noam Brown
ICLR3
2023 Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning
Anton Bakhtin, David J. Wu 0002, Adam Lerer, Jonathan Gray, Athul Paul Jacob, Gabriele Farina, Alexander H. Miller, Noam Brown
ICLR2
2023 Adversarial Diversity in Hanabi
Brandon Cui, Andrei Lupu, Samuel Sokota, Hengyuan Hu, David J. Wu 0002, Jakob N. Foerster
ICLR5
2023 Abstracting Imperfect Information Away from Two-Player Zero-Sum Games
abstract
In their seminal work, Nayyar et al. (2013) showed that imperfect information can be abstracted away from common-payoff games by having players publicly announce their policies as they play. This insight underpins sound solvers and decision-time planning algorithms for common-payoff games. Unfortunately, a naive application of the same insight to two-player zero-sum games fails because Nash equilibria of the game with public policy announcements may not correspond to Nash equilibria of the original game. As a consequence, existing sound decision-time planning algorithms require complicated additional mechanisms that have unappealing properties. The main contribution of this work is showing that certain regularized equilibria do not possess the aforementioned non-correspondence problem—thus, computing them can be treated as perfect-information problems. Because these regularized equilibria can be made arbitrarily close to Nash equilibria, our result opens the door to a new perspective to solving two-player zero-sum games and yields a simplified framework for decision-time planning in two-player zero-sum games, void of the unappealing properties that plague existing decision-time planning approaches.
Samuel Sokota, Ryan D'Orazio, Chun Kai Ling, David J. Wu 0002, J. Zico Kolter, Noam Brown
ICML4
2022 A Fine-Tuning Approach to Belief State Modeling
Samuel Sokota, Hengyuan Hu, David J. Wu 0002, J. Zico Kolter, Jakob N. Foerster, Noam Brown
ICLR3
2022 Modeling Strong and Human-Like Gameplay with KL-Regularized Search
abstract
We consider the task of accurately modeling strong human policies in multi-agent decision-making problems, given examples of human behavior. Imitation learning is effective at predicting human actions but may not match the strength of expert humans (e.g., by sometimes committing blunders), while self-play learning and search techniques such as AlphaZero lead to strong performance but may produce policies that differ markedly from human behavior. In chess and Go, we show that regularized search algorithms that penalize KL divergence from an imitation-learned policy yield higher prediction accuracy of strong humans and better performance than imitation learning alone. We then introduce a novel regret minimization algorithm that is regularized based on the KL divergence from an imitation-learned policy, and show that using this algorithm for search in no-press Diplomacy yields a policy that matches the human prediction accuracy of imitation learning while being substantially stronger.
Athul Paul Jacob, David J. Wu 0002, Gabriele Farina, Adam Lerer, Hengyuan Hu, Anton Bakhtin, Jacob Andreas, Noam Brown
ICML2
2022 Self-Explaining Deviations for Coordination
abstract
Fully cooperative, partially observable multi-agent problems are ubiquitous in the real world. In this paper, we focus on a specific subclass of coordination problems in which humans are able to discover self-explaining deviations (SEDs). SEDs are actions that deviate from the common understanding of what reasonable behavior would be in normal circumstances. They are taken with the intention of causing another agent or other agents to realize, using theory of mind, that the circumstance must be abnormal. We motivate this idea with a real world example and formalize its definition. Next, we introduce an algorithm for improvement maximizing SEDs (IMPROVISED). Lastly, we evaluate IMPROVISED both in an illustrative toy setting and the popular benchmark setting Hanabi, where we show that it can produce so called finesse plays.
Hengyuan Hu, Samuel Sokota, David J. Wu 0002, Anton Bakhtin, Andrei Lupu, Brandon Cui, Jakob N. Foerster
NeurIPS3
2021 No-Press Diplomacy from Scratch
abstract
Prior AI successes in complex games have largely focused on settings with at most hundreds of actions at each decision point. In contrast, Diplomacy is a game with more than 10^20 possible actions per turn. Previous attempts to address games with large branching factors, such as Diplomacy, StarCraft, and Dota, used human data to bootstrap the policy or used handcrafted reward shaping. In this paper, we describe an algorithm for action exploration and equilibrium approximation in games with combinatorial action spaces. This algorithm simultaneously performs value iteration while learning a policy proposal network. A double oracle step is used to explore additional actions to add to the policy proposals. At each state, the target state value and policy for the model training are computed via an equilibrium search procedure. Using this algorithm, we train an agent, DORA, completely from scratch for a popular two-player variant of Diplomacy and show that it achieves superhuman performance. Additionally, we extend our methods to full-scale no-press Diplomacy and for the first time train an agent from scratch with no human data. We present evidence that this agent plays a strategy that is incompatible with human-data bootstrapped agents. This presents the first strong evidence of multiple equilibria in Diplomacy and suggests that self play alone may be insufficient for achieving superhuman performance in Diplomacy.
Anton Bakhtin, David J. Wu 0002, Adam Lerer, Noam Brown
NeurIPS2