Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zelai Xu

dblp:314/3288 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0001-5578-199XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 54% Multi-agent systems · 20% Language models and text generation · 10%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-agent reinforcement learning
3.142025
Learning Global Nash Equilibrium in Team Competitive Games with Generalized Fictitious Cross-Play · J. Mach. Learn. Res. 2025
VolleyBots: A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic Play · NeurIPS 2025
Accelerate Multi-Agent Reinforcement Learning in Zero-Sum Games with Subgame Curriculum Learning · AAAI 2024
Natural language and speech › Language models and text generation
LLM agents
1.622025
Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy Optimization · ICML 2025
Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game · ICML 2024
Robotics › Legged, aerial and field robots
aerial robots
0.912025
VolleyBots: A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic Play · NeurIPS 2025
Machine learning › Reinforcement learning › multi-agent reinforcement learning
best response computation
0.912025
Learning Global Nash Equilibrium in Team Competitive Games with Generalized Fictitious Cross-Play · J. Mach. Learn. Res. 2025
Machine learning › Reinforcement learning › hierarchical reinforcement learning
hierarchical policy
0.912025
VolleyBots: A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic Play · NeurIPS 2025
Knowledge, reasoning and agents › Multi-agent systems › game theory
nash equilibrium
0.912025
Learning Global Nash Equilibrium in Team Competitive Games with Generalized Fictitious Cross-Play · J. Mach. Learn. Res. 2025
Machine learning › Reinforcement learning
policy optimization
0.912025
Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy Optimization · ICML 2025
Machine learning › Reinforcement learning › multi-agent reinforcement learning › equilibrium learning
policy space response oracle
0.912025
Learning Global Nash Equilibrium in Team Competitive Games with Generalized Fictitious Cross-Play · J. Mach. Learn. Res. 2025
Robotics › Motion planning and robot control
robot control
0.912025
VolleyBots: A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic Play · NeurIPS 2025
Machine learning › Reinforcement learning › multi-agent reinforcement learning
self-play
0.912025
Learning Global Nash Equilibrium in Team Competitive Games with Generalized Fictitious Cross-Play · J. Mach. Learn. Res. 2025
Machine learning › Learning paradigms
curriculum learning
0.812024
Accelerate Multi-Agent Reinforcement Learning in Zero-Sum Games with Subgame Curriculum Learning · AAAI 2024
Knowledge, reasoning and agents › Multi-agent systems › imperfect information games
social deduction game
0.812024
Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game · ICML 2024
Knowledge, reasoning and agents › Multi-agent systems › game theory
zero-sum games
0.812024
Accelerate Multi-Agent Reinforcement Learning in Zero-Sum Games with Subgame Curriculum Learning · AAAI 2024
Machine learning › Reinforcement learning › policy optimization
policy gradient
0.612022
Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement Learning · ICML 2022
Machine learning › Reinforcement learning › multi-agent reinforcement learning › value-based multi-agent reinforcement learning
value decomposition
0.612022
Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement Learning · ICML 2022
Machine learning › Reinforcement learning
strategic decision-making
0.212024
Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game · ICML 2024

Methods — techniques the papers use, named apart from their topics

sim-to-real transfer · 0.9on-policy reinforcement learning · 0.9off-policy reinforcement learning · 0.9game-theoretic algorithm · 0.9game theory · 0.9exploitability analysis · 0.9direct preference optimization · 0.9counterfactual regret minimization · 0.9particle-based state sampler · 0.8nash equilibrium approximation · 0.8
YearPublicationVenuePosition
2025 Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy Optimization
abstract
Large language model (LLM) agents have recently demonstrated impressive capabilities in various domains like open-ended conversation and multi-step decision-making. However, it remains challenging for these agents to solve strategic language games, such as Werewolf, which demand both strategic decision-making and free-form language interactions. Existing LLM agents often suffer from intrinsic bias in their action distributions and limited exploration of the unbounded text action space, resulting in suboptimal performance. To address these challenges, we propose Latent Space Policy Optimization (LSPO), an iterative framework that combines game-theoretic methods with LLM fine-tuning to build strategic language agents. LSPO leverages the observation that while the language space is combinatorially large, the underlying strategy space is relatively compact. We first map free-form utterances into a finite latent strategy space, yielding an abstracted extensive-form game. Then we apply game-theoretic methods like Counterfactual Regret Minimization (CFR) to optimize the policy in the latent space. Finally, we fine-tune the LLM via Direct Preference Optimization (DPO) to align with the learned policy. By iteratively alternating between these steps, our LSPO agents progressively enhance both strategic reasoning and language communication. Experiment on the Werewolf game shows that our agents iteratively expand the strategy space with improving performance and outperform existing Werewolf agents, underscoring their effectiveness in free-form language games with strategic interactions.
Zelai Xu, Wanjun Gu, Chao Yu 0005, Yi Wu 0013, Yu Wang 0002
ICML1
2025 VolleyBots: A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic Play
abstract
Robot sports, characterized by well-defined objectives, explicit rules, and dynamic interactions, present ideal scenarios for demonstrating embodied intelligence. In this paper, we present VolleyBots, a novel robot sports testbed where multiple drones cooperate and compete in the sport of volleyball under physical dynamics. VolleyBots integrates three features within a unified platform: competitive and cooperative gameplay, turn-based interaction structure, and agile 3D maneuvering.These intertwined features yield a complex problem combining motion control and strategic play, with no available expert demonstrations.We provide a comprehensive suite of tasks ranging from single-drone drills to multi-drone cooperative and competitive tasks, accompanied by baseline evaluations of representative reinforcement learning (RL), multi-agent reinforcement learning (MARL) and game-theoretic algorithms. Simulation results show that on-policy RL methods outperform off-policy methods in single-agent tasks, but both approaches struggle in complex tasks that combine motion control and strategic play.We additionally design a hierarchical policy which achieves 69.5% win rate against the strongest baseline in the 3 vs 3 task, demonstrating its potential for tackling the complex interplay between low-level control and high-level strategy.To highlight VolleyBots’ sim-to-real potential, we further demonstrate the zero-shot deployment of a policy trained entirely in simulation on real-world drones.
Zelai Xu, Ruize Zhang 0001, Chao Yu 0005, Huining Yuan 0002, Xiangmin Yi, Shilong Ji, Chuqi Wang, Wenbo Ding 0001, Xinlei Chen, Yu Wang 0002
NeurIPS1
2025 Learning Global Nash Equilibrium in Team Competitive Games with Generalized Fictitious Cross-Play
abstract
Self-play (SP) is a popular multi-agent reinforcement learning framework for competitive games. Despite the empirical success, the theoretical properties of SP are limited to two-player settings. For team competitive games where two teams of cooperative agents compete with each other, we show a counter-example where SP cannot converge to a global Nash equilibrium (NE) with high probability. Policy-Space Response Oracles (PSRO) is an alternative framework that finds NEs by iteratively learning the best response (BR) to previous policies. PSRO can be directly extended to team competitive games with unchanged convergence properties by learning team BRs, but its repeated training from scratch makes it hard to scale to complex games. In this work, we propose Generalized Fictitious Cross-Play (GFXP), a novel algorithm that inherits benefits from both frameworks. GFXP simultaneously trains an SP-based main policy and a counter population. The main policy is trained by fictitious self-play and cross-play against the counter population, while the counter policies are trained as the BRs to the main policy's checkpoints. We evaluate GFXP in matrix games and gridworld domains where GFXP achieves the lowest exploitabilities. We further conduct experiments in a challenging football game where GFXP defeats SOTA models with over 94% win rate.
Zelai Xu, Chao Yu 0005, Yancheng Liang, Yi Wu 0013, Yu Wang 0002
J. Mach. Learn. Res.1
2024 Accelerate Multi-Agent Reinforcement Learning in Zero-Sum Games with Subgame Curriculum Learning
abstract
Learning Nash equilibrium (NE) in complex zero-sum games with multi-agent reinforcement learning (MARL) can be extremely computationally expensive. Curriculum learning is an effective way to accelerate learning, but an under-explored dimension for generating a curriculum is the difficulty-to-learn of the subgames –games induced by starting from a specific state. In this work, we present a novel subgame curriculum learning framework for zero-sum games. It adopts an adaptive initial state distribution by resetting agents to some previously visited states where they can quickly learn to improve performance. Building upon this framework, we derive a subgame selection metric that approximates the squared distance to NE values and further adopt a particle-based state sampler for subgame generation. Integrating these techniques leads to our new algorithm, Subgame Automatic Curriculum Learning (SACL), which is a realization of the subgame curriculum learning framework. SACL can be combined with any MARL algorithm such as MAPPO. Experiments in the particle-world environment and Google Research Football environment show SACL produces much stronger policies than baselines. In the challenging hide-and-seek quadrant environment, SACL produces all four emergent stages and uses only half the samples of MAPPO with self-play. The project website is at https://sites.google.com/view/sacl-neurips.
Jiayu Chen 0005, Zelai Xu, Yunfei Li 0005, Chao Yu 0005, Jiaming Song, Huazhong Yang, Fei Fang 0001, Yu Wang 0002, Yi Wu 0013
AAAI2
2024 Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game
abstract
Agents built with large language models (LLMs) have shown great potential across a wide range of domains. However, in complex decision-making tasks, pure LLM-based agents tend to exhibit intrinsic bias in their choice of actions, which is inherited from the model’s training data and results in suboptimal performance. To develop strategic language agents, i.e., agents that generate flexible language actions and possess strong decision-making abilities, we propose a novel framework that powers LLM-based agents with reinforcement learning (RL). We consider Werewolf, a popular social deduction game, as a challenging testbed that emphasizes versatile communication and strategic gameplay. To mitigate the intrinsic bias in language actions, our agents use an LLM to perform deductive reasoning and generate a diverse set of action candidates. Then an RL policy trained to optimize the decision-making ability chooses an action from the candidates to play in the game. Extensive experiments show that our agents overcome the intrinsic bias and outperform existing LLM-based agents in the Werewolf game. We also conduct human-agent experiments and find that our agents achieve human-level performance and demonstrate strong strategic play.
Zelai Xu, Chao Yu 0005, Fei Fang 0001, Yu Wang 0002, Yi Wu 0013
ICML1
2022 Texture BERT for Cross-modal Texture Image Retrieval
abstract
We propose Texture BERT, a model describing visual attributes of texture using natural language. To capture the rich details in texture images, we propose a group-wise compact bilinear pooling method, which represents the texture image by a set of visual patterns. The similarity between the texture image and the corresponding language description is determined by the cross-matching between the set of visual patterns from the texture image and the set of word features from the language description. We also exploit the self-attention transformer layers to provide the cross-modal context and enhance the effectiveness of matching. Our efforts achieve state-of-the-art accuracy on both text retrieval and image retrieval tasks, demonstrating the effectiveness of the proposed Texture BERT model in describing texture through natural language.
Zelai Xu, Ping Li 0001
CIKM1
2022 Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement Learning
abstract
Many advances in cooperative multi-agent reinforcement learning (MARL) are based on two common design principles: value decomposition and parameter sharing. A typical MARL algorithm of this fashion decomposes a centralized Q-function into local Q-networks with parameters shared across agents. Such an algorithmic paradigm enables centralized training and decentralized execution (CTDE) and leads to efficient learning in practice. Despite all the advantages, we revisit these two principles and show that in certain scenarios, e.g., environments with a highly multi-modal reward landscape, value decomposition, and parameter sharing can be problematic and lead to undesired outcomes. In contrast, policy gradient (PG) methods with individual policies provably converge to an optimal solution in these cases, which partially supports some recent empirical observations that PG can be effective in many MARL testbeds. Inspired by our theoretical analysis, we present practical suggestions on implementing multi-agent PG algorithms for either high rewards or diverse emergent behaviors and empirically validate our findings on a variety of domains, ranging from the simplified matrix and grid-world games to complex benchmarks such as StarCraft Multi-Agent Challenge and Google Research Football. We hope our insights could benefit the community towards developing more general and more powerful MARL algorithms.
Chao Yu 0005, Zelai Xu, Yi Wu 0013
ICML3