EDBT 2026 Demo / reviewers in the wild / expert
Zelai Xu
dblp:314/3288
· DBLP profile ↗
7ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0001-5578-199XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Reinforcement learning · 54% Multi-agent systems · 20% Language models and text generation · 10% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
3.1 | 4 | 2025 | Learning Global Nash Equilibrium in Team Competitive Games with Generalized Fictitious Cross-Play · J. Mach. Learn. Res. 2025 VolleyBots: A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic Play · NeurIPS 2025 Accelerate Multi-Agent Reinforcement Learning in Zero-Sum Games with Subgame Curriculum Learning · AAAI 2024 |
Natural language and speech › Language models and text generation
LLM agents |
1.6 | 2 | 2025 | Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy Optimization · ICML 2025 Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game · ICML 2024 |
Robotics › Legged, aerial and field robots
aerial robots |
0.9 | 1 | 2025 | VolleyBots: A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic Play · NeurIPS 2025 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
best response computation |
0.9 | 1 | 2025 | Learning Global Nash Equilibrium in Team Competitive Games with Generalized Fictitious Cross-Play · J. Mach. Learn. Res. 2025 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
hierarchical policy |
0.9 | 1 | 2025 | VolleyBots: A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic Play · NeurIPS 2025 |
Knowledge, reasoning and agents › Multi-agent systems › game theory
nash equilibrium |
0.9 | 1 | 2025 | Learning Global Nash Equilibrium in Team Competitive Games with Generalized Fictitious Cross-Play · J. Mach. Learn. Res. 2025 |
Machine learning › Reinforcement learning
policy optimization |
0.9 | 1 | 2025 | Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy Optimization · ICML 2025 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › equilibrium learning
policy space response oracle |
0.9 | 1 | 2025 | Learning Global Nash Equilibrium in Team Competitive Games with Generalized Fictitious Cross-Play · J. Mach. Learn. Res. 2025 |
Robotics › Motion planning and robot control
robot control |
0.9 | 1 | 2025 | VolleyBots: A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic Play · NeurIPS 2025 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
self-play |
0.9 | 1 | 2025 | Learning Global Nash Equilibrium in Team Competitive Games with Generalized Fictitious Cross-Play · J. Mach. Learn. Res. 2025 |
Machine learning › Learning paradigms
curriculum learning |
0.8 | 1 | 2024 | Accelerate Multi-Agent Reinforcement Learning in Zero-Sum Games with Subgame Curriculum Learning · AAAI 2024 |
Knowledge, reasoning and agents › Multi-agent systems › imperfect information games
social deduction game |
0.8 | 1 | 2024 | Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game · ICML 2024 |
Knowledge, reasoning and agents › Multi-agent systems › game theory
zero-sum games |
0.8 | 1 | 2024 | Accelerate Multi-Agent Reinforcement Learning in Zero-Sum Games with Subgame Curriculum Learning · AAAI 2024 |
Machine learning › Reinforcement learning › policy optimization
policy gradient |
0.6 | 1 | 2022 | Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement Learning · ICML 2022 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › value-based multi-agent reinforcement learning
value decomposition |
0.6 | 1 | 2022 | Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement Learning · ICML 2022 |
Machine learning › Reinforcement learning
strategic decision-making |
0.2 | 1 | 2024 | Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
sim-to-real transfer · 0.9on-policy reinforcement learning · 0.9off-policy reinforcement learning · 0.9game-theoretic algorithm · 0.9game theory · 0.9exploitability analysis · 0.9direct preference optimization · 0.9counterfactual regret minimization · 0.9particle-based state sampler · 0.8nash equilibrium approximation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy OptimizationabstractLarge language model (LLM) agents have recently demonstrated impressive capabilities in various domains like open-ended conversation and multi-step decision-making. However, it remains challenging for these agents to solve strategic language games, such as Werewolf, which demand both strategic decision-making and free-form language interactions. Existing LLM agents often suffer from intrinsic bias in their action distributions and limited exploration of the unbounded text action space, resulting in suboptimal performance. To address these challenges, we propose Latent Space Policy Optimization (LSPO), an iterative framework that combines game-theoretic methods with LLM fine-tuning to build strategic language agents. LSPO leverages the observation that while the language space is combinatorially large, the underlying strategy space is relatively compact. We first map free-form utterances into a finite latent strategy space, yielding an abstracted extensive-form game. Then we apply game-theoretic methods like Counterfactual Regret Minimization (CFR) to optimize the policy in the latent space. Finally, we fine-tune the LLM via Direct Preference Optimization (DPO) to align with the learned policy. By iteratively alternating between these steps, our LSPO agents progressively enhance both strategic reasoning and language communication. Experiment on the Werewolf game shows that our agents iteratively expand the strategy space with improving performance and outperform existing Werewolf agents, underscoring their effectiveness in free-form language games with strategic interactions. Zelai Xu, Wanjun Gu, Chao Yu 0005, Yi Wu 0013, Yu Wang 0002 |
ICML | 1 |
| 2025 | VolleyBots: A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic PlayabstractRobot sports, characterized by well-defined objectives, explicit rules, and dynamic interactions, present ideal scenarios for demonstrating embodied intelligence. In this paper, we present VolleyBots, a novel robot sports testbed where multiple drones cooperate and compete in the sport of volleyball under physical dynamics. VolleyBots integrates three features within a unified platform: competitive and cooperative gameplay, turn-based interaction structure, and agile 3D maneuvering.These intertwined features yield a complex problem combining motion control and strategic play, with no available expert demonstrations.We provide a comprehensive suite of tasks ranging from single-drone drills to multi-drone cooperative and competitive tasks, accompanied by baseline evaluations of representative reinforcement learning (RL), multi-agent reinforcement learning (MARL) and game-theoretic algorithms. Simulation results show that on-policy RL methods outperform off-policy methods in single-agent tasks, but both approaches struggle in complex tasks that combine motion control and strategic play.We additionally design a hierarchical policy which achieves 69.5% win rate against the strongest baseline in the 3 vs 3 task, demonstrating its potential for tackling the complex interplay between low-level control and high-level strategy.To highlight VolleyBots’ sim-to-real potential, we further demonstrate the zero-shot deployment of a policy trained entirely in simulation on real-world drones. Zelai Xu, Ruize Zhang 0001, Chao Yu 0005, Huining Yuan 0002, Xiangmin Yi, Shilong Ji, Chuqi Wang, Wenbo Ding 0001, Xinlei Chen, Yu Wang 0002 |
NeurIPS | 1 |
| 2025 | Learning Global Nash Equilibrium in Team Competitive Games with Generalized Fictitious Cross-PlayabstractSelf-play (SP) is a popular multi-agent reinforcement learning framework for competitive games. Despite the empirical success, the theoretical properties of SP are limited to two-player settings. For team competitive games where two teams of cooperative agents compete with each other, we show a counter-example where SP cannot converge to a global Nash equilibrium (NE) with high probability. Policy-Space Response Oracles (PSRO) is an alternative framework that finds NEs by iteratively learning the best response (BR) to previous policies. PSRO can be directly extended to team competitive games with unchanged convergence properties by learning team BRs, but its repeated training from scratch makes it hard to scale to complex games. In this work, we propose Generalized Fictitious Cross-Play (GFXP), a novel algorithm that inherits benefits from both frameworks. GFXP simultaneously trains an SP-based main policy and a counter population. The main policy is trained by fictitious self-play and cross-play against the counter population, while the counter policies are trained as the BRs to the main policy's checkpoints. We evaluate GFXP in matrix games and gridworld domains where GFXP achieves the lowest exploitabilities. We further conduct experiments in a challenging football game where GFXP defeats SOTA models with over 94% win rate. Zelai Xu, Chao Yu 0005, Yancheng Liang, Yi Wu 0013, Yu Wang 0002 |
J. Mach. Learn. Res. | 1 |
| 2024 | Accelerate Multi-Agent Reinforcement Learning in Zero-Sum Games with Subgame Curriculum LearningabstractLearning Nash equilibrium (NE) in complex zero-sum games with multi-agent reinforcement learning (MARL) can be extremely computationally expensive. Curriculum learning is an effective way to accelerate learning, but an under-explored dimension for generating a curriculum is the difficulty-to-learn of the subgames –games induced by starting from a specific state. In this work, we present a novel subgame curriculum learning framework for zero-sum games. It adopts an adaptive initial state distribution by resetting agents to some previously visited states where they can quickly learn to improve performance. Building upon this framework, we derive a subgame selection metric that approximates the squared distance to NE values and further adopt a particle-based state sampler for subgame generation. Integrating these techniques leads to our new algorithm, Subgame Automatic Curriculum Learning (SACL), which is a realization of the subgame curriculum learning framework. SACL can be combined with any MARL algorithm such as MAPPO. Experiments in the particle-world environment and Google Research Football environment show SACL produces much stronger policies than baselines. In the challenging hide-and-seek quadrant environment, SACL produces all four emergent stages and uses only half the samples of MAPPO with self-play. The project website is at https://sites.google.com/view/sacl-neurips. Jiayu Chen 0005, Zelai Xu, Yunfei Li 0005, Chao Yu 0005, Jiaming Song, Huazhong Yang, Fei Fang 0001, Yu Wang 0002, Yi Wu 0013 |
AAAI | 2 |
| 2024 | Language Agents with Reinforcement Learning for Strategic Play in the Werewolf GameabstractAgents built with large language models (LLMs) have shown great potential across a wide range of domains. However, in complex decision-making tasks, pure LLM-based agents tend to exhibit intrinsic bias in their choice of actions, which is inherited from the model’s training data and results in suboptimal performance. To develop strategic language agents, i.e., agents that generate flexible language actions and possess strong decision-making abilities, we propose a novel framework that powers LLM-based agents with reinforcement learning (RL). We consider Werewolf, a popular social deduction game, as a challenging testbed that emphasizes versatile communication and strategic gameplay. To mitigate the intrinsic bias in language actions, our agents use an LLM to perform deductive reasoning and generate a diverse set of action candidates. Then an RL policy trained to optimize the decision-making ability chooses an action from the candidates to play in the game. Extensive experiments show that our agents overcome the intrinsic bias and outperform existing LLM-based agents in the Werewolf game. We also conduct human-agent experiments and find that our agents achieve human-level performance and demonstrate strong strategic play. Zelai Xu, Chao Yu 0005, Fei Fang 0001, Yu Wang 0002, Yi Wu 0013 |
ICML | 1 |
| 2022 | Texture BERT for Cross-modal Texture Image RetrievalabstractWe propose Texture BERT, a model describing visual attributes of texture using natural language. To capture the rich details in texture images, we propose a group-wise compact bilinear pooling method, which represents the texture image by a set of visual patterns. The similarity between the texture image and the corresponding language description is determined by the cross-matching between the set of visual patterns from the texture image and the set of word features from the language description. We also exploit the self-attention transformer layers to provide the cross-modal context and enhance the effectiveness of matching. Our efforts achieve state-of-the-art accuracy on both text retrieval and image retrieval tasks, demonstrating the effectiveness of the proposed Texture BERT model in describing texture through natural language. Zelai Xu, Ping Li 0001 |
CIKM | 1 |
| 2022 | Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement LearningabstractMany advances in cooperative multi-agent reinforcement learning (MARL) are based on two common design principles: value decomposition and parameter sharing. A typical MARL algorithm of this fashion decomposes a centralized Q-function into local Q-networks with parameters shared across agents. Such an algorithmic paradigm enables centralized training and decentralized execution (CTDE) and leads to efficient learning in practice. Despite all the advantages, we revisit these two principles and show that in certain scenarios, e.g., environments with a highly multi-modal reward landscape, value decomposition, and parameter sharing can be problematic and lead to undesired outcomes. In contrast, policy gradient (PG) methods with individual policies provably converge to an optimal solution in these cases, which partially supports some recent empirical observations that PG can be effective in many MARL testbeds. Inspired by our theoretical analysis, we present practical suggestions on implementing multi-agent PG algorithms for either high rewards or diverse emergent behaviors and empirically validate our findings on a variety of domains, ranging from the simplified matrix and grid-world games to complex benchmarks such as StarCraft Multi-Agent Challenge and Google Research Football. We hope our insights could benefit the community towards developing more general and more powerful MARL algorithms. Chao Yu 0005, Zelai Xu, Yi Wu 0013 |
ICML | 3 |