VLDB 2026 Research / reviewers in the wild / expert
Youpeng Zhao 0001
dblp:259/5839-1
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2025
0000-0002-4610-3545ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CraftFactory: A Conditioned Control Policy Benchmark for Compositional GeneralizationabstractHumans excel at understanding and reasoning about novel, compositionally structured knowledge, largely due to their capacity for compositional generalization—a cognitive skill that has recently been validated in structured neural networks. However, most existing research has focused primarily on semantic translation within canonical language environments, often neglecting the explicit connection to compositional generalization behavior. In contrast, humans typically demonstrate this ability through interaction with their environments rather than solely through internal reasoning. To address this gap, we propose CraftFactory, a benchmark designed for evaluating compositional generalization in an interactive control environment. This benchmark introduces a new challenge for testing compositional generalization in a more realistic and comprehensive manner. CraftFactory stands out due to three key features: (1) it offers an open-ended interactive control environment with thousands of items and flexible actions; (2) it requires advanced compositional inference through various combinations and complex permutations of instructions; and (3) it evaluates compositional generalization intuitively through interactive behavior. By leveraging CraftFactory, we aim to promote the development of more advanced compositional generalization methods, thereby contributing to the broader field of general AI. Jinbing Hou, Youpeng Zhao 0001, Jian Zhao 0018 |
AAAI | 2 |
| 2025 | Generalizable agent modeling for agent collaboration-competition adaptation with multi-retrieval and dynamic generation
Yonggang Jin, Youpeng Zhao 0001, Zipeng Dai, Jian Zhao 0018, Liuyu Xiang, Junge Zhang, Zhaofeng He 0001 |
Neurocomputing | 4 |
| 2025 | CuDA2: An Approach for Incorporating Traitor Agents Into Cooperative Multiagent SystemsabstractCooperative multiagent reinforcement learning (CMARL) strategies are well known to be vulnerable to adversarial perturbations. Previous works on adversarial attacks have primarily focused on glass-box attacks that directly perturb the states or actions of victim agents, often in scenarios with a limited number of attacks. However, gaining complete access to victim agents in real-world environments is exceedingly difficult. To create more realistic adversarial attacks, we introduce a novel method that involves injecting traitor agents into the CMARL system. We model this problem as a traitor Markov decision process (TMDP), where traitors cannot directly attack the victim agents but can influence their formation or positioning through collisions. In TMDP, traitors are trained using the same MARL algorithm as the victim agents, with their reward function set as the negative of the victim agents' reward. Despite this, the training efficiency for traitors remains low because it is challenging for them to directly associate their actions with the victim agents' rewards. To address this issue, we propose the curiosity-driven adversarial attack (CuDA2) framework. CuDA2 enhances the efficiency and aggressiveness of attacks on the specified victim agents' policies while maintaining the optimal policy invariance of the traitors. Specifically, we employ a pretrained random network distillation module, where the extra reward generated by the RND module encourages traitors to explore states unencountered by the victim agents. Extensive experiments on various scenarios from SMAC demonstrate that our CuDA2 framework offers comparable or superior adversarial attack capabilities compared to other baselines. Zhen Chen 0025, Yong Liao 0003, Youpeng Zhao 0001, Zipeng Dai, Jian Zhao 0018 |
IEEE Trans. Games | 3 |
| 2025 | Mini Honor of Kings: A Lightweight Environment for Multiagent Reinforcement LearningabstractGames are widely used as research environments for multiagent reinforcement learning (MARL), but they pose three significant challenges: limited customization, high computational demands, and oversimplification. To address these issues, we introduce the first publicly available map editor for the popular mobile gameHonor of Kingsand design a lightweight environment,Mini Honor of Kings(Mini HoK), for researchers to conduct experiments. Mini HoK is highly efficient, allowing experiments to be run on personal PCs or laptops while still presenting sufficient challenges for existing MARL algorithms. We have tested our environment on common MARL algorithms and demonstrated that these algorithms have yet to surpass the performance of rule based policies, indicating that current MARL methods are not able to solve this environment. This facilitates the dissemination and advancement of MARL methods within the research community. In addition, we hope that more researchers will leverage theHonor of Kingsmap editor to develop innovative and scientifically valuable new maps. Lin Liu 0016, Jian Zhao 0018, Zhengtao Cao, Youpeng Zhao 0001, Zhenbin Ye, Zhaofeng He 0001, Houqiang Li, Xia Lin, Lanxiao Huang |
IEEE Trans. Games | 5 |
| 2024 | DanZero+: Dominating the GuanDan Game Through Reinforcement LearningabstractRecent advancements have propelled artificial intelligence (AI) to showcase expertise in intricate card games, such asMahjong,DouDizhu, andTexas Hold'em. In this work, we aim to develop an AI program for an exceptionally complex and popular card game calledGuanDan. This game involves four players engaging in both competitive and cooperative play throughout a long process, posing great challenges for AI due to its expansive state and action space, long episode length, and complex rules. Employing reinforcement learning techniques, specifically deep Monte Carlo, and a distributed training framework, we first put forward an AI program named DanZero. Evaluation against baseline AI programs based on heuristic rules highlights the outstanding performance of our bot. Besides, in order to further enhance the AI's capabilities, we apply proximal policy optimization toGuanDanon the basis of Danzero. To address the challenges arising from the huge action space, which will significantly impact the performance of policy-based algorithms, we adopt the pretrained model to compress the action space and integrate action features into the model to bolster its generalization capabilities. Using these techniques, we manage to obtain a newGuanDanAI program DanZero+, which achieves a superior performance compared to DanZero. Youpeng Zhao 0001, Yudong Lu, Jian Zhao 0018, Wengang Zhou 0001, Houqiang Li |
IEEE Trans. Games | 1 |
| 2024 | MCMARL: Parameterizing Value Function via Mixture of Categorical Distributions for Multi-Agent Reinforcement LearningabstractIn cooperative multi-agent tasks, a team of agents jointly interact with an environment by taking actions, receiving a team reward and observing the next state. During the interactions, the uncertainty of environment and reward will inevitably induce stochasticity in the long-term returns and the randomness can be exacerbated with the increasing number of agents. However, such randomness is ignored by most of the existing value-based multi-agent reinforcement learning (MARL) methods, which only model the expectation of Q-value for both individual agents and the team. Compared to using the expectations of the long-term returns, it is preferable to directly model the stochasticity by estimating the returns through distributions. With this motivation, this work proposes a novel value-based MARL framework from a distributional perspective,i.e., parameterizing value function viaMixture ofCategorical distributions for MARL. Specifically, we model both individual Q-values and global Q-value with categorical distribution. To integrate categorical distributions, we define five basic operations on the distribution, which allow the generalization of expected value function factorization methods (e.g., VDN and QMIX) to their MCMARL variants. We further prove that our MCMARL framework satisfiesDistributional-Individual-Global-Max(DIGM) principle with respect to the expectation of distribution, which guarantees the consistency between joint and individual greedy action selections in the global Q-value and individual Q-values. Empirically, we evaluate MCMARL on both a stochastic matrix game and a challenging set of StarCraft II micromanagement tasks, showing the efficacy of our framework. Jian Zhao 0018, Mingyu Yang 0003, Youpeng Zhao 0001, Xunhan Hu, Wengang Zhou 0001, Houqiang Li |
IEEE Trans. Games | 3 |
| 2024 | Full DouZero+: Improving DouDizhu AI by Opponent Modeling, Coach-Guided Training and Bidding LearningabstractWith the development of deep reinforcement learning (DRL), much progress in various perfect and imperfect information games has been achieved. Among these games, DouDizhu, a popular card game in China, poses great challenges because of the imperfect information, large state and action space as well as the cooperation issue. In this paper, we put forward an AI system for this game, which adopts opponent modeling and coach-guided training to help agents make better decisions when playing cards. Besides, we take the bidding phase of DouDizhu into consideration, which is usually ignored by existing works, and train a bidding network using Monte-Carlo simulation. As a result, we achieve a full version of our AI system that is applicable to real-world competitions. We conduct extensive experiments to evaluate the effectiveness of the three techniques adopted in our method and demonstrate the superior performance of our AI over the state-of-the-art DouDizhu AI, i.e., DouZero. We upload our AI systems, one is bidding-free and the other is equipped with a bidding network, to Botzone platform and they both rank the first among over 400 and 250 AI programs on the two corresponding leaderboards, respectively. Our codes are available athttps://github.com/submit-paper/Doudizhu_plus. Youpeng Zhao 0001, Jian Zhao 0018, Xunhan Hu, Wengang Zhou 0001, Houqiang Li |
IEEE Trans. Games | 1 |
| 2023 | DanZero: Mastering GuanDan Game with Reinforcement LearningabstractThe use of artificial intelligence (AI) in card games has been a widely researched topic in the field of AI for an extended period. Recent advancements have led to AI programs exhibiting expert-level gameplay in complex card games such as Mahjong, DouDizhu, and Texas Hold’em. This paper aims to develop an AI program, named DanZero, for GuanDan, an exceptionally complex card game that involves four players competing and cooperating in a long process to upgrade their level quickly. Developing AI for GuanDan is challenging due to its large state and action space, long episode length, and uncertainty in the number of players. To address these challenges, we propose DanZero, the first AI program for GuanDan, that employs reinforcement learning using a distributed framework for training. Our framework consists of two processes: the Actor Process and the Learner Process. In the Actor Process, we design state features and generate samples through agents’ self-play. In the Learner Process, we update the model using the Deep Monte-Carlo Method. We trained DanZero for 30 days, utilizing 160 CPUs and 1 GPU to develop the program successfully. We compared DanZero’s performance with eight baseline AI programs based on heuristic rules, and our results indicate DanZero’s exceptional performance. We further tested DanZero with human players and demonstrated its ability to perform at a human level. The code for DanZero can be found in the supplementary material. Yudong Lu, Jian Zhao 0018, Youpeng Zhao 0001, Wengang Zhou 0001, Houqiang Li |
CoG | 3 |
| 2023 | Q-SAT: Value Factorization with Self-Attention for Deep Multi-Agent Reinforcement LearningabstractIn many real-world tasks, a team of agents learn to cooperate with each other under the setting of partial observability and communication constraints, where value factorization has been demonstrated as an effective solution. In a multi-agent system, it's important to capture the inter-connection between agents and push agents to consider more of the relevant teammates. Motivated by the success of self-attention in natural language processing and computer vision, we propose a novel value factorization mechanism, called Q-function Self ATtention (Q-SAT). It models the pairwise action-value functions and connection coefficient between agent pairs explicitly, and pays more attention to the interrelated agents when making decisions. Satisfying the IGM principle, Q-SAT introduces the self-attention into value factorization network. This attention mechanism enables more effective and efficient learning in complex multi-agent environments. Q-SAT can be viewed as a basic building block and is ready to be applied to existing value factorization methods. The experimental results show that Q-SAT captures the connection relationship between agents and significantly improves the learning performance on the challenging StarCraft II micromanagement task and Google Research Football task. Xunhan Hu, Jian Zhao 0018, Youpeng Zhao 0001, Wengang Zhou 0001, Houqiang Li |
IJCNN | 3 |
| 2023 | Multi-Agent First Order Constrained Optimization in Policy SpaceabstractIn the realm of multi-agent reinforcement learning (MARL), achieving high performance is crucial for a successful multi-agent system.
Meanwhile, the ability to avoid unsafe actions is becoming an urgent and imperative problem to solve for real-life applications.
Whereas, it is still challenging to develop a safety-aware method for multi-agent systems in MARL. In this work, we introduce a novel approach called Multi-Agent First Order Constrained Optimization in Policy Space (MAFOCOPS), which effectively addresses the dual objectives of attaining satisfactory performance and enforcing safety constraints. Using data generated from the current policy, MAFOCOPS first finds the optimal update policy by solving a constrained optimization problem in the nonparameterized policy space. Then, the update policy is projected back into the parametric policy space to achieve a feasible policy. Notably, our method is first-order in nature, ensuring the ease of implementation, and exhibits an approximate upper bound on the worst-case constraint violation. Empirical results show that our approach achieves remarkable performance while satisfying safe constraints on several safe MARL benchmarks. Youpeng Zhao 0001, Yaodong Yang 0001, Zhenbo Lu, Wengang Zhou 0001, Houqiang Li |
NeurIPS | 1 |
| 2023 | Improving Deep Reinforcement Learning With Mirror LossabstractRecent years have witnessed the great breakthrough of deep reinforcement learning (DRL) in various artificial intelligence applications, but the training process needs a very large amount of samples and huge computational costs. To alleviate the low sample efficiency issue, one feasible solution is to improve the state representation learning. We uncover that the agents, trained by the original DRL algorithm, face severe performance degradation in mirrored game environments. As mirror symmetry is an important property of the environment, the poor performance in the mirrored situation indicates that the agents are not fully aware of the essence of the environment. In order to handle this problem and make use of the property to attain better state representation, we propose a mirror loss, which serves as an auxiliary module to bring mirror symmetry representation to the DRL agent. It is model-agnostic and prompts the DRL agent to make logically consistent mirrored actions in the mirrored environment. We conduct experiments on OpenAI Gym Atari environments and a more complex reinforcement learning task, Mahjong AI, and the results demonstrate the efficiency and versatility of our method. Jian Zhao 0018, Weide Shu, Youpeng Zhao 0001, Wengang Zhou 0001, Houqiang Li |
IEEE Trans. Games | 3 |
| 2022 | DouZero+: Improving DouDizhu AI by Opponent Modeling and Coach-guided LearningabstractRecent years have witnessed the great breakthrough of deep reinforcement learning (DRL) in various perfect and imperfect information games. Among these games, DouDizhu, a popular card game in China, is very challenging due to the imperfect information, large state and action space as well as elements of collaboration. Recently, a DouDizhu AI system called DouZero has been proposed. Trained using traditional Monte Carlo method with deep neural networks and self-play procedure without the abstraction of human prior knowledge, DouZero has achieved the best performance among all the existing DouDizhu AI programs. In this work, we propose to enhance DouZero by introducing opponent modeling into DouZero. Besides, we propose a novel coach network to further boost the performance of DouZero and accelerate its training process. With the integration of the above two techniques into DouZero, our DouDizhu AI system achieves better performance and ranks top in the Botzone leaderboard among more than 400 AI agents, including DouZero. Youpeng Zhao 0001, Jian Zhao 0018, Xunhan Hu, Wengang Zhou 0001, Houqiang Li |
CoG | 1 |
| 2022 | Coach-assisted multi-agent reinforcement learning framework for unexpected crashed agentsabstractMulti-agent reinforcement learning is difficult to apply in practice, partially because of the gap between simulated and real-world scenarios. One reason for the gap is that simulated systems always assume that agents can work normally all the time, while in practice, one or more agents may unexpectedly “crash” during the coordination process due to inevitable hardware or software failures. Such crashes destroy the cooperation among agents and lead to performance degradation. In this work, we present a formal conceptualization of a cooperative multi-agent reinforcement learning system with unexpected crashes. To enhance the robustness of the system to crashes, we propose a coach-assisted multi-agent reinforcement learning framework that introduces a virtual coach agent to adjust the crash rate during training. We have designed three coaching strategies (fixed crash rate, curriculum learning, and adaptive crash rate) and a re-sampling strategy for our coach agent. To our knowledge, this work is the first to study unexpected crashes in a multi-agent system. Extensive experiments on grid-world and StarCraft II micromanagement tasks demonstrate the efficacy of the adaptive strategy compared with the fixed crash rate strategy and curriculum learning strategy. The ablation study further illustrates the effectiveness of our re-sampling strategy. Jian Zhao 0018, Youpeng Zhao 0001, Weixun Wang, Mingyu Yang 0003, Xunhan Hu, Wengang Zhou 0001, Jianye Hao, Houqiang Li |
Frontiers Inf. Technol. Electron. Eng. | 2 |