EDBT 2026 Demo / reviewers in the wild / expert
Jian Zhao 0018
dblp:70/2932-18
· DBLP profile ↗
26ranked-venue papers
6as first author
25since 2021 · last 2026
0000-0003-4895-990XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Chaos to Cure: A Prefix Heuristics Guided Model-Agnostic Adaptive Detoxification FrameworkabstractThe impressive performance of large language models (LLMs) also brings inherent toxicity risks, prompting the need for effective detoxification to support responsible deployment. Prevailing methods generally follow an inflexible model-specific fashion, addressing only individual models or model families. Moreover, overlooking the underlying toxic risks involved in the input prefix can lead to toxic accumulation during autoregressive generation. Existing methods rely on external strong attribute interventions to address this issue, which further exacerbates contextual semantic inconsistencies and makes it difficult to balance toxicity efficacy and generation quality. To address these concerns, we propose a novel Model-Agnostic Adaptive Detoxification (MAAD) framework. To address accumulating toxicity, we present prefix heuristics that serve as contextual signals, guiding the base LLM toward safer generation. Along this line, we construct an antidote dataset to support a lightweight model, Detoxifier, which steers the base LLM to make in-scope and reliable detoxifying distribution adjustments while preserving fluency and contextual understanding. Designed as an easy-to-deploy module, Detoxifier requires a small amount of data and can be seamlessly applied to various base LLMs with one-off training. Since over-purifying often reduces diversity, we also propose a dynamic truncation method called CW-cutoff sampling to trade off language model quality and diversity. Extensive experiments demonstrate that MAAD strikes a better balance between detoxification effectiveness and generation quality, while also maintaining model utility. Yuhu Shang, Xiang Cheng 0003, Yimeng Ren 0001, Huijia Wu, Xuexiong Luo, Kangkang Lu 0002, Jian Zhao 0018, Zhaofeng He 0001 |
AAAI | 7 |
| 2025 | CraftFactory: A Conditioned Control Policy Benchmark for Compositional GeneralizationabstractHumans excel at understanding and reasoning about novel, compositionally structured knowledge, largely due to their capacity for compositional generalization—a cognitive skill that has recently been validated in structured neural networks. However, most existing research has focused primarily on semantic translation within canonical language environments, often neglecting the explicit connection to compositional generalization behavior. In contrast, humans typically demonstrate this ability through interaction with their environments rather than solely through internal reasoning. To address this gap, we propose CraftFactory, a benchmark designed for evaluating compositional generalization in an interactive control environment. This benchmark introduces a new challenge for testing compositional generalization in a more realistic and comprehensive manner. CraftFactory stands out due to three key features: (1) it offers an open-ended interactive control environment with thousands of items and flexible actions; (2) it requires advanced compositional inference through various combinations and complex permutations of instructions; and (3) it evaluates compositional generalization intuitively through interactive behavior. By leveraging CraftFactory, we aim to promote the development of more advanced compositional generalization methods, thereby contributing to the broader field of general AI. Jinbing Hou, Youpeng Zhao 0001, Jian Zhao 0018 |
AAAI | 3 |
| 2025 | Generalizable agent modeling for agent collaboration-competition adaptation with multi-retrieval and dynamic generation
Yonggang Jin, Youpeng Zhao 0001, Zipeng Dai, Jian Zhao 0018, Liuyu Xiang, Junge Zhang, Zhaofeng He 0001 |
Neurocomputing | 6 |
| 2025 | CuDA2: An Approach for Incorporating Traitor Agents Into Cooperative Multiagent SystemsabstractCooperative multiagent reinforcement learning (CMARL) strategies are well known to be vulnerable to adversarial perturbations. Previous works on adversarial attacks have primarily focused on glass-box attacks that directly perturb the states or actions of victim agents, often in scenarios with a limited number of attacks. However, gaining complete access to victim agents in real-world environments is exceedingly difficult. To create more realistic adversarial attacks, we introduce a novel method that involves injecting traitor agents into the CMARL system. We model this problem as a traitor Markov decision process (TMDP), where traitors cannot directly attack the victim agents but can influence their formation or positioning through collisions. In TMDP, traitors are trained using the same MARL algorithm as the victim agents, with their reward function set as the negative of the victim agents' reward. Despite this, the training efficiency for traitors remains low because it is challenging for them to directly associate their actions with the victim agents' rewards. To address this issue, we propose the curiosity-driven adversarial attack (CuDA2) framework. CuDA2 enhances the efficiency and aggressiveness of attacks on the specified victim agents' policies while maintaining the optimal policy invariance of the traitors. Specifically, we employ a pretrained random network distillation module, where the extra reward generated by the RND module encourages traitors to explore states unencountered by the victim agents. Extensive experiments on various scenarios from SMAC demonstrate that our CuDA2 framework offers comparable or superior adversarial attack capabilities compared to other baselines. Zhen Chen 0025, Yong Liao 0003, Youpeng Zhao 0001, Zipeng Dai, Jian Zhao 0018 |
IEEE Trans. Games | 5 |
| 2025 | Mini Honor of Kings: A Lightweight Environment for Multiagent Reinforcement LearningabstractGames are widely used as research environments for multiagent reinforcement learning (MARL), but they pose three significant challenges: limited customization, high computational demands, and oversimplification. To address these issues, we introduce the first publicly available map editor for the popular mobile gameHonor of Kingsand design a lightweight environment,Mini Honor of Kings(Mini HoK), for researchers to conduct experiments. Mini HoK is highly efficient, allowing experiments to be run on personal PCs or laptops while still presenting sufficient challenges for existing MARL algorithms. We have tested our environment on common MARL algorithms and demonstrated that these algorithms have yet to surpass the performance of rule based policies, indicating that current MARL methods are not able to solve this environment. This facilitates the dissemination and advancement of MARL methods within the research community. In addition, we hope that more researchers will leverage theHonor of Kingsmap editor to develop innovative and scientifically valuable new maps. Lin Liu 0016, Jian Zhao 0018, Zhengtao Cao, Youpeng Zhao 0001, Zhenbin Ye, Zhaofeng He 0001, Houqiang Li, Xia Lin, Lanxiao Huang |
IEEE Trans. Games | 2 |
| 2024 | Coordinate-aligned multi-camera collaboration for active multi-object tracking
Zeyu Fang, Jian Zhao 0018, Mingyu Yang 0003, Zhenbo Lu, Wengang Zhou 0001, Houqiang Li |
Multim. Syst. | 2 |
| 2024 | CTDS: Centralized Teacher With Decentralized Student for Multiagent Reinforcement LearningabstractDue to the partial observability and communication constraints in many multiagent reinforcement learning (MARL) tasks, centralized training with decentralized execution (CTDE) has become one of the most widely used MARL paradigms. In CTDE, centralized information is dedicated to learning the allocation of the team reward with a mixing network while the learning of individualQ-values is usually based on local observations. The insufficient utility of global observation will degrade performance in challenging environments. To this end, this work proposes a novel Centralized Teacher with a Decentralized Student (CTDS) framework, which consists of a teacher model and a student model. Specifically, the teacher model allocates the team reward by learning individualQ-values conditioned on global observation while the student model utilizes the partial observations to approximate theQ-values estimated by the teacher model. In this way, CTDS balances the full utilization of global observation during training and the feasibility of decentralized execution for online inference. Our CTDS framework is generic, which is ready to be applied upon existing CTDE methods to boost their performance. We conduct experiments on a challenging set ofStarCraft IImicromanagement tasks to test the effectiveness of our method and the results show that CTDS outperforms the existing value-based MARL methods. Jian Zhao 0018, Xunhan Hu, Mingyu Yang 0003, Wengang Zhou 0001, Jiangcheng Zhu, Houqiang Li |
IEEE Trans. Games | 1 |
| 2024 | DanZero+: Dominating the GuanDan Game Through Reinforcement LearningabstractRecent advancements have propelled artificial intelligence (AI) to showcase expertise in intricate card games, such asMahjong,DouDizhu, andTexas Hold'em. In this work, we aim to develop an AI program for an exceptionally complex and popular card game calledGuanDan. This game involves four players engaging in both competitive and cooperative play throughout a long process, posing great challenges for AI due to its expansive state and action space, long episode length, and complex rules. Employing reinforcement learning techniques, specifically deep Monte Carlo, and a distributed training framework, we first put forward an AI program named DanZero. Evaluation against baseline AI programs based on heuristic rules highlights the outstanding performance of our bot. Besides, in order to further enhance the AI's capabilities, we apply proximal policy optimization toGuanDanon the basis of Danzero. To address the challenges arising from the huge action space, which will significantly impact the performance of policy-based algorithms, we adopt the pretrained model to compress the action space and integrate action features into the model to bolster its generalization capabilities. Using these techniques, we manage to obtain a newGuanDanAI program DanZero+, which achieves a superior performance compared to DanZero. Youpeng Zhao 0001, Yudong Lu, Jian Zhao 0018, Wengang Zhou 0001, Houqiang Li |
IEEE Trans. Games | 3 |
| 2024 | MCMARL: Parameterizing Value Function via Mixture of Categorical Distributions for Multi-Agent Reinforcement LearningabstractIn cooperative multi-agent tasks, a team of agents jointly interact with an environment by taking actions, receiving a team reward and observing the next state. During the interactions, the uncertainty of environment and reward will inevitably induce stochasticity in the long-term returns and the randomness can be exacerbated with the increasing number of agents. However, such randomness is ignored by most of the existing value-based multi-agent reinforcement learning (MARL) methods, which only model the expectation of Q-value for both individual agents and the team. Compared to using the expectations of the long-term returns, it is preferable to directly model the stochasticity by estimating the returns through distributions. With this motivation, this work proposes a novel value-based MARL framework from a distributional perspective,i.e., parameterizing value function viaMixture ofCategorical distributions for MARL. Specifically, we model both individual Q-values and global Q-value with categorical distribution. To integrate categorical distributions, we define five basic operations on the distribution, which allow the generalization of expected value function factorization methods (e.g., VDN and QMIX) to their MCMARL variants. We further prove that our MCMARL framework satisfiesDistributional-Individual-Global-Max(DIGM) principle with respect to the expectation of distribution, which guarantees the consistency between joint and individual greedy action selections in the global Q-value and individual Q-values. Empirically, we evaluate MCMARL on both a stochastic matrix game and a challenging set of StarCraft II micromanagement tasks, showing the efficacy of our framework. Jian Zhao 0018, Mingyu Yang 0003, Youpeng Zhao 0001, Xunhan Hu, Wengang Zhou 0001, Houqiang Li |
IEEE Trans. Games | 1 |
| 2024 | Full DouZero+: Improving DouDizhu AI by Opponent Modeling, Coach-Guided Training and Bidding LearningabstractWith the development of deep reinforcement learning (DRL), much progress in various perfect and imperfect information games has been achieved. Among these games, DouDizhu, a popular card game in China, poses great challenges because of the imperfect information, large state and action space as well as the cooperation issue. In this paper, we put forward an AI system for this game, which adopts opponent modeling and coach-guided training to help agents make better decisions when playing cards. Besides, we take the bidding phase of DouDizhu into consideration, which is usually ignored by existing works, and train a bidding network using Monte-Carlo simulation. As a result, we achieve a full version of our AI system that is applicable to real-world competitions. We conduct extensive experiments to evaluate the effectiveness of the three techniques adopted in our method and demonstrate the superior performance of our AI over the state-of-the-art DouDizhu AI, i.e., DouZero. We upload our AI systems, one is bidding-free and the other is equipped with a bidding network, to Botzone platform and they both rank the first among over 400 and 250 AI programs on the two corresponding leaderboards, respectively. Our codes are available athttps://github.com/submit-paper/Doudizhu_plus. Youpeng Zhao 0001, Jian Zhao 0018, Xunhan Hu, Wengang Zhou 0001, Houqiang Li |
IEEE Trans. Games | 2 |
| 2024 | Optimizing Camera Motion with MCTS and Target Motion Modeling in Multi-Target Active Object TrackingabstractIn this work, we are dedicated to multi-target active object tracking (AOT), where the goal is to achieve continuous tracking of targets through real-time control of camera. This form of active camera control can be applied to unmanned aerial vehicles (UAV), intelligent robots, and sports events. Our work is conducted in an environment featuring multiple cameras and targets, where our goal is to maximize target coverage. Contrasting with previous research, our work introduces additional degrees of freedom for the cameras, allowing them not only to rotate but also to move along boundary lines. In addition, we model the motion of target to predict the future position of the target in environment. With target’s future position, we use Monte Carlo Tree Search (MCTS) method to find the optimal action of camera. Since the action space is large, we propose to leverage the action selection from multi-agent reinforcement learning (MARL) network to prune the search tree of Monte Carlo Tree Search method, so as to find the optimal action more efficiently. We establish a multi-target 2D environment to simulate several sports games, and experimental results demonstrate that our method can effectively improve the target coverage. The code is available at: http://github.com/HopeChanger/ActiveObjectTracking . Jian Zhao 0018, Mingyu Yang 0003, Wengang Zhou 0001, Houqiang Li |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Implementing First-Person Shooter Game AI in WILD-SCAV with Rule-Enhanced Deep Reinforcement LearningabstractDeep Reinforcement Learning (DRL) algorithms have achieved remarkable performance in various types of games. However, in First-Person Shooter (FPS) games the progress is relatively limited due to the partial observability, sparsity of reward function, complexity in state and action spaces brought by a 3D open environment, and the lack of satisfactory training and testing benchmarks. To this end, Wilderness Scavenger (WILD-SCAV), a novel FPS game environment was developed and used for an online competition held by Inspire-AI and IEEE Conference on Games1, in which we introduced a rule-based agent and wined all three tracks as navigation, supply gathering, and supply battle. In this work, we propose a rule-enhanced deep reinforcement learning algorithm, adopting localization and reconstruction techniques to create 3D occupancy grid maps based on visual observations, and uses them together with heat map and other game variables as the input of a reinforcement learning neural network to conduct basic navigation tasks. Further enhanced and expanded by rules, this navigation agent can handle even more complicated supply gathering and battle tasks. Pre-trained with the previously proposed winning rules in the competition through imitation learning, then trained with distributed PPO algorithm, our agent has better performance over multiple rule-based and RL-based agents, including the agent using winning policy. Zeyu Fang, Jian Zhao 0018, Wengang Zhou 0001, Houqiang Li |
CoG | 2 |
| 2023 | Mastering Curling with RL-revised Decision TreeabstractCurling, also known as "chess on ice", is a popular worldwide sport, which not only tests the physical and mental strength of the participants but also showcases the beauty of movement and stillness and the wisdom of trade-offs. Previously, AI for curling was usually based on decision trees, which required strict artificial prior knowledge and often led to unexpected bugs on extreme occasions. In recent years, however, more and more reinforcement learning algorithms have been proposed in competitive games. Nevertheless, AI derived from RL is very unstable when playing against unseen opponents. In this work, we develop an AI for curling in a novel way, utilizing both decision trees and reinforcement learning. The policy of our AI is defined by a decision tree, and we detect the flaws of it through reinforcement learning. Training a policy model against the decision tree not only helps to mend the flaws of the tree but also provides a way to examine the strength and stability of the tree itself. This approach successfully combines the advantages of RL and decision trees, enhancing the strength and generalization capacity of the policy. Our AI ranked the first among 67 teams in the 2022 RLChina1spring competition. Yuhao Gong, Jian Zhao 0018, Wengang Zhou 0001, Houqiang Li |
CoG | 3 |
| 2023 | DanZero: Mastering GuanDan Game with Reinforcement LearningabstractThe use of artificial intelligence (AI) in card games has been a widely researched topic in the field of AI for an extended period. Recent advancements have led to AI programs exhibiting expert-level gameplay in complex card games such as Mahjong, DouDizhu, and Texas Hold’em. This paper aims to develop an AI program, named DanZero, for GuanDan, an exceptionally complex card game that involves four players competing and cooperating in a long process to upgrade their level quickly. Developing AI for GuanDan is challenging due to its large state and action space, long episode length, and uncertainty in the number of players. To address these challenges, we propose DanZero, the first AI program for GuanDan, that employs reinforcement learning using a distributed framework for training. Our framework consists of two processes: the Actor Process and the Learner Process. In the Actor Process, we design state features and generate samples through agents’ self-play. In the Learner Process, we update the model using the Deep Monte-Carlo Method. We trained DanZero for 30 days, utilizing 160 CPUs and 1 GPU to develop the program successfully. We compared DanZero’s performance with eight baseline AI programs based on heuristic rules, and our results indicate DanZero’s exceptional performance. We further tested DanZero with human players and demonstrated its ability to perform at a human level. The code for DanZero can be found in the supplementary material. Yudong Lu, Jian Zhao 0018, Youpeng Zhao 0001, Wengang Zhou 0001, Houqiang Li |
CoG | 2 |
| 2023 | Neural Dependencies Emerging from Learning Massive CategoriesabstractThis work presents two astonishing findings on neural networks learned for large-scale image classification. 1) Given a well-trained model, the logits predicted for some category can be directly obtained by linearly combining the predictions of a few other categories, which we call neural dependency. 2) Neural dependencies exist not only within a single model, but even between two independently learned models, regardless of their architectures. Towards a theoretical analysis of such phenomena, we demonstrate that identifying neural dependencies is equivalent to solving the Covariance Lasso (CovLasso) regression problem proposed in this paper. Through investigating the properties of the problem solution, we confirm that neural dependency is guaranteed by a redundant logit covariance matrix, which condition is easily met given massive categories, and that neural dependency is highly sparse, implying that one category correlates to only a few others. We further empirically show the potential of neural dependencies in understanding internal data correlations, generalizing models to unseen categories, and improving model robustness with a dependency-derived regularizer. Code to reproduce the results in this paper is available at https://github.com/RuiLiFengiNeural-Dependencies. Ruili Feng, Kecheng Zheng, Kai Zhu 0004, Yujun Shen, Jian Zhao 0018, Deli Zhao, Jingren Zhou 0001, Michael I. Jordan, Zhengjun Zha |
CVPR | 5 |
| 2023 | Q-SAT: Value Factorization with Self-Attention for Deep Multi-Agent Reinforcement LearningabstractIn many real-world tasks, a team of agents learn to cooperate with each other under the setting of partial observability and communication constraints, where value factorization has been demonstrated as an effective solution. In a multi-agent system, it's important to capture the inter-connection between agents and push agents to consider more of the relevant teammates. Motivated by the success of self-attention in natural language processing and computer vision, we propose a novel value factorization mechanism, called Q-function Self ATtention (Q-SAT). It models the pairwise action-value functions and connection coefficient between agent pairs explicitly, and pays more attention to the interrelated agents when making decisions. Satisfying the IGM principle, Q-SAT introduces the self-attention into value factorization network. This attention mechanism enables more effective and efficient learning in complex multi-agent environments. Q-SAT can be viewed as a basic building block and is ready to be applied to existing value factorization methods. The experimental results show that Q-SAT captures the connection relationship between agents and significantly improves the learning performance on the challenging StarCraft II micromanagement task and Google Research Football task. Xunhan Hu, Jian Zhao 0018, Youpeng Zhao 0001, Wengang Zhou 0001, Houqiang Li |
IJCNN | 2 |
| 2023 | DIFFER: Decomposing Individual Reward for Fair Experience Replay in Multi-Agent Reinforcement LearningabstractCooperative multi-agent reinforcement learning (MARL) is a challenging task, as agents must learn complex and diverse individual strategies from a shared team reward. However, existing methods struggle to distinguish and exploit important individual experiences, as they lack an effective way to decompose the team reward into individual rewards. To address this challenge, we propose DIFFER, a powerful theoretical framework for decomposing individual rewards to enable fair experience replay in MARL.
By enforcing the invariance of network gradients, we establish a partial differential equation whose solution yields the underlying individual reward function. The individual TD-error can then be computed from the solved closed-form individual rewards, indicating the importance of each piece of experience in the learning task and guiding the training process. Our method elegantly achieves an equivalence to the original learning framework when individual experiences are homogeneous, while also adapting to achieve more muscular efficiency and fairness when diversity is observed.
Our extensive experiments on popular benchmarks validate the effectiveness of our theory and method, demonstrating significant improvements in learning efficiency and fairness.
Code is available in supplement material. Xunhan Hu, Jian Zhao 0018, Wengang Zhou 0001, Ruili Feng, Houqiang Li |
NeurIPS | 2 |
| 2023 | Improving Deep Reinforcement Learning With Mirror LossabstractRecent years have witnessed the great breakthrough of deep reinforcement learning (DRL) in various artificial intelligence applications, but the training process needs a very large amount of samples and huge computational costs. To alleviate the low sample efficiency issue, one feasible solution is to improve the state representation learning. We uncover that the agents, trained by the original DRL algorithm, face severe performance degradation in mirrored game environments. As mirror symmetry is an important property of the environment, the poor performance in the mirrored situation indicates that the agents are not fully aware of the essence of the environment. In order to handle this problem and make use of the property to attain better state representation, we propose a mirror loss, which serves as an auxiliary module to bring mirror symmetry representation to the DRL agent. It is model-agnostic and prompts the DRL agent to make logically consistent mirrored actions in the mirrored environment. We conduct experiments on OpenAI Gym Atari environments and a more complex reinforcement learning task, Mahjong AI, and the results demonstrate the efficiency and versatility of our method. Jian Zhao 0018, Weide Shu, Youpeng Zhao 0001, Wengang Zhou 0001, Houqiang Li |
IEEE Trans. Games | 1 |
| 2022 | Mastering the Game of 3v3 Snakes with Rule-Enhanced Multi-Agent Reinforcement LearningabstractAs a popular game around the world, Snakes has multiple modes with different settings. In this work, we are dedicated to the 3v3 Snakes, which is characterized by a complex mixture of competition and cooperation. To address this mode of Snakes, most existing AI agents adopt rule based methods, which achieve limited performance due to human’s oversight of some special circumstances. Inspired by the superiority of multi-agent reinforcement learning (MARL), we propose a rule-enhanced multi-agent reinforcement learning algorithm and build a 3v3 Snakes AI. Specifically, we introduce the territory matrix which is commonly utilized in rule based methods to the state features and mask the illegal actions through designed rules. The relationships of individual-team and friends-foes are also merged into reward design. Trained with Distributed PPO and self-play on a single GeForce RTX 2080 GPU for twenty-four hours, our AI achieves state-of-the-art performance and beats human players. On JIDI platform, our agent outperforms the other 132 participating agents and ranks the first for more than 20 consecutive days. Jitao Wang, Dongyun Xue, Jian Zhao 0018, Wengang Zhou 0001, Houqiang Li |
CoG | 3 |
| 2022 | DouZero+: Improving DouDizhu AI by Opponent Modeling and Coach-guided LearningabstractRecent years have witnessed the great breakthrough of deep reinforcement learning (DRL) in various perfect and imperfect information games. Among these games, DouDizhu, a popular card game in China, is very challenging due to the imperfect information, large state and action space as well as elements of collaboration. Recently, a DouDizhu AI system called DouZero has been proposed. Trained using traditional Monte Carlo method with deep neural networks and self-play procedure without the abstraction of human prior knowledge, DouZero has achieved the best performance among all the existing DouDizhu AI programs. In this work, we propose to enhance DouZero by introducing opponent modeling into DouZero. Besides, we propose a novel coach network to further boost the performance of DouZero and accelerate its training process. With the integration of the above two techniques into DouZero, our DouDizhu AI system achieves better performance and ranks top in the Botzone leaderboard among more than 400 AI agents, including DouZero. Youpeng Zhao 0001, Jian Zhao 0018, Xunhan Hu, Wengang Zhou 0001, Houqiang Li |
CoG | 2 |
| 2022 | LDSA: Learning Dynamic Subtask Assignment in Cooperative Multi-Agent Reinforcement LearningabstractCooperative multi-agent reinforcement learning (MARL) has made prominent progress in recent years. For training efficiency and scalability, most of the MARL algorithms make all agents share the same policy or value network. However, in many complex multi-agent tasks, different agents are expected to possess specific abilities to handle different subtasks. In those scenarios, sharing parameters indiscriminately may lead to similar behavior across all agents, which will limit the exploration efficiency and degrade the final performance. To balance the training complexity and the diversity of agent behavior, we propose a novel framework to learn dynamic subtask assignment (LDSA) in cooperative MARL. Specifically, we first introduce a subtask encoder to construct a vector representation for each subtask according to its identity. To reasonably assign agents to different subtasks, we propose an ability-based subtask selection strategy, which can dynamically group agents with similar abilities into the same subtask. In this way, agents dealing with the same subtask share their learning of specific abilities and different subtasks correspond to different specific abilities. We further introduce two regularizers to increase the representation difference between subtasks and stabilize the training by discouraging agents from frequently changing subtasks, respectively. Empirical results show that LDSA learns reasonable and effective subtask assignment for better collaboration and significantly improves the learning performance on the challenging StarCraft II micromanagement benchmark and Google Research Football. Mingyu Yang 0003, Jian Zhao 0018, Xunhan Hu, Wengang Zhou 0001, Jiangcheng Zhu, Houqiang Li |
NeurIPS | 2 |
| 2022 | Coach-assisted multi-agent reinforcement learning framework for unexpected crashed agentsabstractMulti-agent reinforcement learning is difficult to apply in practice, partially because of the gap between simulated and real-world scenarios. One reason for the gap is that simulated systems always assume that agents can work normally all the time, while in practice, one or more agents may unexpectedly “crash” during the coordination process due to inevitable hardware or software failures. Such crashes destroy the cooperation among agents and lead to performance degradation. In this work, we present a formal conceptualization of a cooperative multi-agent reinforcement learning system with unexpected crashes. To enhance the robustness of the system to crashes, we propose a coach-assisted multi-agent reinforcement learning framework that introduces a virtual coach agent to adjust the crash rate during training. We have designed three coaching strategies (fixed crash rate, curriculum learning, and adaptive crash rate) and a re-sampling strategy for our coach agent. To our knowledge, this work is the first to study unexpected crashes in a multi-agent system. Extensive experiments on grid-world and StarCraft II micromanagement tasks demonstrate the efficacy of the adaptive strategy compared with the fixed crash rate strategy and curriculum learning strategy. The ablation study further illustrates the effectiveness of our re-sampling strategy. Jian Zhao 0018, Youpeng Zhao 0001, Weixun Wang, Mingyu Yang 0003, Xunhan Hu, Wengang Zhou 0001, Jianye Hao, Houqiang Li |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2022 | Conditional Sentence Generation and Cross-Modal Reranking for Sign Language TranslationabstractSign Language Translation (SLT) aims to generate spoken language translations from sign language videos. Currently, the available sign language datasets are relatively too small to learn the linguistic properties of spoken language. In this paper, towards effective SLT, we propose a novel framework which takes the advantage of the spoken language grammar learnt from a large corpus of text sentences. Our framework consists of three key modules: word existence verification, conditional sentence generation and cross-modal re-ranking. We first check the existence of words in the vocabulary by a series of binary classification in parallel. After that, the appearing words are assembled and guided by a pretrained spoken language generator to produce multiple candidate sentences in spoken language manner. Last but not least, we select the sentence most semantically similar to the input sign video as the translation result with a crossmodal re-ranking model. We evaluate our framework on two large scale continuous SLT benchmarks,i.e., CSL and RWTHPHOENIX-Weather 2014 T. Experimental results demonstrate that the proposed framework achieves promising performance on both datasets. Jian Zhao 0018, Weizhen Qi, Wengang Zhou 0001, Nan Duan 0001, Ming Zhou 0001, Houqiang Li |
IEEE Trans. Multim. | 1 |
| 2021 | Interest-aware Item Combination PredictionabstractOne of the new upcoming recommendation scenarios is combinatorial item recommendation, where the recommender system aims to recommend a bundle of items for each user request to maximize the aggregated rewards. In this work, we focus on the item combination prediction task with constraints, i.e., predicting the user’s feedback to the nine exposed items given that if the user wants to buy any item in the next list, she/he must purchase all the three items in the current list. Therefore, the users’ response depends on not only the current item but also the items of the next list. Considering the constraint on the users’ feedbacks, we transform the binary combinatorial prediction into the multi-class prediction. To extract users’ preferences from their historical behaviors, we design the attentive interest network and hierarchical GRU network. The final result of the competition demonstrates the effectiveness of our strategy. Shiwen Wu, Jian Zhao 0018 |
IEEE BigData | 2 |
| 2021 | Semantic Boundary Detection With Reinforcement Learning for Continuous Sign Language RecognitionabstractSign language recognition (SLR) is a significant and promising technique to facilitate the communication for the hearing-impaired people. In this paper, we are dedicated to weakly supervised continuous SLR, where for each sign video, there are only ordered gloss labels without temporal boundary along frames. To explicitly align video frames to the sign words in a sign video, we propose a novel semantic boundary detection method based on reinforcement learning for accurate continuous SLR. In our approach, we first propose a multi-scale perception scheme to learn discriminative representation for video clips. Then, we formulate the semantic boundary detection as a reinforcement learning problem. We define the state as the feature representation of a video segment, and the action as the determination of the semantic boundary's location. The reward is computed by the quantitative performance metric between the prediction sentence and the ground truth sentence. The policy network is trained with a policy gradient algorithm. Extensive experiments are conducted on CSL Split II and RWTH-PHOENIX-Weather 2014 datasets, and the results demonstrate the effectiveness and superiority of our method. Chengcheng Wei, Jian Zhao 0018, Wengang Zhou 0001, Houqiang Li |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | State Representation Learning For Effective Deep Reinforcement LearningabstractRecent years have witnessed the great success of deep reinforcement learning (DRL) on a variety of vision games. Although DNN has demonstrated strong power in representation learning, such capacity is under-explored in most DRL works whose focus is usually on optimization solvers. In fact, we discover that the state feature learning is the main obstacle for further improvement of DRL algorithms. To address this issue, we propose a new state representation learning scheme with our Adjacent State Consistency Loss (ASC Loss). The loss is defined based on the hypothesis that there are fewer changes between adjacent states than that of far apart ones, since scenes in videos generally evolve smoothly. In this paper, we exploit ASC loss as an assistant of RL loss in the training phase to boost the state feature learning. We conduct evaluation on Atari games and MuJoCo continuous control tasks, which demonstrates that our method is superior to OpenAI baselines. Jian Zhao 0018, Wengang Zhou 0001, Houqiang Li |
ICME | 1 |