EDBT 2026 Demo / reviewers in the wild / expert
Chao Yu 0004
dblp:36/6789-4
· DBLP profile ↗
47ranked-venue papers
11as first author
27since 2021 · last 2026
0000-0002-4371-3663ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 9 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 8 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CATAL: Causally Disentangled Task Representation Learning for Offline Meta-Reinforcement LearningabstractContext-based Offline Meta Reinforcement Learning (COMRL) has shown promising results in improving the cross-task generalization ability of meta-policies. However, current methods often lead to entangled task representations, in which each latent dimension is influenced by multiple causal factors that govern variations in environment dynamics and reward mechanisms. This entanglement can degrade generalization performance, particularly when multiple causal factors vary simultaneously across tasks. To address this limitation, we propose CAusally disentangled TAsk representation Learning (CATAL) method for COMRL that aims to improve the generalization ability of the meta-policy, where each latent dimension in the task representations aligns to a single causal factor.Theoretically, we show that under mild conditions, the task representations learned by CATAL are causally disentangled. Empirically, extensive results on multi-task MuJoCo benchmarks show that CATAL consistently outperforms existing COMRL baselines in both in-distribution and out-of-distribution generalization. Shan Cong, Chao Yu 0004, Xiangyuan Lan |
AAAI | 2 |
| 2026 | Reliability-Guaranteed and Reward-Seeking Sequence Modeling for Model-Based Offline Reinforcement LearningabstractAs a data-driven learning approach, model-based offline reinforcement learning (MORL) aims to learn a policy by exploiting a dynamics model derived from an existing dataset. Applying conservative quantification to the dynamics model, most existing works on MORL generate trajectories that approximate the real data distribution to facilitate policy learning. However, these methods typically overlook the influence of historical information on environmental dynamics, thus generating unreliable trajectories that fail to align with the true data distribution. In this paper, we propose a new MORL algorithm called Reliability-guaranteed and Reward-seeking Transformer (RT). RT can avoid generating unreliable trajectories through the calculation of cumulative reliability of the trajectories, which is a weighted variational distance between the generated trajectory distribution and the true data distribution. Moreover, by sampling candidate actions with high rewards, RT can efficiently generate high-reward trajectories from the existing offline data, thereby further facilitating policy learning. We theoretically prove the performance guarantees of RT in policy learning, and empirically demonstrate its effectiveness against state-of-the-art model-based methods on several offline benchmark tasks and a large-scale industrial dataset from an on-demand food delivery platform. Shenghong He, Chao Yu 0004, Yile Liang, Xuetao Ding |
AAAI | 2 |
| 2025 | Rapid Learning in Constrained Minimax Games with Negative MomentumabstractIn this paper, we delve into the utilization of the negative momentum technique in constrained minimax games. From an intuitive mechanical standpoint, we introduce a novel framework for momentum buffer updating, which extends the findings of negative momentum from the unconstrained setting to the constrained setting and provides a universal enhancement to the classic game-solver algorithms. Additionally, we provide theoretical guarantees of convergence for our momentum-augmented learning algorithms. We then extend these algorithms to their extensive-form counterparts. Experimental results on both Normal Form Games (NFGs) and Extensive Form Games (EFGs) demonstrate that our momentum techniques can significantly improve algorithm performance, surpassing both their original versions and the SOTA baselines by a large margin. Zijian Fang, Zongkai Liu, Chao Yu 0004, Chaohao Hu |
AAAI | 3 |
| 2025 | Offline Multi-Agent Reinforcement Learning via In-Sample Sequential Policy OptimizationabstractOffline Multi-Agent Reinforcement Learning (MARL) is an emerging field that aims to learn optimal multi-agent policies from pre-collected datasets. Compared to single-agent case, multi-agent setting involves a large joint state-action space and coupled behaviors of multiple agents, which bring extra complexity to offline policy optimization. In this work, we revisit the existing offline MARL methods and show that in certain scenarios they can be problematic, leading to uncoordinated behaviors and out-of-distribution (OOD) joint actions. To address these issues, we propose a new offline MARL algorithm, named In-Sample Sequential Policy Optimization (InSPO). InSPO sequentially updates each agent's policy in an in-sample manner, which not only avoids selecting OOD joint actions but also carefully considers teammates' updated policies to enhance coordination. Additionally, by thoroughly exploring low-probability actions in the behavior policy, InSPO can well address the issue of premature convergence to sub-optimal solutions. Theoretically, we prove InSPO guarantees monotonic policy improvement and converges to quantal response equilibrium (QRE). Experimental results demonstrate the effectiveness of our method compared to current state-of-the-art offline MARL methods. Zongkai Liu, Chao Yu 0004, Xiawei Wu, Yile Liang, Xuetao Ding |
AAAI | 3 |
| 2025 | Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation LearningabstractRecovering a spectrum of diverse policies from a set of expert trajectories is an important research topic in imitation learning. After determining a latent style for a trajectory, previous diverse polices recovering methods usually employ a vanilla behavioral cloning learning objective conditioned on the latent style, treating each state-action pair in the trajectory with equal importance. Based on an observation that in many scenarios, behavioral styles are often highly relevant with only a subset of state-action pairs, this paper presents a new principled method in diverse polices recovering. In particular, after inferring or assigning a latent style for a trajectory, we enhance the vanilla behavioral cloning by incorporating a weighting mechanism based on pointwise mutual information.
This additional weighting reflects the significance of each state-action pair's contribution to learning the style, thus allowing our method to focus on state-action pairs most representative of that style.
We provide theoretical justifications for our new objective, and extensive empirical evaluations confirm the effectiveness of our method in recovering diverse polices from expert data. Jian Yao 0008, Weiming Liu 0004, Hanmin Qin, Hansheng Kong, Kirk Tang, Jiechao Xiong, Chao Yu 0004, Kai Li 0022, Junliang Xing, Hongwu Chen, Juchao Zhuo, Qiang Fu 0016, Haobo Fu |
ICLR | 9 |
| 2025 | Conservative Offline Goal-Conditioned Implicit V-LearningabstractOffline goal-conditioned reinforcement learning (GCRL) learns a goal-conditioned value function to train policies for diverse goals with pre-collected datasets. Hindsight experience replay addresses the issue of sparse rewards by treating intermediate states as goals but fails to complete goal-stitching tasks where achieving goals requires stitching different trajectories. While cross-trajectory sampling is a potential solution that associates states and goals belonging to different trajectories, we demonstrate that this direct method degrades performance in goal-conditioned tasks due to the overestimation of values on unconnected pairs. To this end, we propose Conservative Goal-Conditioned Implicit Value Learning (CGCIVL), a novel algorithm that introduces a penalty term to penalize value estimation for unconnected state-goal pairs and leverages the quasimetric framework to accurately estimate values for connected pairs. Evaluations on OGBench, a benchmark for offline GCRL, demonstrate that CGCIVL consistently surpasses state-of-the-art methods across diverse tasks. Kaiqiang Ke, Zongkai Liu, Shenghong He, Chao Yu 0004 |
ICML | 5 |
| 2025 | Hierarchical task network-enhanced multi-agent reinforcement learning: Toward efficient cooperative strategies
Xuechen Mu, Hankui Zhuo, Chen Chen 0077, Kai Zhang 0012, Chao Yu 0004, Jianye Hao |
Neural Networks | 5 |
| 2025 | Hierarchical Multi-Agent Meta-Reinforcement Learning for Cross-Channel BiddingabstractReal-time bidding (RTB) plays a pivotal role in online advertising ecosystems. Advertisers employ strategic bidding to optimize their advertising impact while adhering to various financial constraints, such as the return-on-investment (ROI) and cost-per-click (CPC). Primarily focusing on bidding with fixed budget constraints, traditional approaches cannot effectively manage the dynamic budget allocation problem where the goal is to achieve global optimization of bidding performance across multiple channels with a shared budget. In this paper, we propose a hierarchical multi-agent reinforcement learning framework for multi-channel bidding optimization. In this framework, the top-level strategy applies a CPC constrained diffusion model to dynamically allocate budgets among the channels according to their distinct features and complex interdependencies, while the bottom-level strategy adopts a state-action decoupled actor-critic method to address the problem of extrapolation errors in offline learning caused by out-of-distribution actions and a context-based meta-channel knowledge learning method to improve the state representation capability of the policy based on the shared knowledge among different channels. Comprehensive experiments conducted on a large scale real-world industrial dataset from the Meituan ad bidding platform demonstrate that our method achieves a state-of-the-art performance. Shenghong He, Chao Yu 0004, Shangqin Mao, Bo Tang 0018, Qianlong Xie |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | RegFTRL: Efficient Equilibrium Learning in Two-Player Zero-Sum GamesabstractRecent literature has witnessed a rising interest in learning Nash equilibrium with a guarantee of last-iterate convergence.In this paper, we introduce a novel approach called Regularized Followthe-Regularized-Leader (RegFTRL), which is an efficient variant of FTRL enriched with an adaptive regularization, for the purpose of learning equilibria in two-player zero-sum games.In the context of normal-form games (NFGs), our proposed RegFTRL algorithm exhibits desirable property of last-iterate linear convergence towards an approximated equilibrium, and converges to an exact Nash equilibrium through adaptive adjustments of the regularization.Moreover, we extend our method to extensive-form games (EFGs) and propose FollowMu, a practical implementation of RegFTRL with a neural network as the function approximator, for model-free learning in sequential non-stationary environments.Finally, empirical results substantiate the theoretical properties of RegFTRL, and demonstrate that FollowMu can achieve favorable performance in EFGs. Zijian Fang, Zongkai Liu, Chao Yu 0004 |
DAI | 3 |
| 2024 | Off-Policy Primal-Dual Safe Reinforcement LearningabstractPrimal-dual safe RL methods commonly perform iterations between the primal update of the policy and the dual update of the Lagrange Multiplier. Such a training paradigm is highly susceptible to the error in cumulative cost estimation since this estimation serves as the key bond connecting the primal and dual update processes. We show that this problem causes significant underestimation of cost when using off-policy methods, leading to the failure to satisfy the safety constraint. To address this issue, we propose conservative policy optimization, which learns a policy in a constraint-satisfying area by considering the uncertainty in cost estimation. This improves constraint satisfaction but also potentially hinders reward maximization. We then introduce local policy convexification to help eliminate such suboptimality by gradually reducing the estimation uncertainty. We provide theoretical interpretations of the joint coupling effect of these two ingredients and further verify them by extensive experiments. Results on benchmark tasks show that our method not only achieves an asymptotic performance comparable to state-of-the-art on-policy methods while using much fewer samples, but also significantly reduces constraint violation during training. Our code is available at https://github.com/ZifanWu/CAL. Zifan Wu, Bo Tang 0018, Chao Yu 0004, Shangqin Mao, Qianlong Xie, Dong Wang 0022 |
ICLR | 4 |
| 2024 | An Offline Adaptation Framework for Constrained Multi-Objective Reinforcement LearningabstractIn recent years, significant progress has been made in multi-objective reinforcement learning (RL) research, which aims to balance multiple objectives by incorporating preferences for each objective. In most existing studies, specific preferences must be provided during deployment to indicate the desired policies explicitly. However, designing these preferences depends heavily on human prior knowledge, which is typically obtained through extensive observation of high-performing demonstrations with expected behaviors. In this work, we propose a simple yet effective offline adaptation framework for multi-objective RL problems without assuming handcrafted target preferences, but only given several demonstrations to implicitly indicate the preferences of expected policies. Additionally, we demonstrate that our framework can naturally be extended to meet constraints on safety-critical objectives by utilizing safe demonstrations, even when the safety thresholds are unknown. Empirical results on offline multi-objective and safe tasks demonstrate the capability of our framework to infer policies that align with real preferences while meeting the constraints implied by the provided demonstrations. Zongkai Liu, Danying Mo, Chao Yu 0004 |
NeurIPS | 4 |
| 2024 | Leveraging Joint-Action Embedding in Multiagent Reinforcement Learning for Cooperative GamesabstractState-of-the-art multi-agent policy gradient (MAPG) methods have demonstrated convincing capability in many cooperative games. However, the exponentially growing joint-action space severely challenges the critic's value evaluation and hinders performance of MAPG methods. To address this issue, we augment Central-Q policy gradient with a joint-action embedding function and propose Mutual-information Maximization MAPG (M3APG). The joint-action embedding function makes joint-actions contain information of state transitions, which will improve the critic's generalization over the joint-action space by allowing it to infer joint-actions' outcomes. We theoretically prove that with a fixed joint-action embedding function, the convergence of M3APG is guaranteed. Experiment results on the StarCraft Multi-Agent Challenge (SMAC) demonstrate that M3APG gives evaluation results with better accuracy and outperform other MAPG basic models across various maps of multiple difficulty levels. We empirically show that our joint-action embedding model can be extended to value-based multi-agent reinforcement learning methods and state-of-the-art MAPG methods. Finally, we run ablation study to show that the usage of mutual information in our method is necessary and effective. Xingzhou Lou, Junge Zhang, Yali Du 0001, Chao Yu 0004, Zhaofeng He 0001, Kaiqi Huang |
IEEE Trans. Games | 4 |
| 2024 | An Offline-Transfer-Online Framework for Cloud-Edge Collaborative Distributed Reinforcement LearningabstractRecent advances in deep reinforcement learning (DRL) have made it possible to train various powerful agents to perform complex tasks in real-time environments. With the next-generation communication technologies, making cloud-edge collaborative artificial intelligence service with evolved DRL agents can be a significant scenario. However, agents with different algorithms and architectures in the same DRL scenario may not be compatible, and training them is either time-consuming or resource-demanding. In this paper, we design a novel cloud-edge collaborative DRL training framework, named Offline-Transfer-Online, which is a new approach that can speed up the convergence of online DRL agents at the edge by interacting with offline agents in the cloud, with the minimum data interchanged and without relying on high-quality offline datasets. Therein, we propose a novel algorithm-independent knowledge distillation algorithm for online RL agents, by leveraging pre-trained models and the interface between agents and the environment to transfer distilled knowledge among multiple heterogeneous agents efficiently. Extensive experiments show that our algorithm can accelerate the convergence of various online agents in a double to decuple speed, with comparable reward achieved in different environments. Tianyu Zeng, Xiaoxi Zhang 0001, Jingpu Duan, Chao Yu 0004, Chuan Wu 0001, Xu Chen 0004 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2023 | Models as Agents: Optimizing Multi-Step Predictions of Interactive Local Models in Model-Based Multi-Agent Reinforcement LearningabstractResearch in model-based reinforcement learning has made significant progress in recent years. Compared to single-agent settings, the exponential dimension growth of the joint state-action space in multi-agent systems dramatically increases the complexity of the environment dynamics, which makes it infeasible to learn an accurate global model and thus necessitates the use of agent-wise local models. However, during multi-step model rollouts, the prediction of one local model can affect the predictions of other local models in the next step. As a result, local prediction errors can be propagated to other localities and eventually give rise to considerably large global errors. Furthermore, since the models are generally used to predict for multiple steps, simply minimizing one-step prediction errors regardless of their long-term effect on other models may further aggravate the propagation of local errors. To this end, we propose Models as AGents (MAG), a multi-agent model optimization framework that reversely treats the local models as multi-step decision making agents and the current policies as the dynamics during the model rollout process. In this way, the local models are able to consider the multi-step mutual affect between each other before making predictions. Theoretically, we show that the objective of MAG is approximately equivalent to maximizing a lower bound of the true environment return. Experiments on the challenging StarCraft II benchmark demonstrate the effectiveness of MAG. Zifan Wu, Chao Yu 0004, Chen Chen 0077, Jianye Hao, Hankui Zhuo |
AAAI | 2 |
| 2023 | Subspace-Aware Exploration for Sparse-Reward Multi-Agent TasksabstractExploration under sparse rewards is a key challenge for multi-agent reinforcement learning problems. One possible solution to this issue is to exploit inherent task structures for an acceleration of exploration. In this paper, we present a novel exploration approach, which encodes a special structural prior on the reward function into exploration, for sparse-reward multi-agent tasks. Specifically, a novel entropic exploration objective which encodes the structural prior is proposed to accelerate the discovery of rewards. By maximizing the lower bound of this objective, we then propose an algorithm with moderate computational cost, which can be applied to practical tasks. Under the sparse-reward setting, we show that the proposed algorithm significantly outperforms the state-of-the-art algorithms in the multiple-particle environment, the Google Research Football and StarCraft II micromanagement tasks. To the best of our knowledge, on some hard tasks (such as 27m_vs_30m}) which have relatively larger number of agents and need non-trivial strategies to defeat enemies, our method is the first to learn winning strategies under the sparse-reward setting. Pei Xu 0003, Junge Zhang, Qiyue Yin, Chao Yu 0004, Yaodong Yang 0001, Kaiqi Huang |
AAAI | 4 |
| 2023 | Safe Offline Reinforcement Learning with Real-Time Budget ConstraintsabstractAiming at promoting the safe real-world deployment of Reinforcement Learning (RL), research on safe RL has made significant progress in recent years. However, most existing works in the literature still focus on the online setting where risky violations of the safety budget are likely to be incurred during training. Besides, in many realworld applications, the learned policy is required to respond to dynamically determined safety budgets (i.e., constraint threshold) in real time. In this paper, we target at the above real-time budget constraint problem under the offline setting, and propose Trajectory-based REal-time Budget Inference (TREBI) as a novel solution that approaches this problem from the perspective of trajectory distribution. Theoretically, we prove an error bound of the estimation on the episodic reward and cost under the offline setting and thus provide a performance guarantee for TREBI. Empirical results on a wide range of simulation tasks and a real-world large-scale advertising application demonstrate the capability of TREBI in solving real-time budget constraint problems under offline settings. Bo Tang 0018, Zifan Wu, Chao Yu 0004, Shangqin Mao, Qianlong Xie, Dong Wang 0022 |
ICML | 4 |
| 2023 | Causal Deep Reinforcement Learning Using Observational DataabstractDeep reinforcement learning (DRL) requires the collection of interventional data, which is sometimes expensive and even unethical in the real world, such as in the autonomous driving and the medical field. Offline reinforcement learning promises to alleviate this issue by exploiting the vast amount of observational data available in the real world. However, observational data may mislead the learning agent to undesirable outcomes if the behavior policy that generates the data depends on unobserved random variables (i.e., confounders). In this paper, we propose two deconfounding methods in DRL to address this problem. The methods first calculate the importance degree of different samples based on the causal inference technique, and then adjust the impact of different samples on the loss function by reweighting or resampling the offline dataset to ensure its unbiasedness. These deconfounding methods can be flexibly combined with existing model-free DRL algorithms such as soft actor-critic and deep Q-learning, provided that a weak condition can be satisfied by the loss functions of these algorithms. We prove the effectiveness of our deconfounding methods and validate them experimentally. Wenxuan Zhu, Chao Yu 0004, Qiang Zhang 0008 |
IJCAI | 2 |
| 2022 | Creativity of AI: Automatic Symbolic Option Discovery for Facilitating Deep Reinforcement LearningabstractDespite of achieving great success in real life, Deep Reinforcement Learning (DRL) is still suffering from three critical issues, which are data efficiency, lack of the interpretability and transferability. Recent research shows that embedding symbolic knowledge into DRL is promising in addressing those challenges. Inspired by this, we introduce a novel deep reinforcement learning framework with symbolic options. This framework features a loop training procedure, which enables guiding the improvement of policy by planning with action models and symbolic options learned from interactive trajectories automatically. The learned symbolic options help doing the dense requirement of expert domain knowledge and provide inherent interpretabiliy of policies. Moreover, the transferability and data efficiency can be further improved by planning with the action models. To validate the effectiveness of this framework, we conduct experiments on two domains, Montezuma's Revenge and Office World respectively, and the results demonstrate the comparable performance, improved data efficiency, interpretability and transferability. Mu Jin, Kebing Jin, Hankui Zhuo, Chen Chen 0077, Chao Yu 0004 |
AAAI | 6 |
| 2022 | Deep Reinforcement Learning for Bandit Arm LocalizationabstractIn the multi-armed bandit (MAB) framework, we investigate the problem of learning the means of distributions that are associated with a finite n umber o f a rms under a monotonic constraint. Different from the traditional MAB, our problem involves a parameter constraint and a limited trial budget (i.e., the number of arm pulls is small). However, the number of training samples can be as large as possible through (infinite) simulations, while each training sample is of limited size. This situation arises when some additional information is provided before the trial starts and each arm pull (or testing) could be of extraordinary cost. For example, in cancer dose-finding clinical trials, higher toxicity probabilities are typically associated with higher dose levels (i.e., the monotonic dose–toxicity constraint), and the loss due to the drug’s toxicity, side-effects or death of patients can be enormous. We formulate this problem in the reinforcement learning (RL) paradigm, which is referred to as a bandit arm localization problem. We propose a novel approach in a double deep Q-learning framework, which is integrated with a state-of-the-art statistical model to preserve the parameter constraint and develop a more effective learning strategy. The double deep Q-learning model can be trained with a large (can be as large as infinite) number of simulated trials, which is the first time to cast dose finding in the RL framework. We evaluate the performance of our approach through extensive simulation studies in realistic settings of phase I clinical trials. The proposed double deep Q-learning is shown to outperform the baseline methods in cancer dose-finding trials. Wenbin Du, Huaqing Jin, Chao Yu 0004, Guosheng Yin |
IEEE Big Data | 3 |
| 2022 | A Unified Diversity Measure for Multiagent Reinforcement LearningabstractPromoting behavioural diversity is of critical importance in multi-agent reinforcement learning, since it helps the agent population maintain robust performance when encountering unfamiliar opponents at test time, or, when the game is highly non-transitive in the strategy space (e.g., Rock-Paper-Scissor). While a myriad of diversity metrics have been proposed, there are no widely accepted or unified definitions in the literature, making the consequent diversity-aware learning algorithms difficult to evaluate and the insights elusive. In this work, we propose a novel metric called the Unified Diversity Measure (UDM) that offers a unified view for existing diversity metrics. Based on UDM, we design the UDM-Fictitious Play (UDM-FP) and UDM-Policy Space Response Oracle (UDM-PSRO) algorithms as efficient solvers for normal-form games and open-ended games. In theory, we prove that UDM-based methods can enlarge the gamescape by increasing the response capacity of the strategy pool, and have convergence guarantee to two-player Nash equilibrium. We validate our algorithms on games that show strong non-transitivity, and empirical results show that our algorithms achieve better performances than strong PSRO baselines in terms of the exploitability and population effectivity. Zongkai Liu, Chao Yu 0004, Yaodong Yang 0002, Zifan Wu |
NeurIPS | 2 |
| 2022 | Plan To Predict: Learning an Uncertainty-Foreseeing Model For Model-Based Reinforcement LearningabstractIn Model-based Reinforcement Learning (MBRL), model learning is critical since an inaccurate model can bias policy learning via generating misleading samples. However, learning an accurate model can be difficult since the policy is continually updated and the induced distribution over visited states used for model learning shifts accordingly. Prior methods alleviate this issue by quantifying the uncertainty of model-generated samples. However, these methods only quantify the uncertainty passively after the samples were generated, rather than foreseeing the uncertainty before model trajectories fall into those highly uncertain regions. The resulting low-quality samples can induce unstable learning targets and hinder the optimization of the policy. Moreover, while being learned to minimize one-step prediction errors, the model is generally used to predict for multiple steps, leading to a mismatch between the objectives of model learning and model usage. To this end, we propose Plan To Predict (P2P), an MBRL framework that treats the model rollout process as a sequential decision making problem by reversely considering the model as a decision maker and the current policy as the dynamics. In this way, the model can quickly adapt to the current policy and foresee the multi-step future uncertainty when generating trajectories. Theoretically, we show that the performance of P2P can be guaranteed by approximately optimizing a lower bound of the true environment return. Empirical results demonstrate that P2P achieves state-of-the-art performance on several challenging benchmark tasks. Zifan Wu, Chao Yu 0004, Chen Chen 0077, Jianye Hao, Hankui Zhuo |
NeurIPS | 2 |
| 2022 | Offline reinforcement learning with representations for actions
Xingzhou Lou, Qiyue Yin, Junge Zhang, Chao Yu 0004, Zhaofeng He 0001, Nengjie Cheng, Kaiqi Huang |
Inf. Sci. | 4 |
| 2022 | Lifelong reinforcement learning with temporal logic formulas and reward machines
Xuejing Zheng, Chao Yu 0004, Minjie Zhang 0001 |
Knowl. Based Syst. | 2 |
| 2022 | Multi-Agent Transfer Reinforcement Learning With Multi-View Encoder for Adaptive Traffic Signal ControlabstractMulti-agent reinforcement learning (MARL) based methods for adaptive traffic signal control (ATSC) have shown promising potentials to solve the heavy traffic problems. The existing MARL methods adopt centralized or distributed strategies. The former only models the environment as an agent and suffers from the exponential growth of action and state space. The latter extends the independent reinforcement learning methods, such as DQN, to multiple interactions directly or propagates information, such as state and policy, without taking their qualities into account. In this paper, we propose a multi-agent transfer reinforcement learning method to enhance the performance of MARL for ATSC, which is termed as multi-agent transfer soft actor-critic with the multi-view encoder (MT-SAC). The MT-SAC combines centralized and distributed strategies. In MT-SAC, we propose a multi-view state encoder and a transfer learning paradigm with guidance. The encoder processes input states from multiple perspectives and uses an attention mechanism to weigh the neighborhood information. While the paradigm enables the agents to handle different conditions for improving generalization abilities by transfer learning. Experimental studies on different scale road networks show that the MT-SAC outperforms the state-of-the-art algorithms and makes the traffic signal controllers more collaborative and robust. Hong-Wei Ge, Dongwan Gao, Liang Sun 0003, Yaqing Hou, Chao Yu 0004, Yuxin Wang 0001, Guozhen Tan |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2021 | Inference-based Hierarchical Reinforcement Learning for Cooperative Multi-agent NavigationabstractThis work aims to address the multi-agent cooperative navigation problem (MCNP), where multiple agents work together to occupy the landmarks in an environment without collision and with minimum time consumption. To this end, we propose an inference-based hierarchical reinforcement learning (IHRL) model, in which the high-level component infers the target allocation scheme among the agents and landmarks using a local message-passing algorithm, while the low-level component trains the sub-policy corresponding to the target assigned by the high-level component using traditional RL algorithms. The highlight of our model lies in the interplay of high-level inference based on the knowledge from learning and low-level learning with the results from inference. In this way, the overall learning efficiency can be improved by integrating more indicative information into the agents’ coordinated learning process. Extensive experiments demonstrate the effectiveness of the proposed model. Lijun Xia, Chao Yu 0004, Zifan Wu |
ICTAI | 2 |
| 2021 | Combining Model-Based and Model-Free Reinforcement Learning Policies for More Efficient Sepsis Treatment
Chao Yu 0004, Qikai Huang, Luhao Wang, Xiangdong Guan |
ISBRA | 2 |
| 2021 | Coordinated Proximal Policy OptimizationabstractWe present Coordinated Proximal Policy Optimization (CoPPO), an algorithm that extends the original Proximal Policy Optimization (PPO) to the multi-agent setting. The key idea lies in the coordinated adaptation of step size during the policy update process among multiple agents. We prove the monotonicity of policy improvement when optimizing a theoretically-grounded joint objective, and derive a simplified optimization objective based on a set of approximations. We then interpret that such an objective in CoPPO can achieve dynamic credit assignment among agents, thereby alleviating the high variance issue during the concurrent update of agent policies. Finally, we demonstrate that CoPPO outperforms several strong baselines and is competitive with the latest multi-agent PPO method (i.e. MAPPO) under typical multi-agent settings, including cooperative matrix games and the StarCraft II micromanagement tasks. Zifan Wu, Chao Yu 0004, Deheng Ye, Junge Zhang, Haiyin Piao, Hankui Zhuo |
NeurIPS | 2 |
| 2020 | D3PG: Decomposed Deep Deterministic Policy Gradient for Continuous Control
Yinzhao Dong, Chao Yu 0004, Hong-Wei Ge |
DAI | 2 |
| 2020 | Two-stage Automatic Image Annotation Based on Latent Semantic Scene ClassificationabstractThe rapid growth of multimedia content makes existing automatic image annotation techniques difficult to satisfy the demands of real-world applications. In this paper, we propose a two-stage automatic image annotation algorithm (TAIA) based on latent semantic scene classification. In the offline training phase, the hidden connectivity of labels is firstly excavated by a directed-weighed graph based on label co-occurrence relation matrix, and then the latent scene categories are detected among the labels by using nonnegative matrix factorization. Further, we propose a multi-view extreme learning machine (MELM) to learn the probability that the multiple visual feature maps to the semantic scenes. In the online annotation phase, the image to be annotated is fed to the scene classifier MELM to identify its relevant scenes. Then k-nearest neighbor based annotator is conducted on the relevant scenes to predict labels for the unannotated images. The TAIA is formulated in such a framework so that the relationship between labels and semantic scenes is fully considered, and the hard classification problem is solved. The experimental results on multiple datasets have demonstrated that the proposed framework TAIA is both effective and efficient. Hong-Wei Ge, Kai Zhang 0050, Yaqing Hou, Chao Yu 0004, Mingde Zhao 0002, Zhen Wang 0004, Liang Sun 0003 |
IJCNN | 4 |
| 2020 | Distributed Multiagent Coordinated Learning for Autonomous Driving in Highways Based on Dynamic Coordination GraphsabstractAutonomous driving is one of the most important AI applications and has attracted extensive interest in recent years. A large number of studies have successfully applied reinforcement learning techniques in various aspects of autonomous driving, ranging from low-level control of driving maneuvers to higher level of strategic decision-making. However, comparatively less progress has been made in investigating how co-existing autonomous vehicles would interact with each other in a common environment and how reinforcement learning can be helpful in such situations by applying multiagent reinforcement learning techniques in the high-level strategic decision-making of the following or overtaking for a group of autonomous vehicles in highway scenarios. Learning to achieve coordination among vehicles in such situations is challenging due to the unique feature of vehicular mobility, which renders it infeasible to directly apply the existing coordinated learning approaches. To solve this problem, we propose using dynamic coordination graph to model the continuously changing topology during vehicles' interactions and come up with two basic learning approaches to coordinate the driving maneuvers for a group of vehicles. Several extension mechanisms are then presented to make these approaches workable in a more complex and realistic setting with any number of vehicles. The experimental evaluation has verified the benefits of the proposed coordinated learning approaches, compared with other approaches that learn without coordination or rely on some traditional mobility models based on some expert driving rules. Chao Yu 0004, Xin Wang 0077, Xin Xu 0001, Minjie Zhang 0001, Hong-Wei Ge, Jiankang Ren, Liang Sun 0003, Bingcai Chen, Guozhen Tan |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2019 | Workload-Aware Harmonic Partitioned Scheduling of Periodic Real-Time Tasks with Constrained DeadlinesabstractMultiprocessor platforms have been widely applied in safety-critical domains to accommodate the increasing computation requirement of modern real-time applications. In this paper, we present a workload-aware harmonic partitioned multiprocessor scheduling scheme for periodic real-time tasks with constrained deadlines under the fixed-priority preemptive scheduling policy. In particular, two grouping metrics effectively integrating both harmonicity and workload characteristic are designed to guide our task partition. With those metrics, our scheme can greatly improve system utilization by taking advantage of the combination of harmonic relationship exploration and workload awareness. Experiments show that our proposed scheme significantly outperforms existing approaches in terms of schedulability. Jiankang Ren, Xiaoyan Su, Guoqi Xie, Chao Yu 0004, Guozhen Tan, Guowei Wu 0001 |
DAC | 4 |
| 2019 | Large-Scale Home Energy Management Using Entropy-Based Collective Multiagent Deep Reinforcement Learning FrameworkabstractSmart grids are contributing to the demand-side management by integrating electronic equipment, distributed energy generation and storage and advanced meters and controllers. With the increasing adoption of electric vehicles and distributed energy generation and storage systems, residential energy management is drawing more and more attention, which is regarded as being critical to demand-supply balancing and peak load reduction. In this paper, we focus on a microgrid scenario in which modern homes interact together under a large-scale setting to better optimize their electricity cost. We first make households form a group with an economic stimulus. Then we formulate the energy expense optimization problem of the household community as a multi-agent coordination problem and present an Entropy-Based Collective Multiagent Deep Reinforcement Learning (EB-C-MADRL) framework to address it. Experiments with various real-world data demonstrate that EB-C-MADRL can reduce both the long-term group power consumption cost and daily peak demand effectively compared with existing approaches. Yaodong Yang 0002, Jianye Hao, Yan Zheng 0002, Chao Yu 0004 |
IJCAI | 4 |
| 2019 | The Price of Governance: A Middle Ground Solution to Coordination in Organizational ControlabstractAchieving coordination is crucial in organizational control. This paper investigates a middle ground solution between decentralized interactions and centralized administrations for coordinating agents beyond inefficient behavior. We first propose the price of governance (PoG) to evaluate how such a middle ground solution performs in terms of effectiveness and cost. We then propose a hierarchical supervision framework to explicitly model the PoG, and define step by step how to realize the core principle of the framework and compute the optimal PoG for a control problem. Two illustrative case studies are carried out to exemplify the applications of the proposed framework and its methodology. Results show that the hierarchical supervision framework is capable of promoting coordination among agents while bounding administrative cost to a minimum in different kinds of organizational control problems. Chao Yu 0004, Guozhen Tan |
IJCAI | 1 |
| 2019 | Multi-Grained Cascade AdaBoost Extreme Learning Machine for Feature RepresentationabstractExtreme learning machine (ELM) has been well recognized for characteristics such as less training parameters, fast training speed and strong generalization ability. Due to its high efficiency, researchers have embedded ELMs into deep learning frameworks to address the problems of high time-consumption and computational complexities that are encountered in the traditional deep neural networks. However, existing ELM-based deep learning algorithms usually neglect the spatial relationship of original data. In this paper, we propose a multi-grained cascade AdaBoost based weighted ELM algorithm (gcAWELM) for feature representation. We use AdaBoost based weighted ELM as a basic module to construct cascade structure for feature learning. Different ensemble ELMs trend to extract varied features. Moreover, multi-grained scanning is employed to exploit the spatial structure of the original data. The gcAWELM can determine the number of cascade levels adaptively and has simpler structure and fewer parameters compared with the traditional deep models. The results on image datasets with different scales show that the gcAWELM can achieve competitive performance for different learning tasks even with the same parameter settings. Hong-Wei Ge, Weiting Sun, Mingde Zhao 0002, Kai Zhang 0050, Liang Sun 0003, Chao Yu 0004 |
IJCNN | 6 |
| 2019 | Execution allowance based fixed priority scheduling for probabilistic real-time systems
Jiankang Ren, Zichuan Xu, Chao Yu 0004, Chi Lin 0001, Guowei Wu 0001, Guozhen Tan |
J. Syst. Softw. | 3 |
| 2018 | Decentralized Multiagent Reinforcement Learning for Efficient Robotic Control by Coordination Graphs
Chao Yu 0004, Jiankang Ren, Hong-Wei Ge, Liang Sun 0003 |
PRICAI (1) | 1 |
| 2018 | Adaptively Shaping Reinforcement Learning Agents via Human Reward
Chao Yu 0004, Tianpei Yang, Wenxuan Zhu, Yuchen Li 0006, Hong-Wei Ge, Jiankang Ren |
PRICAI (1) | 1 |
| 2018 | Efficient and Robust Emergence of Norms through Heuristic Collective LearningabstractIn multiagent systems, social norms serves as an important technique in regulating agents’ behaviors to ensure effective coordination among agents without a centralized controlling mechanism. In such a distributed environment, it is important to investigate how a desirable social norm can be synthesized in a bottom-up manner among agents through repeated local interactions and learning techniques. In this article, we propose two novel learning strategies under the collective learning framework, collective learning EV-l and collective learning EV-g , to efficiently facilitate the emergence of social norms. Extensive simulations results show that both learning strategies can support the emergence of desirable social norms more efficiently and be applicable in a wider range of multiagent interaction scenarios compared with previous work. The influence of different topologies is investigated, which shows that the performance of all strategies is robust across different network topologies. The influences of a number of key factors (neighborhood size, actions space, population size, fixed agents and isolated subpopulations) on norm emergence performance are investigated as well. Jianye Hao, Jun Sun 0001, Guangyong Chen, Chao Yu 0004, Zhong Ming 0001 |
ACM Trans. Auton. Adapt. Syst. | 5 |
| 2017 | APDM: An adaptive multi-priority distributed multichannel MAC protocol for vehicular ad hoc networks in unsaturated conditions
Caixia Song, Guozhen Tan, Chao Yu 0004, Nan Ding 0001, Fuxin Zhang |
Comput. Commun. | 3 |
| 2016 | Accelerating Norm Emergence Through Hierarchical Heuristic LearningabstractSocial norms serve as an important mechanism to regulate the behaviours of agents and to facilitate coordination among them in multiagent systems. One important research question is how a norm can rapidly emerge through repeated local interaction within agent societies under different environments when their coordination space becomes large. To address this problem, we propose a hierarchically heuristic learning strategy (HHLS) under the hierarchical social learning framework. Subordinate agents report their information to their supervisors, while supervisors can generate instructions (rules and suggestions) based on the information collected from their subordinates. Subordinate agents heuristically update their strategies based on both their own experience and the instructions from their supervisors. Extensive experiment evaluations show that HHLS can support the emergence of desirable social norms more efficiently and can be applicable in a much wider range of multiagent interaction scenarios compared with previous work. The influence of key related factors (e.g., different topologies, population, neighbourhood and action space size, cluster size) are also investigated and new insights are obtained as well. Tianpei Yang, Zhaopeng Meng, Jianye Hao, Sandip Sen, Chao Yu 0004 |
ECAI | 5 |
| 2016 | Adaptive Learning for Efficient Emergence of Social Norms in Networked Multiagent Systems
Chao Yu 0004, Hongtao Lv, Sandip Sen, Fenghui Ren, Guozhen Tan |
PRICAI | 1 |
| 2015 | Multiagent Learning of Coordination in Loosely Coupled Multiagent SystemsabstractMultiagent learning (MAL) is a promising technique for agents to learn efficient coordinated behaviors in multiagent systems (MASs). In MAL, concurrent multiple distributed learning processes can make the learning environment nonstationary for each individual learner. Developing an efficient learning approach to coordinate agents' behaviors in this dynamic environment is a difficult problem, especially when agents do not know the domain structure and have only local observability of the environment. In this paper, a coordinated MAL approach is proposed to enable agents to learn efficient coordinated behaviors by exploiting agent independence in loosely coupled MASs. The main feature of the proposed approach is to explicitly quantify and dynamically adapt agent independence during learning so that agents can make a trade-off between a single-agent learning process and a coordinated learning process for an efficient decision making. The proposed approach is employed to solve two-robot navigation problems in different scales of domains. Experimental results show that agents using the proposed approach can learn to act in concert or independently in different areas of the environment, which results in great computational savings and near optimal performance. Chao Yu 0004, Minjie Zhang 0001, Fenghui Ren, Guozhen Tan |
IEEE Trans. Cybern. | 1 |
| 2015 | Emotional Multiagent Reinforcement Learning in Spatial Social DilemmasabstractSocial dilemmas have attracted extensive interest in the research of multiagent systems in order to study the emergence of cooperative behaviors among selfish agents. Understanding how agents can achieve cooperation in social dilemmas through learning from local experience is a critical problem that has motivated researchers for decades. This paper investigates the possibility of exploiting emotions in agent learning in order to facilitate the emergence of cooperation in social dilemmas. In particular, the spatial version of social dilemmas is considered to study the impact of local interactions on the emergence of cooperation in the whole system. A double-layered emotional multiagent reinforcement learning framework is proposed to endow agents with internal cognitive and emotional capabilities that can drive these agents to learn cooperative behaviors. Experimental results reveal that various network topologies and agent heterogeneities have significant impacts on agent learning behaviors in the proposed framework, and under certain circumstances, high levels of cooperation can be achieved among the agents. Chao Yu 0004, Minjie Zhang 0001, Fenghui Ren, Guozhen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2014 | Coordinated learning by exploiting sparse interaction in multiagent systemsabstractSUMMARY Multiagent learning provides a promising paradigm to study how autonomous agents learn to achieve coordinated behavior in multiagent systems. In multiagent learning, the concurrency of multiple distributed learning processes makes the environment nonstationary for each individual learner. Developing an efficient learning approach to coordinate agents’ behavior in this dynamic environment is a difficult problem especially when agents do not know the domain structure and at the same time have only local observability of the environment. In this paper, a coordinated learning approach is proposed to enable agents to learn where and how to coordinate their behavior in loosely coupled multiagent systems where the sparse interactions of agents constrain coordination to some specific parts of the environment. In the proposed approach, an agent first collects statistical information to detect those states where coordination is most necessary by considering not only the potential contributions from all the domain states but also the direct causes of the miscoordination in a conflicting state. The agent then learns to coordinate its behavior with others through its local observability of the environment according to different scenarios of state transitions. To handle the uncertainties caused by agents’ local observability, an optimistic estimation mechanism is introduced to guide the learning process of the agents. Empirical studies show that the proposed approach can achieve a better performance by improving the average agent reward compared with an uncoordinated learning approach and by reducing the computational complexity significantly compared with a centralized learning approach. Copyright © 2012 John Wiley & Sons, Ltd. Chao Yu 0004, Minjie Zhang 0001, Fenghui Ren |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | Collective Learning for the Emergence of Social Norms in Networked Multiagent SystemsabstractSocial norms such as social rules and conventions play a pivotal role in sustaining system order by regulating and controlling individual behaviors toward a global consensus in large-scale distributed systems. Systematic studies of efficient mechanisms that can facilitate the emergence of social norms enable us to build and design robust distributed systems, such as electronic institutions and norm-governed sensor networks. This paper studies the emergence of social norms via learning from repeated local interactions in networked multiagent systems. A collective learning framework, which imitates the opinion aggregation process in human decision making, is proposed to study the impact of agent local collective behaviors on the emergence of social norms in a number of different situations. In the framework, each agent interacts repeatedly with all of its neighbors. At each step, an agent first takes a best-response action toward each of its neighbors and then combines all of these actions into a final action using ensemble learning methods. Extensive experiments are carried out to evaluate the framework with respect to different network topologies, learning strategies, numbers of actions, influences of nonlearning agents, and so on. Experimental results reveal some significant insights into the manipulation and control of norm emergence in networked multiagent systems achieved through local collective behaviors. Chao Yu 0004, Minjie Zhang 0001, Fenghui Ren |
IEEE Trans. Cybern. | 1 |
| 2013 | Emotional Multiagent Reinforcement Learning in Social Dilemmas
Chao Yu 0004, Minjie Zhang 0001, Fenghui Ren |
PRIMA | 1 |
| 2012 | Exploiting Independent Relationships in Multiagent Systems for Coordinated Learning
Chao Yu 0004, Minjie Zhang 0001, Fenghui Ren |
PRICAI | 1 |