VLDB 2026 Research / reviewers in the wild / expert
Yuanheng Zhu
dblp:118/5095
· DBLP profile ↗
57ranked-venue papers
14as first author
35since 2021 · last 2026
0000-0001-5384-423XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 10 first-author · 27 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tacit mechanism: Bridging pre-training of individuality to multi-agent adversarial coordination
Shiqing Yao, Jiajun Chai, Haixin Yu, Yongzhe Chang, Tiantian Zhang 0002, Yuanheng Zhu, Xueqian Wang 0001 |
Neural Networks | 6 |
| 2026 | CMIP: Combining Constructive Model With Improvement Policy for Large-Scale Min-Max Multiple Traveling Salesman ProblemabstractThe min-max multiple traveling salesman problem (min-max mTSP) is a significant variant of the min-max routing problem, focusing on minimizing the longest subtour cost among multiple salesmen working cooperatively. This problem is highly relevant in real-world scenarios but is notoriously challenging, especially as the scale increases with numerous salesmen covering thousands of cities. This paper presents a novel approach for solving large-scale min-max mTSP. Our method, based on deep reinforcement learning, introduces a novel two-stage process. In the first stage, we generate an initial solution using a constructive model incorporating global and local attention mechanisms through a gated network. Additionally, we employ multi-task training on a single constructive model across various mTSP problems with differing numbers of salesmen, using weighted task balancing to balance the multi-task learning process. In the second stage, the initial solution is iteratively refined using improvement policy, which re-optimizes the current subtours to form a new better one. To the best of our knowledge, our method is the first capable of handling problems with up to 10,000 nodes. The experimental results demonstrate that our approach achieves the best solution on 71% of the problems in randomly uniform datasets, outperforming all existing methods. Our code is available athttps://github.com/1hhix/CMIP Binbin Zuo, Weifan Li, Jiankuo Zhao, Tianxiang Bai, Linqian Yang, Yuanheng Zhu |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2025 | ARAC: Adaptive Regularized Multi-Agent Soft Actor-Critic in Graph-Structured Adversarial GamesabstractIn graph-structured multi-agent reinforcement learning (MARL) adversarial tasks such as pursuit and confrontation, agents must coordinate under highly dynamic interactions, where sparse rewards hinder efficient policy learning. We propose Adaptive Regularized Multi-Agent Soft Actor-Critic (ARAC), which integrates an attention-based graph neural network (GNN) for modeling agent dependencies with an adaptive divergence regularization mechanism. The GNN enables expressive representation of spatial relations and state features in graph environments. Divergence regularization can serve as policy guidance to alleviate the sparse reward problem, but it may lead to suboptimal convergence when the reference policy itself is imperfect. The adaptive divergence regularization mechanism enables the framework to exploit reference policies for efficient exploration in the early stages, while gradually reducing reliance on them as training progresses to avoid inheriting their limitations. Experiments in pursuit and confrontation scenarios demonstrate that ARAC achieves faster convergence, higher final success rates, and stronger scalability across varying numbers of agents compared with MARL baselines, highlighting its effectiveness in complex graph-structured environments. Ruochuan Shi, Runyu Lu, Yuanheng Zhu, Dongbin Zhao |
DAI | 3 |
| 2025 | RLAE: Reinforcement Learning-Assisted Ensemble for LLMsabstractEnsembling large language models (LLMs) can effectively combine diverse strengths of different models, offering a promising approach to enhance performance across various tasks.However, existing methods typically rely on fixed weighting strategies that fail to adapt to the dynamic, context-dependent characteristics of LLM capabilities.In this work, we propose Reinforcement Learning-Assisted Ensemble for LLMs (RLAE), a novel framework that reformulates LLM ensemble through the lens of a Markov Decision Process (MDP).Our approach introduces a RL agent that dynamically adjusts ensemble weights by considering both input context and intermediate generation states, with the agent being trained using rewards that directly correspond to the quality of final outputs.We implement RLAE using both single-agent and multi-agent reinforcement learning algorithms (RLAE PPO and RLAE MAPPO ), demonstrating substantial improvements over conventional ensemble methods.Extensive evaluations on a diverse set of tasks show that RLAE outperforms existing approaches by up to 3.3% accuracy points, offering a more effective framework for LLM ensembling.Furthermore, our method exhibits superior generalization capabilities across different tasks without the need for retraining, while simultaneously achieving lower time latency.The source code is available at here. Yuqian Fu, Yuanheng Zhu, Jiajun Chai, Guojun Yin, Dongbin Zhao |
EMNLP | 2 |
| 2025 | Empowering LLM Agents with Zero-Shot Optimal Decision-Making through Q-learningabstractLarge language models (LLMs) are trained on extensive text data to gain general comprehension capability. Current LLM agents leverage this ability to make zero- or few-shot decisions without reinforcement learning (RL) but fail in making optimal decisions, as LLMs inherently perform next-token prediction rather than maximizing rewards. In contrast, agents trained via RL could make optimal decisions but require extensive environmental interaction. In this work, we develop an algorithm that combines the zero-shot capabilities of LLMs with the optimal decision-making of RL, referred to as the Model-based LLM Agent with Q-Learning (MLAQ). MLAQ employs Q-learning to derive optimal policies from transitions within memory. However, unlike RL agents that collect data from environmental interactions, MLAQ constructs an imagination space fully based on LLM to perform imaginary interactions for deriving zero-shot policies. Our proposed UCB variant generates high-quality imaginary data through interactions with the LLM-based world model, balancing exploration and exploitation while ensuring a sub-linear regret bound. Additionally, MLAQ incorporates a mixed-examination mechanism to filter out incorrect data. We evaluate MLAQ in benchmarks that present significant challenges for existing LLM agents. Results show that MLAQ achieves a optimal rate of over 90\% in tasks where other methods struggle to succeed. Additional experiments are conducted to reach the conclusion that introducing model-based RL into LLM agents shows significant potential to improve optimal decision-making ability. Our interactive website is available at http://mlaq.site. Jiajun Chai, Yuqian Fu, Dongbin Zhao, Yuanheng Zhu |
ICLR | 5 |
| 2025 | INS: Interaction-aware Synthesis to Enhance Offline Multi-agent Reinforcement LearningabstractData scarcity in offline multi-agent reinforcement learning (MARL) is a key challenge for real-world applications. Recent advances in offline single-agent reinforcement learning (RL) demonstrate the potential of data synthesis to mitigate this issue.
However, in multi-agent systems, interactions between agents introduce additional challenges. These interactions complicate the synthesis of multi-agent datasets, leading to data distortion when inter-agent interactions are neglected. Furthermore, the quality of the synthetic dataset is often constrained by the original dataset. To address these challenges, we propose **INteraction-aware Synthesis (INS)**, which synthesizes high-quality multi-agent datasets using diffusion models. Recognizing the sparsity of inter-agent interactions, INS employs a sparse attention mechanism to capture these interactions, ensuring that the synthetic dataset reflects the underlying agent dynamics. To overcome the limitation of diffusion models requiring continuous variables, INS implements a bit action module, enabling compatibility with both discrete and continuous action spaces. Additionally, we incorporate a select mechanism to prioritize transitions with higher estimated values, further enhancing the dataset quality. Experimental results across multiple datasets in MPE and SMAC environments demonstrate that INS consistently outperforms existing methods, resulting in improved downstream policy performance and superior dataset metrics. Notably, INS can synthesize high-quality data using only 10% of the original dataset, highlighting its efficiency in data-limited scenarios. Yuqian Fu, Yuanheng Zhu, Jiajun Chai, Dongbin Zhao |
ICLR | 2 |
| 2025 | Divergence-Regularized Discounted Aggregation: Equilibrium Finding in Multiplayer Partially Observable Stochastic GamesabstractThis paper presents Divergence-Regularized Discounted Aggregation (DRDA), a multi-round learning system for solving partially observable stochastic games (POSGs). DRDA is based on action values and applicable to multiplayer POSGs, which can unify normal-form games (NFGs), extensive-form games (EFGs) with perfect recall, and Markov games (MGs). In each single round, DRDA can be viewed as a discounted variant of Follow the Regularized Leader (FTRL) under a general value function for POSGs. While previous studies on discounted FTRL have demonstrated its last-iterate convergence towards quantal response equilibrium (QRE) in NFGs, this paper extends the theoretical results to POSGs under divergence regularization and generalizes the QRE concept of Nash distribution. The linear last-iterate convergence of single-round DRDA to its rest point is proved under the assumption on the hypomonotonicity of the game. When the rest point is unique, it induces the unique Nash distribution defined in the POSG, which has a bounded deviation from Nash equilibrium (NE). Under multiple learning rounds, DRDA keeps replacing the base policy for divergence regularization with the policy at the rest point in the previous round. It is further proved that the limit point of multi-round DRDA must be an exact NE (rather than a QRE). In experiments, discrete-time DRDA can converge to NE at a near-exponential rate in (multiplayer) NFGs and outperform the existing baselines for EFGs, MGs, and typical POSGs. Runyu Lu, Yuanheng Zhu, Dongbin Zhao |
ICLR | 2 |
| 2025 | Constrained Exploitability Descent: An Offline Reinforcement Learning Method for Finding Mixed-Strategy Nash EquilibriumabstractThis paper proposes Constrained Exploitability Descent (CED), a model-free offline reinforcement learning (RL) algorithm for solving adversarial Markov games (MGs). CED combines the game-theoretical approach of Exploitability Descent (ED) with policy constraint methods from offline RL. While policy constraints can perturb the optimal pure-strategy solutions in single-agent scenarios, we find the side effect less detrimental in adversarial games, where the optimal policy can be a mixed-strategy Nash equilibrium. We theoretically prove that, under the uniform coverage assumption on the dataset, CED converges to a stationary point in deterministic two-player zero-sum Markov games. We further prove that the min-player policy at the stationary point follows the property of mixed-strategy Nash equilibrium in MGs. Compared to the model-based ED method that optimizes the max-player policy, our CED method no longer relies on a generalized gradient. Experiments in matrix games, a tree-form game, and an infinite-horizon soccer game verify that CED can find an equilibrium policy for the min-player as long as the offline dataset guarantees uniform coverage. Besides, CED achieves a significantly lower NashConv compared to an existing pessimism-based method and can gradually improve the behavior policy even under non-uniform data coverages. When combined with neural networks, CED also outperforms behavior cloning and offline self-play in a large-scale two-team robotic combat game. Runyu Lu, Yuanheng Zhu, Dongbin Zhao |
ICML | 2 |
| 2025 | DipLLM: Fine-Tuning LLM for Strategic Decision-making in DiplomacyabstractDiplomacy is a complex multiplayer game that re- quires both cooperation and competition, posing significant challenges for AI systems. Traditional methods rely on equilibrium search to generate extensive game data for training, which demands substantial computational resources. Large Lan- guage Models (LLMs) offer a promising alterna- tive, leveraging pre-trained knowledge to achieve strong performance with relatively small-scale fine-tuning. However, applying LLMs to Diplo- macy remains challenging due to the exponential growth of possible action combinations and the intricate strategic interactions among players. To address this challenge, we propose DipLLM, a fine-tuned LLM-based agent that learns equilib- rium policies for Diplomacy. DipLLM employs an autoregressive factorization framework to sim- plify the complex task of multi-unit action assign- ment into a sequence of unit-level decisions. By defining an equilibrium policy within this frame- work as the learning objective, we fine-tune the model using only 1.5% of the data required by the state-of-the-art Cicero model, surpassing its per- formance. Our results demonstrate the potential of fine-tuned LLMs for tackling complex strategic decision-making in multiplayer games. Kaixuan Xu, Jiajun Chai, Yuqian Fu, Yuanheng Zhu, Dongbin Zhao |
ICML | 5 |
| 2025 | Learning Pre-Trained Tacit Behavior for Efficient Multi-Agent Adversarial Coordination
Shiqing Yao, Jiajun Chai, Haixin Yu, Yongzhe Chang, Yuanheng Zhu, Xueqian Wang 0001 |
AAMAS | 5 |
| 2025 | Offline Goal-Conditioned Reinforcement Learning with Elastic-Subgoal Diffused Policy Learning
Yaocheng Zhang, Yuanheng Zhu, Yuqian Fu, Songjun Tu, Dongbin Zhao |
AAMAS | 2 |
| 2025 | Equilibrium Policy Generalization: A Reinforcement Learning Framework for Cross-Graph Zero-Shot Generalization in Pursuit-Evasion GamesabstractEquilibrium learning in adversarial games is an important topic widely examined in the fields of game theory and reinforcement learning (RL). Pursuit-evasion game (PEG), as an important class of real-world games from the fields of robotics and security, requires exponential time to be accurately solved. When the underlying graph structure varies, even the state-of-the-art RL methods require recomputation or at least fine-tuning, which can be time-consuming and impair real-time applicability. This paper proposes an Equilibrium Policy Generalization (EPG) framework to effectively learn a generalized policy with robust cross-graph zero-shot performance. In the context of PEGs, our framework is generally applicable to both pursuer and evader sides in both no-exit and multi-exit scenarios. These two generalizability properties, to our knowledge, are the first to appear in this domain. The core idea of the EPG framework is to train an RL policy across different graph structures against the equilibrium policy for each single graph. To construct an equilibrium oracle for single-graph policies, we present a dynamic programming (DP) algorithm that provably generates pure-strategy Nash equilibrium with near-optimal time complexity. To guarantee scalability with respect to pursuer number, we further extend DP and RL by designing a grouping mechanism and a sequence model for joint policy decomposition, respectively. Experimental results show that, using equilibrium guidance and a distance feature proposed for cross-graph PEG training, the EPG framework guarantees desirable zero-shot performance in various unseen real-world graphs. Besides, when trained under an equilibrium heuristic proposed for the graphs with exits, our generalized pursuer policy can even match the performance of the fine-tuned policies from the state-of-the-art PEG methods. Runyu Lu, Peng Zhang 0127, Ruochuan Shi, Yuanheng Zhu, Dongbin Zhao, Yang Liu 0066, Dong Wang 0004, Cesare Alippi |
NeurIPS | 4 |
| 2025 | Learning and Planning Multi-Agent Tasks via an MoE-based World ModelabstractMulti-task multi-agent reinforcement learning (MT-MARL) aims to develop a single model capable of solving a diverse set of tasks. However, existing methods often fall short due to the substantial variation in optimal policies across tasks, making it challenging for a single policy model to generalize effectively. In contrast, we find that many tasks exhibit **bounded similarity** in their underlying dynamics—highly similar within certain groups (e.g., door-open/close) diverge significantly between unrelated tasks (e.g., door-open \& object-catch). To leverage this property, we reconsider the role of modularity in multi-task learning, and propose **M3W**, a novel approach that applies mixture-of-experts (MoE) to world model instead of policy, enabling both learning and planning. For learning, it uses a SoftMoE-based dynamics model alongside a SparseMoE-based predictor to facilitate knowledge reuse across similar tasks while avoiding gradient conflicts across dissimilar tasks. For planning, it evaluates and optimizes actions using the predicted rollouts from the world model, without relying directly on a explicit policy model, thereby overcoming the limitations of policy-centric methods. As the first MoE-based multi-task world model, M3W demonstrates superior performance, sample efficiency, and multi-task adaptability, as validated on Bi-DexHands with 14 tasks and MA-Mujoco with 24 tasks. The demos and anonymous code are available at \url{https://github.com/zhaozijie2022/m3w-marl}. Zijie Zhao 0001, Zhongyue Zhao, Kaixuan Xu, Yuqian Fu, Jiajun Chai, Yuanheng Zhu, Dongbin Zhao |
NeurIPS | 6 |
| 2025 | Multi-Task Multi-Agent Reinforcement Learning With Task-Entity Transformers and Value Decomposition TrainingabstractMulti-task multi-agent reinforcement learning aims to control multiple agents to perform well on multiple tasks. It encounters three core challenges: the varying number of agents and entities, the disparities in cooperative behaviors among different tasks, and the training imbalance caused by varying task difficulty levels. To address these issues, we propose a novel framework named Task-Entity Transformer Qmix (TETQmix), which employs pretrained language models for task encoding, utilizes proposed Task-Entity Transformer to handle observations across various tasks, and adjusts task learning weights to achieve balanced multi-task training. Task-Entity Transformer not only enables handling multi-task scenarios with varying numbers of agents and entities, but also leverages cross-attention modules to integrate observation and task embeddings, so that each agent can obtain individual values and decisions for multiple tasks. We then utilize a transformer-based mixer to monotonically combine the individual values, and train the whole network’s parameters using temporal-difference errors. To facilitate multi-task training, we define task regret as the difference between the current-stage return and the candidate best one, and adjust the learning weight of each task based on its task regret. Experiments are conducted on both simulated multi-particle environments and real-world multi-robot systems. Compared with existing baselines, our method not only is superior in multi-task learning efficiency, but also shows promising transfer ability on unseen tasks. Note to Practitioners—The flexibility of multi-agent systems makes them quite fit to multiple tasks. Compared to designing different decision models for different tasks, it is more convenient if one can use just one decision model to resolve multiple tasks. Besides, it can make the maximum utilization of trajectory data coming from similar tasks when the data are integrated for multi-task decision model training. Natural language provides a powerful tool to describe the task context and emphasize the similarities or differences among different tasks. Pretrained language models can encode the task context, based on which the decision model can adjust its output distribution for different tasks and even synthesize the decisions from existing and similar tasks to achieve promising zero-shot and few-shot transfer performance for unseen tasks. With our proposed TETQmix, practitioners are able to realize multi-task capability in multi-agent systems and increase the generalization in a variety of scenarios. Yuanheng Zhu, Shangjing Huang, Binbin Zuo, Dongbin Zhao, Changyin Sun 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | Boosting On-Policy Actor-Critic With Shallow Updates in CriticabstractDeep reinforcement learning (DRL) benefits from the representation power of deep neural networks (NNs), to approximate the value function and policy in the learning process. Batch reinforcement learning (BRL) benefits from stable training and data efficiency with fixed representation and enjoys solid theoretical analysis. This work proposes least-squares deep policy gradient (LSDPG), a hybrid approach that combines least-squares reinforcement learning (RL) with online DRL to achieve the best of both worlds. LSDPG leverages a shared network to share useful features between policy (actor) and value function (critic). LSDPG learns policy, value function, and representation separately. First, LSDPG views deep NNs of the critic as a linear combination of representation weighted by the weights of the last layer and performs policy evaluation with regularized least-squares temporal difference (LSTD) methods. Second, arbitrary policy gradient algorithms can be applied to improve the policy. Third, an auxiliary task is used to periodically distill the features from the critic into the representation. Unlike most DRL methods, where the critic algorithms are often used in a nonstationary situation, i.e., the policy to be evaluated is changing, the critic in LSDPG is working on a stationary case in each iteration of the critic update. We prove that, under some conditions, the critic converges to the regularized TD fixpoint of current policy, and the actor converges to the local optimal policy. The experimental results on challenging Procgen benchmark illustrate the improvement of sample efficiency of LSDPG over proximal policy optimization and phasic policy gradient (PPG). Luntong Li, Yuanheng Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Last-Iterate Convergence to Approximate Nash Equilibria in Multiplayer Imperfect Information GamesabstractImperfect information and multiple players are the two common features of real-world games. However, few of the existing game-theoretic methods are applicable to multiplayer imperfect information games (IIGs) when it comes to finding Nash equilibria. Moreover, the commonly used methods that rely on average-iterate convergence are not conducive to deep reinforcement learning (DRL), which is widely applied to large-scale problems, as it is costly to preserve average policies under function approximation. To deal with these problems, we construct a continuous-time dynamic named imperfect-information exponential-decay score-based learning (IESL) by considering the concept of Nash distribution [a type of quantal response equilibrium (QRE)] in IIGs. Theoretically, we prove the last-iterate convergence of IESL to approximate Nash equilibria in multiplayer IIGs under the assumption of individual concavity. Empirically, we verify that IESL converges in six poker scenarios, with the ultimate NashConv lower than that of the comparative methods (including counterfactual regret minimization (CFR), replicator dynamics (RDs), and their variants) in multiplayer Leduc hold'em. When compared with the existing equilibrium-finding algorithms in multiplayer normal-form games (NFGs), IESL also demonstrates a more stable performance. In addition, we observe a trade-off between the difficulty of IESL's last-iterate convergence and the NashConv of the convergent policies, which aligns with our convergence analysis based on the hypomonotonicity of the game. Runyu Lu, Yuanheng Zhu, Dongbin Zhao, Yu Liu 0005, You He 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Meta Learning Task Representation in Multiagent Reinforcement Learning: From Global Inference to Local InferenceabstractMultiagent meta reinforcement learning (MAMRL) enables multiagent systems (MASs) to adapt to multiple tasks. However, partial observability poses a significant challenge by hindering efficient task inference from agents' limited local experiences. To address this, we propose MG2L, a novel algorithm featuring a global-to-local (G2L) training scheme based on mutual information optimization (MIO). We first extend the centralized training and decentralized execution (CTDE) framework to MAMRL, and introduce a multilevel task encoder for joint global and local task inference. Building on this encoder, the MG2L scheme employs tailored loss functions to optimize task representations. For global inference, the MAS learns a centralized global representation by maximizing the MI between the representation and the task context. For local inference, we formulate conditional MI reduction to quantify the G2L gap. Agents then learn the local representation by minimizing this reduction. The MG2L scheme effectively harmonizes centralized training with decentralized execution, offering a versatile solution for MAMRL challenges. Additionally, we integrate a permutation-invariant attention (PIA) module into the task encoder to reduce sensitivity to behavior policy variations. Extensive experiments-including comparative analyses, ablation studies, meta-test evaluations, and visualizations-demonstrate MG2L's effectiveness. The implementation of MG2L is publicly available at https://github.com/zhaozijie2022/mg2l. Zijie Zhao 0001, Yuqian Fu, Jiajun Chai, Yuanheng Zhu, Dongbin Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Discretizing Continuous Action Space With Unimodal Probability Distributions for On-Policy Reinforcement LearningabstractFor on-policy reinforcement learning (RL), discretizing action space for continuous control can easily express multiple modes and is straightforward to optimize. However, without considering the inherent ordering between the discrete atomic actions, the explosion in the number of discrete actions can possess undesired properties and induce a higher variance for the policy gradient (PG) estimator. In this article, we introduce a straightforward architecture that addresses this issue by constraining the discrete policy to be unimodal using Poisson probability distributions. This unimodal architecture can better leverage the continuity in the underlying continuous action space using explicit unimodal probability distributions. We conduct extensive experiments to show that the discrete policy with the unimodal probability distribution provides significantly faster convergence and higher performance for on-policy RL algorithms in challenging control tasks, especially in highly complex tasks such as Humanoid. We provide theoretical analysis on the variance of the PG estimator, which suggests that our attentively designed unimodal discrete policy can retain a lower variance and yield a stable learning process. Yuanyang Zhu, Zhi Wang 0001, Yuanheng Zhu, Chunlin Chen 0001, Dongbin Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | LDR: Learning Discrete Representation to Improve Noise Robustness in Multiagent TasksabstractIn real-world applications of multiagent reinforcement learning (MARL), agents often face inaccurate environments due to unavoidable noise, presenting a challenge to their robustness. However, limited prior work focuses on addressing such noise in observations, hindering the deployment of multiagent systems. In this article, we propose a method named learning discrete representation (LDR) to improve robustness against noise in multiagent tasks. Specifically, LDR employs a quantization module with a segment mechanism to encode observations and teammate actions, generating discrete representations from learnable codebooks. These representations are subsequently processed via a combiner for decision-making. Through discretization, LDR is able to mitigate the impact of minor noise on decision-making. To enhance the learning efficiency, we incorporate a set-input block that treats the joint observations of agents as a permutation-invariant set, thereby reducing the complexity of the joint observation space. Additionally, we theoretically analyze the expressiveness of discrete representation and the boundedness of discrete distortion. We evaluate the proposed method on StarCraft II micromanagement tasks and multiagent MuJoCo with noisy observations. Empirical results demonstrate that LDR outperforms existing algorithms, improving robustness in noisy cooperative MARL tasks while maintaining superior performance in clean observations. Yuqian Fu, Yuanheng Zhu, Jiajun Chai, Dongbin Zhao |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2024 | Meta-DT: Offline Meta-RL as Conditional Sequence Modeling with World Model DisentanglementabstractA longstanding goal of artificial general intelligence is highly capable generalists that can learn from diverse experiences and generalize to unseen tasks. The language and vision communities have seen remarkable progress toward this trend by scaling up transformer-based models trained on massive datasets, while reinforcement learning (RL) agents still suffer from poor generalization capacity under such paradigms. To tackle this challenge, we propose Meta Decision Transformer (Meta-DT), which leverages the sequential modeling ability of the transformer architecture and robust task representation learning via world model disentanglement to achieve efficient generalization in offline meta-RL. We pretrain a context-aware world model to learn a compact task representation, and inject it as a contextual condition to the causal transformer to guide task-oriented sequence generation. Then, we subtly utilize history trajectories generated by the meta-policy as a self-guided prompt to exploit the architectural inductive bias. We select the trajectory segment that yields the largest prediction error on the pretrained world model to construct the prompt, aiming to encode task-specific information complementary to the world model maximally. Notably, the proposed framework eliminates the requirement of any expert demonstration or domain knowledge at test time. Experimental results on MuJoCo and Meta-World benchmarks across various dataset types show that Meta-DT exhibits superior few and zero-shot generalization capacity compared to strong baselines while being more practical with fewer prerequisites. Our code is available at https://github.com/NJU-RL/Meta-DT. Zhi Wang 0001, Yuanheng Zhu, Dongbin Zhao, Chunlin Chen 0001 |
NeurIPS | 4 |
| 2024 | NVIF: Neighboring Variational Information Flow for Cooperative Large-Scale Multiagent Reinforcement LearningabstractCommunication-based multiagent reinforcement learning (MARL) has shown promising results in promoting cooperation by enabling agents to exchange information. However, the existing methods have limitations in large-scale multiagent systems due to high information redundancy, and they tend to overlook the unstable training process caused by the online-trained communication protocol. In this work, we propose a novel method called neighboring variational information flow (NVIF), which enhances communication among neighboring agents by providing them with the maximum information set (MIS) containing more information than the existing methods. NVIF compresses the MIS into a compact latent state while adopting neighboring communication. To stabilize the overall training process, we introduce a two-stage training mechanism. We first pretrain the NVIF module using a randomly sampled offline dataset to create a task-agnostic and stable communication protocol, and then use the pretrained protocol to perform online policy training with RL algorithms. Our theoretical analysis indicates that NVIF-proximal policy optimization (PPO), which combines NVIF with PPO, has the potential to promote cooperation with agent-specific rewards. Experiment results demonstrate the superiority of our method in both heterogeneous and homogeneous settings. Additional experiment results also demonstrate the potential of our method for multitask learning. Jiajun Chai, Yuanheng Zhu, Dongbin Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Policy Representation Opponent Shaping via Contrastive Learning
Yuanheng Zhu |
ICONIP (9) | 2 |
| 2023 | NeuronsMAE: A Novel Multi-Agent Reinforcement Learning Environment for Cooperative and Competitive Multi-Robot TasksabstractMulti-agent reinforcement learning (MARL) has achieved remarkable success in various challenging problems. Meanwhile, more and more benchmarks have emerged and provided some standards to evaluate the algorithms in different fields. On the one hand, the virtual MARL environments lack knowledge of real-world tasks and actuator abilities. On the other hand, the current task-specified multi-robot platform has poor support for the universality of multi-agent reinforcement learning algorithms and lacks support for transferring from simulation to the real environment. Bridging the gap between the virtual MARL environments and the real multi-robot platform becomes the key to promoting the practicability of MARL algorithms. This paper proposes a novel MARL environment for real multi-robot tasks named NeuronsMAE (Neurons Multi-Agent Environment). This environment supports cooperative and competitive multi-robot tasks and is configured with rich parameter interfaces to study the multi-agent policy transfer from simulation to reality. With this platform, we evaluate various popular MARL algorithms and build a new MARL benchmark for multi-robot tasks. We hope that this platform will facilitate the research and application of MARL algorithms for real robot tasks. Information about the benchmark and the open-source code are released at https://github.com/DRL-CASIA/NeuronsMAE. Guangzheng Hu, Haoran Li 0010, Yuanheng Zhu, Dongbin Zhao |
IJCNN | 4 |
| 2023 | Advantage Constrained Proximal Policy Optimization in Multi-Agent Reinforcement LearningabstractWe investigate the integration of value-based and policy gradient methods in multi-agent reinforcement learning (MARL). The Individual-Global-Max (IGM) principle plays an important role in value-based MARL, as it ensures consistency between joint and local action values. IGM is difficult to guarantee in multi-agent policy gradient methods due to stochastic exploration and conflicting gradient directions. In this paper, we propose a novel multi-agent policy gradient algorithm called Advantage Constrained Proximal Policy Optimization (ACPPO). ACPPO calculates each agent's current local state-action advantage based on their advantage network and estimates the joint state-action advantage based on multi-agent advantage decomposition lemma. According to the consistency of the estimated joint-action advantage and local advantage, the coefficient of each agent constrains the joint-action advantage. ACPPO, unlike previous policy gradient MARL algorithms, does not require an additional sampled baseline to reduce variance or a sequential scheme to improve accuracy. The proposed method is evaluated using the continuous matrix game, the Starcraft Multi-Agent Challenge, and the Multi-Agent MuJoCo task. ACPPO outperforms baselines such as MAPPO, MADDPG, and HATRPO, according to the results. Weifan Li, Yuanheng Zhu, Dongbin Zhao |
IJCNN | 2 |
| 2023 | Enhanced Rolling Horizon Evolution Algorithm With Opponent Model Learning: Results for the Fighting Game AI CompetitionabstractThe Fighting Game AI Competition (FTGAIC) provides a challenging benchmark for two-player video game artificial intelligence. The challenge arises from the large action space, diverse styles of characters and abilities, and the real-time nature of the game. In this article, we propose a novel algorithm that combines the rolling horizon evolution algorithm (RHEA) with opponent model learning. The approach is readily applicable to any two-player video game. In contrast to conventional RHEA, an opponent model is proposed and is optimized by supervised learning with cross-entropy and reinforcement learning with policy gradient and Q-learning respectively, based on history observations from opponent. The model is learned during the live gameplay. With the learned opponent model, the extended RHEA is able to make more realistic plans based on what the opponent is likely to do. This tends to lead to better results. We compared our approach directly with the bots from the FTGAIC 2018 competition and found our method to significantly outperform all of them for all three characters. Furthermore, our proposed bot with the policy gradient based opponent model is the only one without using Monte Carlo tree search among the top five bots in the 2019 competition in which it achieved second place, while using much less domain knowledge than the winner. Zhentao Tang, Yuanheng Zhu, Dongbin Zhao, Simon M. Lucas |
IEEE Trans. Games | 2 |
| 2023 | Empirical Policy Optimization for n-Player Markov GamesabstractIn single-agent Markov decision processes, an agent can optimize its policy based on the interaction with the environment. In multiplayer Markov games (MGs), however, the interaction is nonstationary due to the behaviors of other players, so the agent has no fixed optimization objective. The challenge becomes finding equilibrium policies for all players. In this research, we treat the evolution of player policies as a dynamical process and propose a novel learning scheme for Nash equilibrium. The core is to evolve one's policy according to not just its current in-game performance, but an aggregation of its performance over history. We show that for a variety of MGs, players in our learning scheme will provably converge to a point that is an approximation to Nash equilibrium. Combined with neural networks, we develop an empirical policy optimization algorithm, which is implemented in a reinforcement-learning framework and runs in a distributed way, with each player optimizing its policy based on own observations. We use two numerical examples to validate the convergence property on small-scale MGs, and a pong example to show the potential on large games. Yuanheng Zhu, Weifan Li, Mengchen Zhao, Jianye Hao, Dongbin Zhao |
IEEE Trans. Cybern. | 1 |
| 2023 | UNMAS: Multiagent Reinforcement Learning for Unshaped Cooperative ScenariosabstractMultiagent reinforcement learning methods, such as VDN, QMIX, and QTRAN, that adopt centralized training with decentralized execution (CTDE) framework have shown promising results in cooperation and competition. However, in some multiagent scenarios, the number of agents and the size of the action set actually vary over time. We call these unshaped scenarios, and the methods mentioned above fail in performing satisfyingly. In this article, we propose a new method, called Unshaped Networks for Multiagent Systems (UNMAS), which adapts to the number and size changes in multiagent systems. We propose the self-weighting mixing network to factorize the joint action-value. Its adaption to the change in agent number is attributed to the nonlinear mapping from each-agent Q value to the joint action-value with individual weights. Besides, in order to address the change in an action set, each agent constructs an individual action-value network that is composed of two streams to evaluate the constant environment-oriented subset and the varying unit-oriented subset. We evaluate UNMAS on various StarCraft II micromanagement scenarios and compare the results with several state-of-the-art MARL algorithms. The superiority of UNMAS is demonstrated by its highest winning rates especially on the most difficult scenario 3s5z_vs_3s6z. The agents learn to perform effectively cooperative behaviors, while other MARL algorithms fail. Animated demonstrations and source code are provided in https://sites.google.com/view/unmas. Jiajun Chai, Weifan Li, Yuanheng Zhu, Dongbin Zhao, Kewu Sun, Jishiyu Ding |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Event-Triggered Communication Network With Limited-Bandwidth Constraint for Multi-Agent Reinforcement LearningabstractCommunicating agents with each other in a distributed manner and behaving as a group are essential in multi-agent reinforcement learning. However, real-world multi-agent systems suffer from restrictions on limited bandwidth communication. If the bandwidth is fully occupied, some agents are not able to send messages promptly to others, causing decision delay and impairing cooperative effects. Recent related work has started to address the problem but still fails in maximally reducing the consumption of communication resources. In this article, we propose an event-triggered communication network (ETCNet) to enhance communication efficiency in multi-agent systems by communicating only when necessary. For different task requirements, two paradigms of the ETCNet framework, event-triggered sending network (ETSNet) and event-triggered receiving network (ETRNet), are proposed for learning efficient sending and receiving protocols, respectively. Leveraging the information theory, the limited bandwidth is translated to the penalty threshold of an event-triggered strategy, which determines whether an agent at each step participates in communication or not. Then, the design of the event-triggered strategy is formulated as a constrained Markov decision problem and reinforcement learning finds the feasible and optimal communication protocol that satisfies the limited bandwidth constraint. Experiments on typical multi-agent tasks demonstrate that ETCNet outperforms other methods in reducing bandwidth occupancy and still preserves the cooperative performance of multi-agent systems at the most. Guangzheng Hu, Yuanheng Zhu, Dongbin Zhao, Mengchen Zhao, Jianye Hao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | A Hierarchical Deep Reinforcement Learning Framework for 6-DOF UCAV Air-to-Air CombatabstractUnmanned combat air vehicle (UCAV) combat is a challenging scenario with high-dimensional continuous state and action space and highly nonlinear dynamics. In this article, we propose a general hierarchical framework to resolve the within-vision-range (WVR) air-to-air combat problem under six dimensions of degree (6-DOF) dynamics. The core idea is to divide the whole decision-making process into two loops and use reinforcement learning (RL) to solve them separately. The outer loop uses a combat policy to decide the macro command according to the current combat situation. Then the inner loop uses a control policy to answer the macro command by calculating the actual input signals for the aircraft. We design the Markov decision-making process for the control policy and the Markov game between two aircraft. We present a two-stage training mechanism. For the control policy, we design an effective reward function to accurately track various macro behaviors. For the combat policy, we present a fictitious self-play mechanism to improve the combat performance by combating against the historical combat policies. Experiment results show that the control policy can achieve better tracking performance than conventional methods. The fictitious self-play mechanism can learn competitive combat policy, which can achieve high winning rates against conventional methods. Jiajun Chai, Wenzhang Chen, Yuanheng Zhu, Zong-xin Yao, Dongbin Zhao |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2022 | LILAC: Learning a Leader for Cooperative Reinforcement LearningabstractIn cooperative multi-agent reinforcement learning,role-based learning promises to reach satisfactory policy learning through the decomposition of complicated tasks using roles. Different roles are responsible for different aspects of the task. However, how this group of roles can be quickly identified is not clear. To address this problem, we propose a novel framework, LearnIng a LeAder for Cooperative reinforcement learning (LILAC), which introduces a leader to integrate information to assign roles. Leaders take a broad view of the whole task and feed the integrated information into a Gaussian mixture model to sample role embedding distribution. It enables LILAC to assign appropriate roles to different agents and improves cooperative performance. In order to evaluate the cooperation of multiple agents, a mixing network, inputted by individual local utility networks, is constructed to estimate the global action value. Two loss functions, temporal difference loss and mean divergence loss, are adopted by LILAC to learn network parameters and to encourage diversity of policies for different roles. By virtue of the leader module, LILAC outperforms the StarCraft II micromanagement benchmark in our experiments, especially on challenging tasks. Yuqian Fu, Jiajun Chai, Yuanheng Zhu, Dongbin Zhao |
CoG | 3 |
| 2022 | Learning Continuous 3-DoF Air-to-Air Close-in Combat Strategy using Proximal Policy OptimizationabstractAir-to-air close-in combat is based on many basic fighter maneuvers and can be largely modeled as an algorithmic function of inputs. This paper studies autonomous close-in combat, to learn new strategy that can adapt to different circumstances to fight against an opponent. Current methods for learning close-in combat strategy are largely limited to discrete action sets whether in the form of rules, actions or sub-polices. In contrast, we consider one-on-one air combat game with continuous action space and present a deep reinforcement learning method based on proximal policy optimization (PPO) that learns close-in combat strategy from observations in an end-to-end manner. The state space is designed to promote the learning efficiency of PPO. We also design a minimax strategy for the game. Simulation results show that the learned PPO agent is able to defeat the minimax opponent with about 97% win rate. Luntong Li, Jiajun Chai, Zhen Liu 0020, Yuanheng Zhu, Jianqiang Yi |
CoG | 5 |
| 2022 | Decentralized Event-Driven Constrained Control Using Adaptive Critic DesignsabstractWe study the decentralized event-driven control problem of nonlinear dynamical systems with mismatched interconnections and asymmetric input constraints. To begin with, by introducing a discounted cost function for each auxiliary subsystem, we transform the decentralized event-driven constrained control problem into a group of nonlinear$H_{2}$-constrained optimal control problems. Then, we develop the event-driven Hamilton–Jacobi–Bellman equations (ED-HJBEs), which arise in the nonlinear$H_{2}$-constrained optimal control problems. Meanwhile, we demonstrate that all the solutions of the ED-HJBEs together keep the overall system stable in the sense of uniform ultimate boundedness (UUB). To solve the ED-HJBEs, we build a critic-only architecture under the framework of adaptive critic designs. The architecture only employs critic neural networks and updates their weight vectors via the gradient descent method. After that, based on the Lyapunov approach, we prove that the UUB stability of all signals in the closed-loop auxiliary subsystems is assured. Finally, simulations of an illustrated nonlinear interconnected plant are provided to validate the present designs. Xiong Yang 0001, Yuanheng Zhu, Qinglai Wei |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Online Minimax Q Network Learning for Two-Player Zero-Sum Markov GamesabstractThe Nash equilibrium is an important concept in game theory. It describes the least exploitability of one player from any opponents. We combine game theory, dynamic programming, and recent deep reinforcement learning (DRL) techniques to online learn the Nash equilibrium policy for two-player zero-sum Markov games (TZMGs). The problem is first formulated as a Bellman minimax equation, and generalized policy iteration (GPI) provides a double-loop iterative way to find the equilibrium. Then, neural networks are introduced to approximate Q functions for large-scale problems. An online minimax Q network learning algorithm is proposed to train the network with observations. Experience replay, dueling network, and double Q-learning are applied to improve the learning process. The contributions are twofold: 1) DRL techniques are combined with GPI to find the TZMG Nash equilibrium for the first time and 2) the convergence of the online learning algorithm with a lookup table and experience replay is proven, whose proof is not only useful for TZMGs but also instructive for single-agent Markov decision problems. Experiments on different examples validate the effectiveness of the proposed algorithm on TZMG problems. Yuanheng Zhu, Dongbin Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Proximal Policy Optimization with Elo-based Opponent Selection and Combination with Enhanced Rolling Horizon Evolution AlgorithmabstractTwo-player zero-sum video game is a basic and important problem in game artificial intelligence. In 2020, enhanced rolling horizon evolution algorithm with policy gradient (ERHEAPI) beat heuristics, Monte-Carlo tree search and other methods to win the championship of Fighting Game Artificial Intelligence Competition (FTGAIC). However, the performance of ERHEAPI in the first round was not good. In this paper, we present an effective method noted as ERHEAPPO that combines proximal policy optimization (PPO) and enhanced rolling horizon evolution algorithm (ERHEA) with opponent model learning to further improve performance. We train the PPO agent and find that the Elo-based opponent selection can improve the sample efficiency. We compare the performance of the proposed ERHEAPPO with ERHEAPI. The experimental results demonstrate the effectiveness of ERHEAPPO. Rongqin Liang, Yuanheng Zhu, Zhentao Tang, Mu Yang |
CoG | 2 |
| 2021 | Optimal Feedback Control of Pedestrian Flow in Heterogeneous CorridorsabstractMaintaining the orderliness and efficiency of pedestrian flow through an architectural area is critical for the evacuation process. Especially, clogs and jams are easily triggered in width-changing areas. In this article, we consider pedestrian movement in heterogeneous corridors and design an optimal feedback control to regulate pedestrian flow. Flow characteristics are first studied based on microscopic social-force simulations. A Gaussian process describes the relationship between flow variables with the observation data. The macroscopic model for flow in heterogeneous corridors is developed. To avoid jams, discharges among these corridors are balanced with the narrowest corridor as the primary concern. At the equilibrium, a continuous-time nonlinear control system is formulated, and the adaptive dynamic programming learns the optimal feedback controller. Policy iteration (PI) and neural networks are combined together, and the convergence of neural-network-based PI is demonstrated by analyzing its equivalence to the Gauss–Newton method. Batch normalization is introduced to stabilize the learning process. Simulated experiments demonstrate that the control design can effectively regulate pedestrian flow for both macroscopic and microscopic models.Note to Practitioners—The development of video-processing techniques provides a powerful tool to detect human behavior in real time. In crowd events, the pedestrian movement must be regulated; otherwise, it is easy to fall into the faster-is-slower effect. It is especially important for evacuation routes with different widths. In this article, the optimal feedback control is studied to regulate pedestrian flow in heterogeneous corridors. It takes flow densities as state and produces commands that are composed of entrance influx and free-flow velocities. These commands can be executed with the support of speakers, displays, or the recently developed interactive robots. To avoid congestion, discharges of different corridors are balanced, and the system is optimally stabilized at equilibrium. Based on our work, engineers are able to design pedestrian flow control and achieve optimal evacuation in arbitrary heterogeneous corridors. Yuanheng Zhu, Dongbin Zhao, Haibo He |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2020 | An Improved Minimax-Q Algorithm Based on Generalized Policy Iteration to Solve a Chaser-Invader GameabstractIn this paper, we use reinforcement learning and zero-sum games to solve a Chaser-Invader game, which is actually a Markov game (MG). Different from the single agent Markov Decision Process (MDP), MG can realize the interaction of multiple agents, which is an extension of game theory to a MDP environment. This paper proposes an improved algorithm based on the classical Minimax-Q algorithm. First, in order to solve the problem where Minimax-Q algorithm can only be applied for discrete and simple environment, we use Deep Q-network instead of traditional Q-learning. Second, we propose a generalized policy iteration to solve the zero-sum game. This method makes the agent use linear programming method to solve the Nash equilibrium action at each moment. Finally, through comparative experiments, we prove that the improved algorithm can perform as well as Monte Carlo Tree Search in simple environments and better than Monte Carlo Tree Search in complex environments. Minsong Liu, Yuanheng Zhu, Dongbin Zhao |
IJCNN | 2 |
| 2020 | Cooperative Multi-Agent Deep Reinforcement Learning with Counterfactual RewardabstractIn partially observable fully cooperative games, agents generally tend to maximize global rewards with joint actions, so it is difficult for each agent to deduce their own contribution. To address this credit assignment problem, we propose a multi-agent reinforcement learning algorithm with counterfactual reward mechanism, which is termed as CoRe algorithm. CoRe computes the global reward difference in condition that the agent does not take its actual action but takes other actions, while other agents fix their actual actions. This approach can determine each agent's contribution for the global reward. We evaluate CoRe in a simplified Pig Chase game with a decentralised Deep Q Network (DQN) framework. The proposed method helps agents learn end-to-end collaborative behaviors. Compared with other DQN variants with global reward, CoRe significantly improves learning efficiency and achieves better results. In addition, CoRe shows excellent performances in various size game environments. Kun Shao, Yuanheng Zhu, Zhentao Tang, Dongbin Zhao |
IJCNN | 2 |
| 2020 | LMI-Based Synthesis of String-Stable Controller for Cooperative Adaptive Cruise ControlabstractController synthesis is a challenging problem in cooperative adaptive cruise control (CACC). Especially the requirement of string stability makes it even harder to choose appropriate control parameters. This paper applies a time-domain definition to string stability and converts the problem to the H∞control of a time-delay system. Based on the proposed control structure, the H∞norm and stability criteria of CACC are satisfied by a set of constraints in terms of a Lyapunov-Krasovskii functional candidate. These constraints are further reduced to linear matrix inequalities so that feasible solutions can be easily and efficiently computed. Simulations on an identified model validate the performance of our method in both frequency and time domains. Yuanheng Zhu, Haibo He, Dongbin Zhao |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2020 | Invariant Adaptive Dynamic Programming for Discrete-Time Optimal ControlabstractFor systems that can only be locally stabilized, control laws and their effective regions are both important. In this paper, invariant policy iteration is proposed to solve the optimal control of discrete-time systems. At each iteration, a given policy is evaluated in its invariantly admissible region, and a new policy and a new region are updated for the next iteration. Theoretical analysis shows the method is regionally convergent to the optimal value and the optimal policy. Combined with sum-of-squares polynomials, the method is able to achieve the near-optimal control of a class of discrete-time systems. An invariant adaptive dynamic programming algorithm is developed to extend the method to scenarios where system dynamics is not available. Online data are utilized to learn the near-optimal policy and the invariantly admissible region. Simulated experiments verify the effectiveness of our method. Yuanheng Zhu, Dongbin Zhao, Haibo He |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2019 | Optimal Pedestrian Evacuation in Building with Consecutive Differential Dynamic ProgrammingabstractFast and efficient evacuation of pedestrians from an enclosed area is a difficult but crucial issue in modern society. In this paper, the optimization of evacuation from a building is studied. A graph is adopted to describe the building layout with nodes representing areas and edges representing connections. The dynamics of the evacuation process in the graph is formulated by a nonlinear discrete-time model at a macroscopic level. To find the optimal evacuation plan, a consecutive differential dynamic programming is developed. It inherits the differential dynamic programming property that solves the value and optimal policy locally. Additionally, it consecutively executes actions for multiple steps in the trajectory, which is beneficial to reduce computational burden and lower optimization difficulty. Simulations on a four-storey building layout demonstrates our method is efficient and suitable for on-site evacuation plan making. Yuanheng Zhu, Haibo He, Dongbin Zhao, Zhongsheng Hou |
IJCNN | 1 |
| 2018 | Driving Control with Deep and Reinforcement Learning in The Open Racing Car Simulator
Yuanheng Zhu, Dongbin Zhao |
ICONIP (3) | 1 |
| 2018 | Visual Navigation with Actor-Critic Deep Reinforcement LearningabstractVisual navigation in complex environments is crucial for intelligent agents. In this paper, we propose an efficient deep reinforcement learning (DRL) method to tackle visual navigation tasks. We present the synchronous advantage actor-critic (A2C) with generalized advantage estimator (GAE) algorithm. The A2C enables agents to learn from multiple processes, which significantly reduces the training time. The GAE used to estimate the advantage function improves the policy gradient estimates. We focus on visual navigation tasks in ViZDoom, and train agents in two health gathering scenarios. The experimental results show this method successfully teaches our agents to navigate in these scenarios. The A2C with GAE agent reaches the highest score in the first task, and a competitive score in the second task. In addition, this agent has better average scores and lower variances in both tasks. Kun Shao, Dongbin Zhao, Yuanheng Zhu |
IJCNN | 3 |
| 2018 | Policy Iteration for H∞ Optimal Control of Polynomial Nonlinear Systems via Sum of Squares ProgrammingabstractSum of squares (SOS) polynomials have provided a computationally tractable way to deal with inequality constraints appearing in many control problems. It can also act as an approximator in the framework of adaptive dynamic programming. In this paper, an approximate solution to the optimal control of polynomial nonlinear systems is proposed. Under a given attenuation coefficient, the Hamilton-Jacobi-Isaacs equation is relaxed to an optimization problem with a set of inequalities. After applying the policy iteration technique and constraining inequalities to SOS, the optimization problem is divided into a sequence of feasible semidefinite programming problems. With the converged solution, the attenuation coefficient is further minimized to a lower value. After iterations, approximate solutions to the smallest -gain and the associated optimal controller are obtained. Four examples are employed to verify the effectiveness of the proposed algorithm. Yuanheng Zhu, Dongbin Zhao, Xiong Yang 0001 |
IEEE Trans. Cybern. | 1 |
| 2017 | Data-driven adaptive dynamic programming for continuous-time fully cooperative games with partially constrained inputs
Dongbin Zhao, Yuanheng Zhu |
Neurocomputing | 3 |
| 2017 | Iterative Adaptive Dynamic Programming for Solving Unknown Nonlinear Zero-Sum Game Based on Online Dataabstractcontrol is a powerful method to solve the disturbance attenuation problems that occur in some control systems. The design of such controllers relies on solving the zero-sum game (ZSG). But in practical applications, the exact dynamics is mostly unknown. Identification of dynamics also produces errors that are detrimental to the control performance. To overcome this problem, an iterative adaptive dynamic programming algorithm is proposed in this paper to solve the continuous-time, unknown nonlinear ZSG with only online data. A model-free approach to the Hamilton-Jacobi-Isaacs equation is developed based on the policy iteration method. Control and disturbance policies and value are approximated by neural networks (NNs) under the critic-actor-disturber structure. The NN weights are solved by the least-squares method. According to the theoretical analysis, our algorithm is equivalent to a Gauss-Newton method solving an optimization problem, and it converges uniformly to the optimal solution. The online data can also be used repeatedly, which is highly efficient. Simulation results demonstrate its feasibility to solve the unknown nonlinear ZSG. When compared with other algorithms, it saves a significant amount of online measurement time. Yuanheng Zhu, Dongbin Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Event-Triggered H∞ Control for Continuous-Time Nonlinear System via Concurrent LearningabstractIn this paper, the H∞optimal control problem for a class of continuous-time nonlinear systems is investigated using event-triggered method. First, the H∞optimal control problem is formulated as a two-player zero-sum (ZS) differential game. Then, an adaptive triggering condition is derived for the ZS game with an event-triggered control policy and a time-triggered disturbance policy. The event-triggered controller is updated only when the triggering condition is not satisfied. Therefore, the communication between the plant and the controller is reduced. Furthermore, a positive lower bound on the minimal intersample time is provided to avoid Zeno behavior. For implementation purpose, the event-triggered concurrent learning algorithm is proposed, where only one critic neural network (NN) is used to approximate the value function, the control policy and the disturbance policy. During the learning process, the traditional persistence of excitation condition is relaxed using the recorded data and instantaneous data together. Meanwhile, the stability of closed-loop system and the uniform ultimate boundedness (UUB) of the critic NN's parameters are proved by using Lyapunov technique. Finally, simulation results verify the feasibility to the ZS game and the corresponding H∞control problem. Dongbin Zhao, Yuanheng Zhu |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2016 | Model-free reinforcement learning for nonlinear zero-sum games with simultaneous explorationsabstractIn this paper, the continuous-time unknown nonlinear zero-sum game is investigated using a model-free online learning method. First, motivated by model-based policy iteration, an iterative equation without any knowledge of system dynamics is derived by introducing simultaneous explorations. Then, the model-free reinforcement learning based on the derived iterative equation is developed to approach the solution of the Hamilton-Jacobi-Isaacs equation. For the online implementation purpose, three neural networks are constructed to approach the value function, control and disturbance policies, respectively. Finally, a simulation example is provided to demonstrate the effectiveness of the proposed scheme. Dongbin Zhao, Yuanheng Zhu |
IJCNN | 3 |
| 2016 | Convolutional fitted Q iteration for vision-based control problemsabstractIn this paper a deep reinforcement learning (DRL) method is proposed to solve the control problem which takes raw image pixels as input states. A convolutional neural network (CNN) is used to approximate Q functions, termed as Q-CNN. A pretrained network, which is the result of a classification challenge on a vast set of natural images, initializes the parameters of Q-CNN. Such initialization assigns Q-CNN with the features of image representation, so it is more concentrated on the control tasks. The weights are tuned under the scheme of fitted Q iteration (FQI), which is an offline reinforcement learning method with the stable convergence property. To demonstrate the performance, a modified Food-Poison problem is simulated. The agent determines its movements based on its forward view. In the end the algorithm successfully learns a satisfied policy which has better performance than the results of previous researches. Dongbin Zhao, Yuanheng Zhu, Le Lv, Yaran Chen |
IJCNN | 2 |
| 2016 | Experience Replay for Optimal Control of Nonzero-Sum Game Systems With Unknown DynamicsabstractIn this paper, an approximate online equilibrium solution is developed for an N -player nonzero-sum (NZS) game systems with completely unknown dynamics. First, a model identifier based on a three-layer neural network (NN) is established to reconstruct the unknown NZS games systems. Moreover, the identifier weight vector is updated based on experience replay technique which can relax the traditional persistence of excitation condition to a simplified condition on recorded data. Then, the single-network adaptive dynamic programming (ADP) with experience replay algorithm is proposed for each player to solve the coupled nonlinear Hamilton- (HJ) equations, where only the critic NN weight vectors are required to tune for each player. The feedback Nash equilibrium is provided by the solution of the coupled HJ equations. Based on the experience replay technique, a novel critic NN weights tuning law is proposed to guarantee the stability of the closed-loop system and the convergence of the value functions. Furthermore, a Lyapunov-based stability analysis shows that the uniform ultimate boundedness of the closed-loop system is achieved. Finally, two simulation examples are given to verify the effectiveness of the proposed control scheme. Dongbin Zhao, Ding Wang 0001, Yuanheng Zhu |
IEEE Trans. Cybern. | 4 |
| 2015 | Thermal comfort control based on MEC algorithm for HVAC systemsabstractThis paper combines an efficient reinforcement learning algorithm named Multisamples in Each Cell (MEC) with a building thermal comfort control problem. It implements the efficient exploration rule and makes high use of observed samples. A grid is utilized to partition the continuous state into cells that are used to store samples. A near-upper Q function is obtained based on the samples in each cell. The value iteration technique is designed to derive the near optimal control policy. The algorithm can efficiently balance exploration and exploitation. The entire implementation process needs no model of systems. The thermal comfort criterion, predicted mean vote, is introduced to evaluate zone thermal comfort status. A two story, multi-zone small office building equipped with a variable air volume direct expansion cooling system is built in EnergyPlus to establish an EnergyPlus-MATLAB co-simulation platform. A MEC thermal comfort control simulation is implemented to validate the high performance property compared with Q-learning. Dong Li 0016, Dongbin Zhao, Yuanheng Zhu, Zhongpu Xia |
IJCNN | 3 |
| 2015 | Convergence analysis and application of fuzzy-HDP for nonlinear discrete-time HJB systems
Yuanheng Zhu, Dongbin Zhao, Derong Liu 0001 |
Neurocomputing | 1 |
| 2015 | A data-based online reinforcement learning algorithm satisfying probably approximately correct principle
Yuanheng Zhu, Dongbin Zhao |
Neural Comput. Appl. | 1 |
| 2015 | MEC - A Near-Optimal Online Reinforcement Learning Algorithm for Continuous Deterministic SystemsabstractIn this paper, the first probably approximately correct (PAC) algorithm for continuous deterministic systems without relying on any system dynamics is proposed. It combines the state aggregation technique and the efficient exploration principle, and makes high utilization of online observed samples. We use a grid to partition the continuous state space into different cells to save samples. A near-upper Q operator is defined to produce a near-upper Q function using samples in each cell. The corresponding greedy policy effectively balances between exploration and exploitation. With the rigorous analysis, we prove that there is a polynomial time bound of executing nonoptimal actions in our algorithm. After finite steps, the final policy reaches near optimal in the framework of PAC. The implementation requires no knowledge of systems and has less computation complexity. Simulation studies confirm that it is a better performance than other similar PAC algorithms. Dongbin Zhao, Yuanheng Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2014 | A data-based online reinforcement learning algorithm with high-efficient explorationabstractAn online reinforcement learning algorithm is proposed in this paper to directly utilizes online data efficiently for continuous deterministic systems without system parameters. The dependence on some specific approximation structures is crucial to limit the wide application of online reinforcement learning algorithms. We utilize the online data directly with the kd-tree technique to remove this limitation. Moreover, we design the algorithm in the Probably Approximately Correct principle. Two examples are simulated to verify its good performance. Yuanheng Zhu, Dongbin Zhao |
ADPRL | 1 |
| 2014 | Full-range adaptive cruise control based on supervised adaptive dynamic programming
Dongbin Zhao, Zhaohui Hu, Zhongpu Xia, Cesare Alippi, Yuanheng Zhu, Ding Wang 0001 |
Neurocomputing | 5 |
| 2013 | Online Model-Free RLSPI Algorithm for Nonlinear Discrete-Time Non-affine Systems
Yuanheng Zhu, Dongbin Zhao |
ICONIP (2) | 1 |
| 2012 | Neural and fuzzy dynamic programming for under-actuated systemsabstractThis paper aims to integrate the fuzzy control with adaptive dynamic programming (ADP) scheme, to provide an optimized fuzzy control performance, together with faster convergence of ADP for the help of the fuzzy prior knowledge. ADP usually consists of two neural networks, one is the Actor as the controller, the other is the Critic as the performance evaluator. A fuzzy controller applied in many fields can be used instead as the Actor to speed up the learning convergence, because of its simplicity and prior information on fuzzy membership and rules. The parameters of the fuzzy rules are learned by ADP scheme to approach optimal control performance. The feature of fuzzy controller makes the system steady and robust to system states and uncertainties. Simulations on under-actuated systems, a cart-pole plant and a pendubot plant, are implemented. It is verified that the proposed scheme is capable of balancing under-actuated systems and has a wider control zone. Dongbin Zhao, Yuanheng Zhu, Haibo He |
IJCNN | 2 |