Zhiqiang Pu

dblp:154/6080 · DBLP profile ↗
← Back
62ranked-venue papers
4as first author
50since 2021 · last 2026
0000-0002-4841-4048ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 54 · 2 first-author · 44 since 2021Systems, architecture and hardware · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Unreal-MAP: Unreal-Engine-Based General Platform for Multi-agent Reinforcement Learning
abstract
In this paper, we propose Unreal Multi-Agent Playground (Unreal-MAP), an MARL general platform based on the Unreal-Engine (UE). Unreal-MAP allows users to freely create multi-agent tasks using the vast visual and physical resources available in the UE community, and deploy state-of-the-art (SOTA) MARL algorithms within them. Unreal-MAP is user-friendly in terms of deployment, modification, and visualization, and all its components are open-source. We also develop an experimental framework compatible with algorithms ranging from rule-based to learning-based provided by third-party frameworks. Lastly, we deploy several SOTA algorithms in example tasks developed via Unreal-MAP, and conduct corresponding experimental analyses including a sim2real demo. We believe Unreal-MAP can play an important role in the MARL field by closely integrating existing algorithms with user-customized tasks, thus advancing the field of MARL.
Qingxu Fu, Zhiqiang Pu, Tenghai Qiu
AAAI3
2026 SSEditor: Controllable mask-to-scene generation with diffusion model
Jiahao Pang, Zhiqiang Pu, Yanyan Liang 0001
Knowl. Based Syst.3
2026 Flexible Modal Mixture-of-Experts With Inter-Modal Knowledge Distillation for Face Anti-Spoofing
Hui Ma 0018, Ajian Liu 0001, Ning Li 0035, Boyun Wang, Hang Zou 0002, Yuan Zhang 0023, Jing Huang 0017, Zhiqiang Pu, Jun Wan 0001, Zhanchuan Cai, Zhen Lei 0001, Yanyan Liang 0001
IEEE Trans. Inf. Forensics Secur.9
2025 Vision-Based Generic Potential Function for Policy Alignment in Multi-Agent Reinforcement Learning
abstract
Guiding the policy of multi-agent reinforcement learning to align with human common sense is a difficult problem, largely due to the complexity of modeling common sense as a reward, especially in complex and long-horizon multi-agent tasks. Recent works have shown the effectiveness of reward shaping, such as potential-based rewards, to enhance policy alignment. The existing works, however, primarily rely on experts to design rule-based rewards, which are often labor-intensive and lack a high-level semantic understanding of common sense. To solve this problem, we propose a hierarchical vision-based reward shaping method. At the bottom layer, a visual-language model (VLM) serves as a generic potential function, guiding the policy to align with human common sense through its intrinsic semantic understanding. To help the policy adapts to uncertainty and changes in long-horizon tasks, the top layer features an adaptive skill selection module based on a visual large language model (vLLM). The module uses instructions, video replays, and training records to dynamically select suitable potential function from a pre-designed pool. Besides, our method is theoretically proven to preserve the optimal policy. Extensive experiments conducted in the Google Research Football environment demonstrate that our method not only achieves a higher win rate but also effectively aligns the policy with human common sense.
Shijie Wang 0006, Zhiqiang Pu, Siyao Zhao, Xiaolin Ai
AAAI3
2025 OMoE: Diversifying Mixture of Low-Rank Adaptation by Orthogonal Finetuning
abstract
Building mixture-of-experts (MoE) architecture for low-rank adaptation (LoRA) is emerging as a potential direction in parameter-efficient fine-tuning (PEFT) for its modular design and remarkable performance. However, simply stacking the number of experts cannot guarantee significant improvement. In this work, we first conduct qualitative analysis to indicate that experts collapse to similar representations in vanilla MoE, limiting the capacity of modular design and computational efficiency. Ulteriorly, our analysis reveals that the performance of previous MoE variants may be limited by a lack of diversity among experts. Motivated by these findings, we propose Orthogonal Mixture-of-Experts (OMoE), a resource-efficient MoE variant that trains experts in an orthogonal manner to promote diversity. In OMoE, a Gram-Schmidt process is leveraged to enforce that the experts’ representations lie within the Stiefel manifold. By applying orthogonal constraints directly to the architecture, OMoE keeps the learning objective unchanged, without compromising optimality. Our method is simple and alleviates memory bottlenecks, as it incurs minimal experts compared to vanilla MoE models. Experiments on diverse benchmarks demonstrate that OMoE can consistently achieve stable and efficient performance improvement when compared with the state-of-the-art methods while significantly reducing the number of required experts.
Jinyuan Feng, Zhiqiang Pu, Xiaolin Ai, Huimu Wang
ECAI2
2025 Stochastic Trajectory Prediction Under Unstructured Constraints
abstract
Trajectory prediction facilitates effective planning and decision-making, while constrained trajectory prediction integrates regulation into prediction. Recent advances in constrained trajectory prediction focus on structured constraints by constructing optimization objectives. However, handling unstructured constraints is challenging due to the lack of differentiable formal definitions. To address this, we propose a novel method for constrained trajectory prediction using a conditional generative paradigm, named Controllable Trajectory Diffusion (CTD). The key idea is that any trajectory corresponds to a degree of conformity to a constraint. By quantifying this degree and treating it as a condition, a model can implicitly learn to predict trajectories under unstructured constraints. CTD employs a pre-trained scoring model to predict the degree of conformity (i.e., a score), and uses this score as a condition for a conditional diffusion model to generate trajectories. Experimental results demonstrate that CTD achieves high accuracy on the ETH/UCY and SDD benchmarks. Qualitative analysis confirms that CTD ensures adherence to unstructured constraints and can predict trajectories that satisfy combinatorial constraints.
Zhiqiang Pu, Shijie Wang 0006, Boyin Liu, Huimu Wang, Yanyan Liang 0001, Jianqiang Yi
ICRA2
2025 Diversity-Driven Offline-to-Online Multi-Player Policy Learning for Football Matches
abstract
Due to the complexity of football matches, high-quality player policies must exhibit diverse behaviors for effective collaboration. However, in football tasks where online interactions are time-consuming and costly, it is difficult to balance training efficiency and the emergence of diversity when learning policies from scratch. Therefore, this paper proposes a novel diversity-driven offline-to-online (DOTO) multi-player policy learning method for football matches. Specifically, we design a transformer-actor-critic module that is universally applicable to both offline and online stages, enabling seamless adaptation. This module utilizes expert data for offline pre-training, which serves as the initialization for online fine-tuning. Subsequently, we introduce an across offline and online adaptive intention clustering guidance module, which learns intention representations of agents and performs intention-based clustering to construct adaptive loss functions to guide policy diversity. Besides, to mitigate the bootstrapping errors of distribution shift from offline to online, we implement an online partial random initialization mechanism, balancing the inherent conservativeness of offline learning with the flexible exploration of online learning. Extensive experiments in Google Research Football environment show that DOTO not only increases win rate to 50% which surpasses all state-of-the-art methods in the hardest 11 vs. 11 task, but also generates richer policy diversity.
Shijie Wang 0006, Zhiqiang Pu, Huimu Wang, Jianqiang Yi
IJCNN2
2025 CLGA: A Collaborative LLM Framework for Dynamic Goal Assignment in Multi-Robot Systems
abstract
Goal assignment is a critical challenge in multi-robot systems. The emergence of large language models (LLMs) has enabled the use of natural language commands for tackling goal assignment problems. However, applying LLMs directly to these tasks presents two limitations: 1) limited accuracy and 2) excessive decision delays due to their autoregressive nature, hindering adaptability to unexpected changes. To address these issues, inspired by dual-process theory, we propose a framework called Collaborative LLMs for dynamic Goal Assignment (CLGA). Specifically, we leverage LLMs for pre-planning tasks and invoke an external solver to generate an initial goal assignment solution, ensuring solution accuracy. During execution, small-scale models enable real-time adjustments to respond to dynamic environmental changes. This approach integrates the strengths of slow, precise pre-planning and fast, adaptive online adjustments, allowing agents to efficiently handle real-world challenges. Additionally, we introduce a benchmark dataset for NLP-based goal assignment to advance research in this domain. Simulation and real-world experiments demonstrate that CLGA significantly enhances task execution efficiency and flexibility in multi-robot systems. The prompt, experimental videos, and datasets associated with this work are available at https://sites.google.com/view/project-clga/.
Xin Yu 0009, Yandong Wang 0002, Rongye Shi, Gangzheng Ai, Zhiqiang Pu, Wenjun Wu 0001
IROS7
2025 Distributed Nash equilibrium for pursuit-evasion game with one evader and multiple pursuers
Libing Wang, Zhiqiang Pu
Sci. China Inf. Sci.4
2025 A Policy Resonance Approach to Solve the Problem of Responsibility Diffusion in Multiagent Reinforcement Learning
abstract
State-of-the-art (SOTA) multiagent reinforcement algorithms distinguish themselves in many ways from their single-agent equivalences. However, most of them still totally inherit the single-agent exploration-exploitation strategy. Naively inheriting this strategy from single-agent algorithms causes potential collaboration failures, in which the agents blindly follow mainstream behaviors and reject taking minority responsibility. We name this problem the responsibility diffusion (RD) as it shares similarities with the same-name social psychology effect. In this work, we start by theoretically analyzing the cause of this RD problem, which can be traced back to the exploration-exploitation dilemma of multiagent systems (especially large-scale multiagent systems). We address this RD problem by proposing a policy resonance (PR) approach which modifies the collaborative exploration strategy of agents by refactoring the joint agent policy while keeping individual policies approximately invariant. Next, we show that SOTA algorithms can equip this approach to promote the collaborative performance of agents in complex cooperative tasks. Experiments are performed in multiple test benchmark tasks to illustrate the effectiveness of this approach.
Qingxu Fu, Tenghai Qiu, Jianqiang Yi, Zhiqiang Pu, Xiaolin Ai, Wanmai Yuan
IEEE Trans. Neural Networks Learn. Syst.4
2025 Cognition-Oriented Multiagent Reinforcement Learning
abstract
Inspired by psychological insights into individual behavior, we propose a novel cognition-oriented multiagent reinforcement learning (CORL) framework. CORL equips agents with two distinct types of cognition-situational and self-cognition-derived from local observations. To enhance the informativeness and precision of these cognition types, we introduce two information-theoretical regularizers: one to align situational cognition with the global state and the other to align self-cognition with each agent's identity for improved role differentiation and team coordination. In addition, the centralized training and decentralized execution framework is adopted to train the policy network. Our simulations demonstrate that CORL effectively harnesses local observations for enriched cooperation, leading to pronounced performance improvements, particularly in challenging tasks.
Tenghai Qiu, Shiguang Wu 0001, Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi, Yuqian Zhao 0001, Biao Luo 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Neural Formation A*: A Knowledge-Data Hybrid-Driven Path Planning Algorithm for Multi-agent Formation Cooperation
Qi'ang Cai, Xiaolin Ai, Zhiqiang Pu, Jianqiang Yi, Feng Lv
ICANN (4)4
2024 Extraction and Transfer of General Deep Feature in Reinforcement Learning
abstract
Knowledge transfer from teacher agent to student agent in reinforcement learning addresses the issue of sample inefficiency resulting from random exploration. However, existing approaches usually assume that the policy or state-action pairs (demonstrations) of teacher is accessible, which may not be met in reality due to privacy and other limitations. This paper proposes a novel knowledge transfer method based solely on the state sequences of teacher. Specifically, We encode raw state into general deep feature through learning state evaluation and domain discrimination tasks based on teachers’ experience, and share the encoder among different student agents in various scenarios. We validate our method in Starcraft II and Google Research Football environment. The experimental results show that our method significantly accelerates the learning process of agents and improves their convergence performance, showcasing the "depth" of the feature. Additionally, the feature successfully generalize to teachers’ unseen scenarios and unfamiliar roles, demonstrating the "generality" of the feature.
Min Chen 0038, Yi Pan 0009, Zhiqiang Pu, Jianqiang Yi, Shijie Wang 0006, Boyin Liu
IJCNN3
2024 Heterogeneous Observation Aggregation Network for Multi-agent Reinforcement Learning
abstract
Learning effective policies is challenging for a multi-agent system in partially observable environments, where agents need to extract relevant features from local observations. Most approaches in multi-agent reinforcement learning (MARL) are limited to feature extraction for homogenous agents. They struggle to deal with local observations in heterogeneous multi-agent scenarios, where agents have different observation spaces and are necessitated to process semantically varied information. To address this issue, we analyze the observational heterogeneity of multi-agent systems, and propose a heterogeneous-graph-based approach for feature extraction in MARL. We model agent observations as heterogeneous graphs, and design a heterogeneous observation aggregation network (HOA-Net) for processing these graph-based observations. HOA-Net is specifically designed to address various forms of observational heterogeneity. It employs class-specific weighting networks and computes across-class attentions for observed entities, effectively reducing the number of learnable parameters. The proposed method is evaluated on SMAC and an Unreal-Engine-based heterogeneous multi-agent testbed. Experimental results demonstrate that our method significantly outperforms other baselines in effectively aggregating an agent’s observation, and finally enhancing the performance of heterogeneous multi-agent systems.
Xiaolin Ai, Zhiqiang Pu, Tenghai Qiu, Jianqiang Yi
IJCNN3
2024 Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement Learning
abstract
Reinforcement learning (RL) has emerged as a pivotal technique for fine-tuning large language models (LLMs) on specific tasks. However, prevailing RL fine-tuning methods predominantly rely on PPO and its variants. Though these algorithms are effective in general RL settings, they often exhibit suboptimal performance and vulnerability to distribution collapse when applied to the fine-tuning of LLMs. In this paper, we propose CORY, extending the RL fine-tuning of LLMs to a sequential cooperative multi-agent reinforcement learning framework, to leverage the inherent coevolution and emergent capabilities of multi-agent systems. In CORY, the LLM to be fine-tuned is initially duplicated into two autonomous agents: a pioneer and an observer. The pioneer generates responses based on queries, while the observer generates responses using both the queries and the pioneer’s responses. The two agents are trained together. During training, the agents exchange roles periodically, fostering cooperation and coevolution between them. Experiments evaluate CORY's performance by fine-tuning GPT-2 and Llama-2 under subjective and objective reward functions on the IMDB Review and GSM8K datasets, respectively. Results show that CORY outperforms PPO in terms of policy optimality, resistance to distribution collapse, and training robustness, thereby underscoring its potential as a superior methodology for refining LLMs in real-world applications.
Zhiqiang Pu, Boyin Liu, Xiaolin Ai, Yanyan Liang 0001, Min Chen 0038
NeurIPS3
2024 An Affine-Based Maneuver Control Method for Multi-Agent Cooperative Transportation System Over Switching Formations
abstract
This paper addresses the maneuver control prob-lem for the multi-agent cooperative transportation systems (MACTSs) with double-integrator dynamics over switching formations. The switching formations consist of a pair of operations: one is the switching directed graphs and the other is the corresponding formation configurations. In real cooperative transportation scenarios, varying interaction re-lationships necessitate distinct system configurations. But it is challenging for researchers to design control law with the evolutions of not only the communication topology but also the system configuration. Drawing inspiration from advancements on switching topologies, we utilize the characteristics of the stress matrix to build up a new LMI inequality to design a novel class of distributed affine-based controller. The global convergence will be achieved as long as the feedback gain matrix and the switching signal satisfy three specific conditions. And we give the corresponding algorithm to calculate the required control parameters. A Lyapunov function is constructed to demonstrate the system's stability and simulation example is provided in detail at the end of this paper to validate our method's efficacy.
Xiaolin Ai, Zhiqiang Pu, Feng Lv
SMC3
2024 Fuzzy Feedback Multiagent Reinforcement Learning for Adversarial Dynamic Multiteam Competitions
abstract
A large proportion of recent studies on cooperative Multi-Agent Reinforcement Learning (MARL) focus on the policylearning process in scenarios with stationary opponents (or without opponents). This paper, instead, investigates a different challenge of achieving team superiority in dynamic competitions among competitors that evolve dynamically with MARL. We aim to enhance the competitiveness of such MARL learners by enabling them to adjust their own learning settings dynamically, so as to take quick counter-measures against the policy shift of competitor learners, or to learn faster to suppress the opponents. We propose a Competitive Auto-Multiagent Learner with Fuzzy Feedback (CALF) with two essential highlights: (1) CALF establishes feedback controllers to achieve real-time adjustments based on fuzzy logic, using human-readable fuzzy rules to provide significant explainability and flexibility; (2) CALF integrates Bayesian Optimization to search and optimize the feedback fuzzy logic rules automatically. CALF can be used to apply real-time adjustments for MARL hyperparameters and intrinsic rewards. We also give solid empirical results to show that CALF significantly promotes team competitiveness in adversarial competitions, spanning from small-scale tasks involving 2 teams to large-scale tasks involving 3 teams and hundreds of agents. Furthermore, CALF exhibits superior competitiveness when engaging in competition with established competitors like Qmix, Qtran, and Qplex in dynamic competitive environments. Moreover, the experiments also demonstrate that the integration of the fuzzy logic with Bayesian Optimization offers considerable transferability and explainability, enabling a CALF-implemented learner optimized from one scenario to be transferred to other distinct scenarios.
Qingxu Fu, Zhiqiang Pu, Yi Pan 0009, Tenghai Qiu, Jianqiang Yi
IEEE Trans. Fuzzy Syst.2
2024 Multiexperience-Assisted Efficient Multiagent Reinforcement Learning
abstract
Recently, multiagent reinforcement learning (MARL) has shown great potential for learning cooperative policies in multiagent systems (MASs). However, a noticeable drawback of current MARL is the low sample efficiency, which causes a huge amount of interactions with environment. Such amount of interactions greatly hinders the real-world application of MARL. Fortunately, effectively incorporating experience knowledge can assist MARL to quickly find effective solutions, which can significantly alleviate the drawback. In this article, a novel multiexperience-assisted reinforcement learning (MEARL) method is proposed to improve the learning efficiency of MASs. Specifically, monotonicity-constrained reward shaping is innovatively designed using expert experience to provide additional individual rewards to guide multiagent learning efficiently, with the invariance guarantee of the team optimization objective. Furthermore, a reward distribution estimator is specially developed to model an implicated reward distribution of environment by using transition experience from environment, containing collected samples (state-action pair, reward, and next state). This estimator can predict the expectation reward of each agent for the taken action to accurately estimate the state value function and accelerate its convergence. Besides, the performance of MEARL is evaluated on two multiagent environment platforms: our designed unmanned aerial vehicle combat (UAV-C) and StarCraft II Micromanagement (SCII-M). Simulation results demonstrate that the proposed MEARL can greatly improve the learning efficiency and performance of MASs and is superior to the state-of-the-art methods in multiagent tasks.
Zhen Liu 0020, Jianqiang Yi, Shiguang Wu 0001, Zhiqiang Pu
IEEE Trans. Neural Networks Learn. Syst.5
2023 Improving Generalization of Multi-agent Reinforcement Learning Through Domain-Invariant Feature Extraction
Yifan Xu 0018, Zhiqiang Pu, Feimo Li, Xinghua Chai
ICANN (6)2
2023 Lazy Agents: A New Perspective on Solving Sparse Reward Problem in Multi-agent Reinforcement Learning
abstract
Sparse reward remains a valuable and challenging problem in multi-agent reinforcement learning (MARL). This paper addresses this issue from a new perspective, i.e., lazy agents. We empirically illustrate how lazy agents damage learning from both exploration and exploitation. Then, we propose a novel MARL framework called Lazy Agents Avoidance through Influencing External States (LAIES). Firstly, we examine the causes and types of lazy agents in MARL using a causal graph of the interaction between agents and their environment. Then, we mathematically define the concept of fully lazy agents and teams by calculating the causal effect of their actions on external states using the do-calculus process. Based on definitions, we provide two intrinsic rewards to motivate agents, i.e., individual diligence intrinsic motivation (IDI) and collaborative diligence intrinsic motivation (CDI). IDI and CDI employ counterfactual reasoning based on the external states transition model (ESTM) we developed. Empirical results demonstrate that our proposed method achieves state-of-the-art performance on various tasks, including the sparse-reward version of StarCraft multi-agent challenge (SMAC) and Google Research Football (GRF). Our code is open-source and available at https://github.com/liuboyin/LAIES.
Boyin Liu, Zhiqiang Pu, Yi Pan 0009, Jianqiang Yi, Yanyan Liang 0001
ICML2
2023 All for Goals: a Stylized Automated Analysis Framework in Football Matches
abstract
Automated analysis in football matches is meaningful for player and team evaluation. However, most related works ignore match style and team strength. In this paper, a novel stylized automated analysis framework termed All for Goals (AFG) is proposed for football matches, which considers match style and team strength to better quantify the relationship of all match states and player actions respectively with potential goals. AFG is composed of an automatic labeling module, a potential goal prediction module, and a state and player evaluation module. Specifically, in the automatic labeling module, relevant samples are given the same label to avoid manual labeling. In the potential goal prediction module, we introduce the Pretrain-Finetune paradigm. Based on labeled data, an average model learning to identify scoring difficulty is obtained in the first pre-training procedure, and tuned models learning specific styles are obtained in the second fine-tuning procedure. In the state and player evaluation module, the evaluation mechanisms of state, on-ball action, and off-ball running based on potential goal prediction result are designed for match review and tactics mining. Finally, we validate the rationality and validity of AFG on multiple tasks. On the goal prediction task, the models show high recall rates and remarkable difference in style. On real-time situation analysis, credit assignment for football events, and off-ball running analysis tasks, the evaluation mechanisms give the results consistent with football domain knowledge.
Min Chen 0038, Zhiqiang Pu, Yi Pan 0009, Jianqiang Yi, Yixiong Cui, Lida Du
IJCNN2
2023 Learning Superior Cooperative Policy in Adversarial Multi-Team Reinforcement Learning
abstract
Multi-agent Reinforcement Learning (MARL) has become a powerful tool for addressing multi-agent challenges. Existing studies have explored numerous models to use MARL to solve single-team cooperation (competition) problems and adversarial problems with opponents controlled by static knowledge-based policies. However, most studies in the literature often ignore adversarial multi-team problems involving dynamically evolving opponents. We investigate adversarial multi-team problems where all participating teams use MARL learners to learn policies against each other. Two objectives are achieved in this study. Firstly, we design an adversarial team-versus-team learning framework to generate cooperative multi-agent policies to compete against opponents without preprogrammed opponent partners or any supervision. Secondly, we explore the key factors to achieve win-rate superiority during dynamic competitions. Then we put forward a novel FeedBack MARL (FBMARL) algorithm that takes advantage of feedback loops to adjust optimizer hyper-parameters based on real-time game statistics. Finally, the effectiveness of our FBMARL model is tested in a benchmark environment named Multi-Team Decentralized Collective Assault (MT-DCA). The results demonstrate that our feedback MARL model can achieve superior performance over baseline competitor MARL learners in 2-team and 3-team dynamic competitions.
Qingxu Fu, Tenghai Qiu, Zhiqiang Pu, Jianqiang Yi, Xiaolin Ai, Wanmai Yuan
IJCNN3
2023 Causal Mean Field Multi-Agent Reinforcement Learning
abstract
Scalability remains a challenge in multi-agent reinforcement learning and is currently under active research. A framework named mean-field reinforcement learning (MFRL) could alleviate the scalability problem by employing the Mean Field Theory to turn a many-agent problem into a two-agent problem. However, this framework lacks the ability to identify essential interactions under nonstationary environments. Causality contains relatively invariant mechanisms behind interactions, though environments are nonstationary. Therefore, we propose an algorithm called causal mean-field Q-learning (CMFQ) to address the scalability problem. CMFQ is ever more robust toward the change of the number of agents though inheriting the compressed representation of MFRL's action-state space. Firstly, we model the causality behind the decision-making process of MFRL into a structural causal model (SCM). Then the essential degree of each interaction is quantified via intervening on the SCM. Furthermore, we design the causality-aware compact representation for behavioral information of agents as the weighted sum of all behavioral information according to their causal effects. We test CMFQ in a mixed cooperative-competitive game and a cooperative game. The result shows that our method has excellent scalability performance in both training in environments containing a large number of agents and testing in environments containing much more agents.
Zhiqiang Pu, Yi Pan 0009, Boyin Liu, Junlong Gao
IJCNN2
2023 Heterogeneous-graph Attention Reinforcement Learning for Football Matches
abstract
Football player's decision-making problem is quite challenging because of the essential feature of football matches: many players with different and complex cooperative or competitive relationships. To better leverage these relationships, this paper proposes a player-policy learning method with heterogeneous-graph attention reinforcement learning (PPL-HGARL) to enable an active player closest to the ball to learn effective policies for playing football matches. Specifically, a multi-head feature representation module is designed to reconstruct raw observations using prior expert knowledge. Furthermore, a heterogeneous-graph player-relation attention network is designed by use of graph among different roles, in order to model cooperative and competitive relations among players. The network learns proper state representation for the active player, making the player pay attention to other important players. Besides, an actor-critic algorithm is adopted to train the policies efficiently. Competing with rule-based opponents of different difficulty levels in the Google Research Football environment, the active player achieves excellent results, which validates the effectiveness and superiority of the proposed method.
Shijie Wang 0006, Yi Pan 0009, Zhiqiang Pu, Jianqiang Yi, Yanyan Liang 0001
IJCNN3
2023 Learning Cooperative Policies with Graph Networks in Distributed Swarm Systems
abstract
Deriving efficient cooperative policies in uncertain dynamic environments poses huge challenges for a distributed swarm system due to the limited capability of the agents and the complex dynamics of the environment. In this paper, a novel distributed method based on deep reinforcement learning using observation-level and communication-level graph networks is proposed to learn cooperative policies for the distributed swarm system. Specifically, a relational directed graph attention neural network is designed to model observation-level graphs composed of heterogeneous relational graphs among each agent and each type of entities (e.g., obstacles, other teammates, opponents), for extracting different relational representations. Moreover, a relevant directed graph attention network is presented to cut off the ineffective communication among irrelevant agents, and model a relevant communication topology between each agent and relevant homogeneous neighbor agents as an communication-level graph, for promoting efficient inter-agent interactions. Furthermore, a distributed actor-critic algorithm with full parameter sharing is implemented to learn cooperative swarm policies by using distributed critics, which avoids the curse of dimensionality under a centralized critic. Various simulation results validate the effectiveness and generalization of the proposed method, and demonstrate that the proposed method outperforms existing state-of-the-art methods on coverage and pursuit tasks.
Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi, Xiaolin Ai, Wanmai Yuan
IJCNN3
2023 Deconfounded Opponent Intention Inference for Football Multi-Player Policy Learning
abstract
Due to the high complexity of a football match, the opponents' strategies are variable and unknown. Thus predicting the opponents' future intentions accurately based on current situation is crucial for football players' decision-making. To better anticipate the opponents and learn more effective strategies, a deconfounded opponent intention inference (DOII) method for football multi-player policy learning is proposed in this paper. Specifically, opponents' intentions are inferred by an opponent intention supervising module. Furthermore, for some confounders which affect the causal relationship among the players and the opponents, a decon-founded trajectory graph module is designed to mitigate the influence of these confounders and increase the accuracy of the inferences about opponents' intentions. Besides, an opponent-based incentive module is designed to improve the players' sensitivity to the opponents' intentions and further to train reasonable players' strategies. Representative results indicate that DOII can effectively improve the performance of players' strategies in the Google Research Football environment, which validates the superiority of the proposed method.
Shijie Wang 0006, Yi Pan 0009, Zhiqiang Pu, Boyin Liu, Jianqiang Yi
IROS3
2023 Learning to Play Football From Sports Domain Perspective: A Knowledge-Embedded Deep Reinforcement Learning Framework
abstract
Applying deep reinforcement learning to football games has recently received extensive attention. However, this remains challenging due to the excessively high complexity of the football environment, such as high-dynamical game states, sparse rewards, and multiple roles with different capabilities. Existing works aim to address these problems without considering abundant domain knowledge of football. In this article, a football knowledge-embedded learning framework is proposed. Specifically, the pitch control concept is innovatively introduced to design a knowledge-embedded state representation. As a result, a novel pitch control model is designed that quantitatively provides space influence values of a single player, the whole team, and the ball. Different from existing models, this model additionally considers each player's various capabilities, including flexibility, explosive force, and stamina. Furthermore, the deformable convolution network is adopted for state representation extracting, which is used to process the geometric transformation of the players' positions and spatial influence values generated by the pitch control model. Then, based on this comprehensive state representation, a proximal policy optimization-based reinforcement learning scheme is adopted to generate the final policy. Finally, extensive simulations, including learning against a fixed opponent and learning from self-play, clearly show the effectiveness and adaptability of our proposed framework.
Boyin Liu, Zhiqiang Pu, Huimu Wang, Jianqiang Yi, Jiachen Mi
IEEE Trans. Games2
2023 Cognition-Driven Multiagent Policy Learning Framework for Promoting Cooperation
abstract
Many attempts have been made to promote cooperation for multiagent systems. However, several issues that draw less attentions but may dramatically degrade the cooperation performance still exist, such as redundant information interactions among neighbors, and difficulties in understanding complex and dynamic environments from high-level cognition. To address these limitations, a cognition-driven multiagent policy (CDMAP) learning framework is proposed in this article. It includes a cognition difference network (CDN), a coupling cognition network (CCN), and a policy optimization network (PON). CDN is designed based on a variational autoencoder, where a concept of cognition difference is defined to prune redundant interactions among agents for more efficient communication. Based on the pruned topology, CCN captures the hidden representations of the surrounding environment. Several coupling graph attention layers are incorporated in CCN, each layer with different but coupling adjacent matrices, yielding a comprehensive state understanding from multiple representation spaces. Based on the captured hidden states, PON generates the final policies, where QMIX is adopted as a value factorization method to alleviate the credit-assignment problem. At last, CDMAP is evaluated through two representative multiagent games including Google Research Football andStarCraft II. The results demonstrate its superior effectiveness compared with existing methods.
Zhiqiang Pu, Huimu Wang, Boyin Liu, Jianqiang Yi
IEEE Trans. Games1
2023 Peer Incentive Reinforcement Learning for Cooperative Multiagent Games
abstract
Social learning, especially social incentives, is extremely important for humans to achieve a high level of coordination. Inspired by this, we introduce this concept into cooperative multiagent reinforcement learning (MARL), to implicitly address the credit assignment problem and promote the interagent direct interactions for cooperations among agents in cooperative multiagent games. In this article, we propose a novel intrinsic reward method with peer incentives (IRPI) based on actor–critic policy gradient. This method can enable agents to incentivize each other for their cooperations through using causal influence among them. Specifically, a novel intrinsic reward mechanism is innovatively designed to empower each agent the ability to give positive or negative rewards to other peer agents' actions through considering the causal influence of the other agents on it. The mechanism is realized by a feedforward neural network through utilizing causal influence between the agents. The causal influence of one agent on another is inferred via counterfactual reasoning using the joint action-value function in MARL. The quality of the influence is assessed via counterfactual reasoning using the individual value function in MARL. Simulations are carried out on two popular multiagent game testbeds: Starcraft II Micromanagement and Multiagent Particle Environments. Simulation results demonstrate that the proposed IRPI can enhance cooperations among the agents to achieve better performance compared with a number of state-of-the-art MARL methods in a variety of cooperative multiagent games.
Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi
IEEE Trans. Games3
2023 Deep-Reinforcement-Learning-Based Multitarget Coverage With Connectivity Guaranteed
abstract
Deriving a distributed, time-efficient, and connectivity-guaranteed coverage policy in multitarget environment poses huge challenges for a multirobot team with limited coverage and limited communication. In particular, the robot team needs to cover multiple targets while preserving connectivity. In this article, a novel deep-reinforcement-learning-based approach is proposed to take both multitarget coverage and connectivity preservation into account simultaneously, which consists of four parts: a hierarchical observation attention representation, an interaction attention representation, a two-stage policy learning, and a connectivity-guaranteed policy filtering. The hierarchical observation attention representation is designed for each robot to extract the latent features of the relations from its neighboring robots and the targets. To promote the cooperation behavior among the robots, the interaction attention representation is designed for each robot to aggregate information from its neighboring robots. Moreover, to speed up the training process and improve the performance of the learned policy, the two-stage policy learning is presented using two reward functions based on algebraic connectivity and coverage rate. Furthermore, the learned policy is filtered to strictly guarantee the connectivity based on a model of connectivity maintenance. Finally, the effectiveness of the proposed method is validated by numerous simulations. Besides, our method is further deployed to an experimental platform based on quadrotor unmanned aerial vehicles and omnidirectional vehicles. The experiments illustrate the practicability of the proposed method.
Shiguang Wu 0001, Zhiqiang Pu, Tenghai Qiu, Jianqiang Yi
IEEE Trans. Ind. Informatics2
2023 Attention Enhanced Reinforcement Learning for Multi agent Cooperation
abstract
In this article, a novel method, called attention enhanced reinforcement learning (AERL), is proposed to address issues including complex interaction, limited communication range, and time-varying communication topology for multi agent cooperation. AERL includes a communication enhanced network (CEN), a graph spatiotemporal long short-term memory network (GST-LSTM), and parameters sharing multi-pseudo critic proximal policy optimization (PS-MPC-PPO). Specifically, CEN based on graph attention mechanism is designed to enlarge the agents' communication range and to deal with complex interaction among the agents. GST-LSTM, which replaces the standard fully connected (FC) operator in LSTM with graph attention operator, is designed to capture the temporal dependence while maintaining the spatial structure learned by CEN. PS-MPC-PPO, which extends proximal policy optimization (PPO) in multi agent systems with parameters' sharing to scale to environments with a large number of agents in training, is designed with multi-pseudo critics to mitigate the bias problem in training and accelerate the convergence process. Simulation results for three groups of representative scenarios including formation control, group containment, and predator-prey games demonstrate the effectiveness and robustness of AERL.
Zhiqiang Pu, Huimu Wang, Zhen Liu 0020, Jianqiang Yi, Shiguang Wu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 A Deep Reinforcement Learning Approach Combined With Model-Based Paradigms for Multiagent Formation Control With Collision Avoidance
abstract
Generating collision-free formation control strategy for multiagent systems faces huge challenges in collaborative navigation tasks, especially in a highly dynamic and uncertain environment. Two typical methodologies for solving this problem are the conventional model-based paradigm and the data-driven paradigm, particularly the widely used deep reinforcement learning (DRL) method. However, both the model-based and data-driven paradigms encounter inherent drawbacks. In this paper, we present two novel general schemes that combine these two paradigms together in an online mode. Specifically, the two paradigms are combined in a parallel and a serial structure in these two schemes, respectively. In the parallel scheme, the outputs of the model-based and DRL-based controllers are lumped together. In the serial scheme, the output of the model-based controller is fed as an input of the DRL-based controller. The interpretation of the two combined schemes is suggested from a control-oriented perspective, where the parallel DRL controller is viewed as a complementary uncertainty compensator and the serial DRL controller is taken as an inverse dynamics estimator. Finally, comprehensive simulations are conducted to demonstrate the superiority of the proposed schemes, and the effectiveness is further verified by deploying our schemes to a physical experiment platform based on a set of three-wheeled omnidirectional robots.
Zhiqiang Pu, Xiaolin Ai, Tenghai Qiu, Jianqiang Yi
IEEE Trans. Syst. Man Cybern. Syst.1
2022 Concentration Network for Reinforcement Learning of Large-Scale Multi-Agent Systems
abstract
When dealing with a series of imminent issues, humans can naturally concentrate on a subset of these concerning issues by prioritizing them according to their contributions to motivational indices, e.g., the probability of winning a game. This idea of concentration offers insights into reinforcement learning of sophisticated Large-scale Multi-Agent Systems (LMAS) participated by hundreds of agents. In such an LMAS, each agent receives a long series of entity observations at each step, which can overwhelm existing aggregation networks such as graph attention networks and cause inefficiency. In this paper, we propose a concentration network called ConcNet. First, ConcNet scores the observed entities considering several motivational indices, e.g., expected survival time and state value of the agents, and then ranks, prunes, and aggregates the encodings of observed entities to extract features. Second, distinct from the well-known attention mechanism, ConcNet has a unique motivational subnetwork to explicitly consider the motivational indices when scoring the observed entities. Furthermore, we present a concentration policy gradient architecture that can learn effective policies in LMAS from scratch. Extensive experiments demonstrate that the presented architecture has excellent scalability and flexibility, and significantly outperforms existing methods on LMAS benchmarks.
Qingxu Fu, Tenghai Qiu, Jianqiang Yi, Zhiqiang Pu, Shiguang Wu 0001
AAAI4
2022 Knowledge Transfer from Situation Evaluation to Multi-agent Reinforcement Learning
Min Chen 0038, Zhiqiang Pu, Yi Pan 0009, Jianqiang Yi
ICONIP (4)2
2022 Tacit Commitments Emergence in Multi-agent Reinforcement Learning
Boyin Liu, Zhiqiang Pu, Junlong Gao, Jianqiang Yi
ICONIP (1)2
2022 Multi-Target Encirclement with Collision Avoidance via Deep Reinforcement Learning using Relational Graphs
abstract
In this paper, we propose a novel decentralized method based on deep reinforcement learning using robot-level and target-level relational graphs, to solve the problem of multi-target encirclement with collision avoidance (MECA). Specifically, the robot-level relational graphs, composed of three heterogeneous relational graphs between each robot and other robots, targets and obstacles, are modeled and learned through using graph attention networks (GATs) for extracting different spatial relational representations. Moreover, for each target within the observation of each robot, a target-level relational graph is built with GAT to construct spatial relations from the robot. Furthermore, the movement of each target is modeled by the target-level relational graph and learned through supervised learning for predicting the trajectory of the target. In addition, a knowledge-embedded compound reward function is defined to solve the multi-objective problem in MECA, and guide the policy learning for deriving the behavior of MECA. An actor-critic training algorithm based on the centralized training and decentralized execution framework is adopted to train the policy network. Simulation and real-world experiment results demonstrate the effectiveness and generalization of our method.
Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi
ICRA3
2022 A Cooperation Graph Approach for Multiagent Sparse Reward Reinforcement Learning
abstract
Multiagent reinforcement learning (MARL) can solve complex cooperative tasks. However, the efficiency of existing MARL methods relies heavily on well-defined reward functions. Multiagent tasks with sparse reward feedback are especially challenging not only because of the credit distribution problem, but also due to the low probability of obtaining positive reward feedback. In this paper, we design a graph network called Cooperation Graph (CG). The Cooperation Graph is the combination of two simple bipartite graphs, namely, the Agent Clustering subgraph (ACG) and the Cluster Designating subgraph (CDG). Next, based on this novel graph structure, we propose a Cooperation Graph Multiagent Reinforcement Learning (CG-MARL) algorithm, which can efficiently deal with the sparse reward problem in multiagent tasks. In CG-MARL, agents are directly controlled by the Cooperation Graph. And a policy neural network is trained to manipulate this Cooperation Graph, guiding agents to achieve cooperation in an implicit way. This hierarchical feature of CG-MARL provides space for customized cluster-actions, an extensible interface for introducing fundamental cooperation knowledge. In experiments, CG-MARL shows state-of-the-art performance in sparse reward multiagent benchmarks, including the anti-invasion interception task and the multi-cargo delivery task.
Qingxu Fu, Tenghai Qiu, Zhiqiang Pu, Jianqiang Yi, Wanmai Yuan
IJCNN3
2022 Multi-Agent Local Information Reconstruction for Situational Cognition
abstract
Learning an effective strategy is challenging for agents in partially observable environment, where the agents can only observe a part of environment information and make decisions based on local information. Hence, how to effectively utilize local information to achieve efficient cooperation among the agents is particularly important. The agents can establish an understanding of themselves and their surrounding environment based on their historical observation information. However, the understanding lacking global information is local or limited, resulting in low performance in some complex tasks. To solve the problem, a situational cognition learning framework is proposed for each agent based on local information, with which each agent can reconstruct a cognition about itself and its surrounding environment and map it into a high dimensional representation space. In particular, situational cognition is modeled as a random variable under the condition of local trajectory. In addition, an information regularizer is introduced to ensure that the situational cognition is complete and accurate through maximizing the mutual information between the situational cognition and global information, conditioned on the local trajectory of the agent. Various simulations are conducted and show that the proposed framework significantly promotes cooperation among the agents and improves performance.
Shiguang Wu 0001, Zhiqiang Pu, Tenghai Qiu, Jianqiang Yi
IJCNN2
2022 Intrinsic Reward with Peer Incentives for Cooperative Multi-Agent Reinforcement Learning
abstract
In this paper, we propose a novel Intrinsic Reward method with Peer Incentives (IRPI) to promote the inter-agent direct interactions and implicitly address the credit assignment problem in cooperative multi-agent reinforcement learning (MARL). The IRPI method can build mutual incentives between agents by using their causal effect, to realize their advanced cooperation. Specifically, a new intrinsic reward mechanism is conducted, which equips each agent with the ability to reward other agent by using the causal effect between them. Moreover, the mechanism is built through a neural network and learned by using causal effect between the agents. Furthermore, the counterfactual reasoning is used to infer the causal effect between the agents using the joint action-state value function, and then assess the quality of the effect using individual state value function in MARL. Simulational results in Starcraft II Micromanagement demonstrate that the proposed IRPI can enhance cooperation among the RL agents to achieve better performance than some state-of-the-art MARL methods in various cooperative multi-aaent tasks.
Zhen Liu 0020, Shiguang Wu 0001, Zhiqiang Pu, Jianqiang Yi
IJCNN4
2022 Multi-UAV Cooperative Short-Range Combat via Attention-Based Reinforcement Learning using Individual Reward Shaping
abstract
In this paper, we propose a novel distributed method based on attention-based deep reinforcement learning using individual reward shaping, for multiple unmanned aerial vehicles (UAVs) cooperative short-range combat mission. Specifically, a two-level attention distributed policy, composed of observation-level and communication-level attention networks, is designed to enable each UAV to selectively focus on important environmental features and messages, for enhancing the effectiveness of the cooperative policy. Moreover, due to the high complexity and stochasticity of the UAV combat mission, the learning of UAVs is tricky and low efficient. To embed knowledge to accelerate the policy learning, a potential-based individual reward function is constructed by implicitly translating the individual reward into the specific form of dynamic action potentials. In addition, an actor-critic training algorithm based on the centralized training and decentralized execution framework is adopted to train the policy network of UAV maneuver decision. We build a three-dimensional UAV simulation and training platform based on Unity for multi-UAV short-range combat missions. Simulation results demonstrate the effectiveness of the proposed method and the superiority of the attention policy and individual reward shaping.
Tenghai Qiu, Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi, Jinying Zhu, Ruiguang Hu
IROS4
2022 Finite-Time Command-Filtered Composite Adaptive Neural Control of Uncertain Nonlinear Systems
abstract
This article presents a new command-filtered composite adaptive neural control scheme for uncertain nonlinear systems. Compared with existing works, this approach focuses on achieving finite-time convergent composite adaptive control for the higher-order nonlinear system with unknown nonlinearities, parameter uncertainties, and external disturbances. First, radial basis function neural networks (NNs) are utilized to approximate the unknown functions of the considered uncertain nonlinear system. By constructing the prediction errors from the serial-parallel nonsmooth estimation models, the prediction errors and the tracking errors are fused to update the weights of the NNs. Afterward, the composite adaptive neural backstepping control scheme is proposed via nonsmooth command filter and adaptive disturbance estimation techniques. The proposed control scheme ensures that high-precision tracking performances and NN approximation performances can be achieved simultaneously. Meanwhile, it can avoid the singularity problem in the finite-time backstepping framework. Moreover, it is proved that all signals in the closed-loop control system can be convergent in finite time. Finally, simulation results are given to illustrate the effectiveness of the proposed control scheme.
Jinlin Sun, Haibo He, Jianqiang Yi, Zhiqiang Pu
IEEE Trans. Cybern.4
2022 Fixed-Time Adaptive Fuzzy Control for Uncertain Nonstrict-Feedback Systems With Time-Varying Constraints and Input Saturations
abstract
This article investigates the adaptive fuzzy tracking control design for uncertain nonstrict-feedback nonlinear systems with time-varying constraints and asymmetric input saturations. It is known that time-varying output constraint, time-varying error constraints, and input saturations are commonly seen issues in many practical engineering systems due to inherent physical limitations and performance requirements. To deal with these problems and improve the convergence rate of the control system, a novel fixed-time convergent adaptive fuzzy control scheme is proposed for the considered uncertain nonlinear systems via the time-varying barrier Lyapunov function (BLF) technique. First, to deal with the unknown nonlinearities and uncertainties in the system dynamic model, fuzzy logic systems are utilized to approximate the unknown functions of the considered model. Then, the problem of asymmetric input amplitude and rate saturations is handled by constructing a unified smooth characterizing function and an innovative auxiliary design signal system. Subsequently, considering physical limitations and performance requirements, a novel time-varying BLF-based adaptive fuzzy backstepping control scheme is designed for the uncertain nonstrict-feedback nonlinear systems to realize superior tracking performances and keep the states staying in predefined time-varying compact regions during operations. Further, rigorous theoretical analyses have been conducted to show that the proposed control scheme achieves fixed-time convergence of all signals in the closed-loop control system. Finally, representative simulation results verify the effectiveness of the proposed control scheme.
Jinlin Sun, Jianqiang Yi, Zhiqiang Pu
IEEE Trans. Fuzzy Syst.3
2021 Semantic Perception Swarm Policy with Deep Reinforcement Learning
Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi
ICONIP (3)3
2021 Multi-target Coverage with Connectivity Maintenance using Knowledge-incorporated Policy Framework
abstract
This paper considers a multi-target coverage problem where a robot team aims to efficiently cover multi-targets while maintaining connectivity in a distributed manner. A novel knowledge-incorporated policy framework is proposed to derive a distributed, efficient, and connectivity guaranteed coverage policy. In particular, a knowledge-guided policy network (KGPnet) is designed, which consists of observation attention representation, interaction attention representation, and knowledge-guided policy learning. Giving credit to the KGPnet, the connectivity guaranteed coverage policy can be applied to different number targets. Moreover, based on the knowledge of the algebraic connectivity and coverage rate, a comprehensive reward is designed to guide the training of the behavior of multi-target coverage with connectivity maintenance. Furthermore, since the policy learned through deep reinforcement learning (DRL) can not guarantee the connectivity of the robot team, a knowledge-nested policy filtering is designed to filter dis-connectivity policies to satisfy the connectivity constraint based on the knowledge model of connectivity maintenance. Various simulations are conducted to verify the effectiveness of the proposed method. Besides, numerous real-world experiments with three-wheel omnidirectional cars and a motion capture system are presented to demonstrate the practicability of the proposed method.
Shiguang Wu 0001, Zhiqiang Pu, Zhen Liu 0020, Tenghai Qiu, Jianqiang Yi
ICRA2
2021 Multi-Agent Cognition Difference Reinforcement Learning for Multi-Agent Cooperation
abstract
Multi-agent cooperation is one of the most attractive research fields in multi-agent systems. There are many attempts made by researchers in this field to promote the cooperation behavior. However, in partially-observable environments, a large number of agents and complex interactions among the agents cause huge difficulty for policy learning. Moreover, redundant communication contents caused by many agents make effective features hard to be extracted, which prevents the policy from converging. To address the limitations above, a novel method called multi-agent cognition difference reinforcement learning (MACD-RL) is proposed in this paper. The key feature of MACD-RL lies in cognition difference network (CDN) and a soft communication network (SCN). CDN is designed to allow each agent to choose its neighbors (communication targets) adaptively with its environment cognition difference. SCN is designed to handle the complex interactions among the agents with soft attention mechanism. The results of simulations including mixed cooperative and competitive tasks demonstrate that the effectiveness and robustness of the proposed model.
Huimu Wang, Tenghai Qiu, Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi, Wanmai Yuan
IJCNN4
2021 Multi-agent Collaborative Learning with Relational Graph Reasoning in Adversarial Environments
abstract
This paper proposes a collaborative policy framework via relational graph reasoning for multi-agent systems to accomplish adversarial tasks. A relational graph reasoning module consisting of an agent graph reasoning module and an opponent graph module, is designed to enable each agent to learn mixture state representation to enhance the effectiveness of the policy. In particular, for each agent, the agent graph reasoning module is designed to infer different underlying influences from different opponents and generate agent-level state representation. The opponent graph reasoning module is creatively designed for the opponents to reason relations from their surrounding objects including the agents and the opponents based on their latent features and then predict the future state of the opponents. It forms an opponent-level state representation. Besides, in order to effectively predict the state of the opponents, an intrinsic reward based on prediction error is designed to motivate the policy learning. Furthermore, interactions among agents are utilized to transmit messages and fuse information to promote the cooperative behaviors among the agents. Finally, various representative simulations on two multi-agent adversarial tasks are conducted to demonstrate the superiority and effectiveness of the proposed framework by comparison with existing methods.
Shiguang Wu 0001, Tenghai Qiu, Zhiqiang Pu, Jianqiang Yi
IROS3
2021 Fixed-time adaptive observer-based time-varying formation control for multi-agent systems with directed topologies
Tianyi Xiong, Zhou Gu, Jianqiang Yi, Zhiqiang Pu
Neurocomputing4
2021 Robust Adaptive Tracking Control for Hypersonic Vehicle Based on Interval Type-2 Fuzzy Logic System and Small-Gain Approach
abstract
This paper presents a novel robust adaptive tracking control method for a hypersonic vehicle in a cruise flight stage based on interval type-2 fuzzy-logic system (IT2-FLS) and small-gain approach. After the input-output linearization, the vehicle model can be decomposed into two uncertain subsystems by considering matching disturbances and parametric uncertainties. For each subsystem, an interval type-2 Takagi-Sugeno-Kang fuzzy logic system (IT2-TSK-FLS) is then employed to approximate the unavailable model information. Following the idea of a small-gain approach, a composite feedback form for each subsystem is constructed, based on which the final robust adaptive tracking control law is developed. Rigorous stability analysis shows that all signals in the derived closed-loop system are kept uniformly ultimately bounded (UUB). The main contribution of this paper is that the proposed control law for the hypersonic vehicle is with only two adaptive parameters in total which can greatly alleviate the computation and storage burden in practice; meanwhile its superiority over the conventional minimal-learning-parameter (MLP)-based one is specifically illustrated. Comparative numerical simulations of three cases demonstrate the effectiveness of our proposed control method with respect to complicated uncertainties.
Xinlong Tao, Jianqiang Yi, Zhiqiang Pu, Tianyi Xiong
IEEE Trans. Cybern.3
2021 Formation Control With Collision Avoidance Through Deep Reinforcement Learning Using Model-Guided Demonstration
abstract
Generating collision-free, time-efficient paths in an uncertain dynamic environment poses huge challenges for the formation control with collision avoidance (FCCA) problem in a leader-follower structure. In particular, the followers have to take both formation maintenance and collision avoidance into account simultaneously. Unfortunately, most of the existing works are simple combinations of methods dealing with the two problems separately. In this article, a new method based on deep reinforcement learning (RL) is proposed to solve the problem of FCCA. Especially, the learning-based policy is extended to the field of formation control, which involves a two-stage training framework: an imitation learning (IL) and later an RL. In the IL stage, a model-guided method consisting of a consensus theory-based formation controller and an optimal reciprocal collision avoidance strategy is designed to speed up training and increase efficiency. In the RL stage, a compound reward function is presented to guide the training. In addition, we design a formation-oriented network structure to perceive the environment. Long short-term memory is adopted to enable the network structure to perceive the information of obstacles of an uncertain number, and a transfer training approach is adopted to improve the generalization of the network in different scenarios. Numerous representative simulations are conducted, and our method is further deployed to an experimental platform based on a multiomnidirectional-wheeled car system. The effectiveness and practicability of our proposed method are validated through both the simulation and experiment results.
Zezhi Sui, Zhiqiang Pu, Jianqiang Yi, Shiguang Wu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2021 On the Principle and Applications of Conditional Disturbance Negation
abstract
It is commonly believed that observer-based compensation is an effective way for disturbance rejection. A less talked about fact is that such disturbance rejection control technique may also degrade control performance. In this article, we present a typical cross-coupling system to reveal this problem and, more importantly, propose a new design principle of conditional disturbance negation (CDN) to eliminate its potential drawbacks of disturbance observer-based compensation. Qualitative analysis is first given for a general form of such cross-coupling systems, indicating the necessity of CDN. The analysis and control design principle of CDN is then exemplified through two applications. A numerical linear application produces abundant quantitative results through the powerful transfer function and frequency domain tools. A more complex nonlinear flexible air-breathing hypersonic vehicle application shows how conventional compensation deteriorates the couplings between rigid and flexible modes, and validates the effectiveness of CDN through comprehensive model analysis and simulation results. The proposed CDN design principle also arouses awareness of the importance of: 1) understanding the characteristics of the plant to be controlled and 2) recognizing the critical role the information plays in engineering practice.
Zhiqiang Pu, Jinlin Sun, Jianqiang Yi
IEEE Trans. Syst. Man Cybern. Syst.1
2020 STGA-LSTM: A Spatial-Temporal Graph Attentional LSTM Scheme for Multi-agent Cooperation
Huimu Wang, Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi
ICONIP (2)3
2020 Multi-agent Cooperation and Competition with Two-Level Attention Network
Shiguang Wu 0001, Zhiqiang Pu, Jianqiang Yi, Huimu Wang
ICONIP (2)2
2020 Multi-Robot Cooperative Target Encirclement through Learning Distributed Transferable Policy
abstract
Making efficient motion decisions for a multi-robot system is a challenging problem in target encirclement with collision avoidance. Specifically, each robot with local communication has to consider cooperative target encirclement and collision avoidance simultaneously. In this paper, a distributed transferable policy network framework based on deep reinforcement learning is proposed to solve the problem of multi-robot cooperative target encirclement with collision avoidance. The proposed policy network framework is able to process the information of uncertain number of robots and obstacles, which is a desirable property for multi-robot systems. In particular, graph attention communication mechanism is adopted to model multi-robot interactions as a graph and extract cooperative information from the graph. Long short-term memory is used to accept the states of uncertain number of obstacles. In addition, a compound reward is designed to lead the training of the behavior of target encirclement with collision avoidance. Curriculum learning is implemented to speed up the process of this training. Simulation results validate the effectiveness of the proposed algorithm. Moreover, we further show that the learned policy can directly transfer to different scenarios along with good generalization.
Zhen Liu 0020, Shiguang Wu 0001, Zhiqiang Pu, Jianqiang Yi
IJCNN4
2020 Adaptive Fuzzy Nonsmooth Backstepping Output-Feedback Control for Hypersonic Vehicles With Finite-Time Convergence
abstract
It is commonly believed that uncertainties are obstacles to the tracking performances of flexible air-breathing hypersonic vehicles (FAHVs). In addition, the estimation of unmeasured states in an FAHV makes this issue much more complicated. To deal with these difficulties, this article explores a novel adaptive fuzzy nonsmooth backstepping output-feedback control scheme for the FAHV with closed-loop finite-time convergence. First, to approximate the unknown dynamics of the FAHV, some function approximators are constructed by utilizing an interval type-2 (IT2) fuzzy logic system. On this basis, a fixed-time convergent adaptive IT2 fuzzy observer is designed to estimate the unmeasured flight path angle and the unmeasured angle of attack of the FAHV accurately and rapidly. Consequently, the estimation errors of the designed observer are convergent in a fixed time independent of their initial estimation errors. Furthermore, based on the estimated states, velocity and altitude tracking controllers are designed by using a finite-time adaptive IT2 fuzzy nonsmooth backstepping control technique, which avoids the problem of “explosion of complexity” in traditional backstepping methods. Subsequently, rigorous Lyapunov stability analysis is conducted to show the finite-time convergence of the closed-loop signals of the FAHV control system. Finally, the robustness and superiority of the proposed control scheme is validated by several simulations in representative scenarios.
Jinlin Sun, Jianqiang Yi, Zhiqiang Pu, Zhen Liu 0020
IEEE Trans. Fuzzy Syst.3
2020 Fixed-Time Control With Uncertainty and Measurement Noise Suppression for Hypersonic Vehicles via Augmented Sliding Mode Observers
abstract
Uncertainty and measurement noise are main obstacles that limit the tracking control performances of flexible air-breathing hypersonic vehicles (FAHVs). In this article, we propose a novel fixed-time convergent nonsmooth backstepping control scheme for FAHV via augmented sliding mode observers (ASMOs) to overcome these obstacles. The ASMOs are first designed for the FAHV dynamics by employing the measured states corrupted by noises as inputs. On one hand, the ASMOs can simultaneously estimate the uncertainties and filter out the measurement noises. On the other hand, the observation error of each ASMO can be convergent within a fixed time independent of its initial observation error. Then, based on the estimation results, the altitude and velocity tracking controllers are developed by using fixed-time nonsmooth backstepping technique. Afterwards, a Lyapunov-based stability analysis is given to illustrate the fixed-time convergence of the closed-loop signals of the FAHV control system. Finally, comparative simulations are conducted to illustrate the superiority of the proposed control scheme.
Jinlin Sun, Zhiqiang Pu, Jianqiang Yi, Zhen Liu 0020
IEEE Trans. Ind. Informatics2
2020 Fixed-Time Sliding Mode Disturbance Observer-Based Nonsmooth Backstepping Control for Hypersonic Vehicles
abstract
This paper presents a novel fixed-time sliding mode disturbance observer (SMDO)-based robust backstepping cruise tracking control scheme with closed-loop finite-time convergence for flexible air-breathing hypersonic vehicles (FAHVs). In order to enhance the control system's robustness, a fixed-time SMDO is designed to compensate for the flexibility effects, model uncertainties, and external disturbances in FAHVs. In consequence, fixed convergence time of disturbance observation is achieved independently of initial estimation errors. Furthermore, velocity and altitude continuous finite-time tracking controllers are constructed by incorporating the SMDO and nonsmooth backstepping technique. To solve the problem of “explosion of complexity” in the conventional backstepping approach, nonsmooth filters are specifically constructed to generate the derivatives of virtual control laws. A Lyapunov-based stability analysis is conducted to show the finite-time convergence of the closed-loop FAHV control system. Finally, several representative numerical simulations are given to illustrate the effectiveness and superiority of the proposed control strategy.
Jinlin Sun, Jianqiang Yi, Zhiqiang Pu, Xiangmin Tan
IEEE Trans. Syst. Man Cybern. Syst.3
2019 Formation Control with Collision Avoidance through Deep Reinforcement Learning
abstract
Generating collision free, time efficient paths for followers is a challenging problem in formation control with collision avoidance. Specifically, the followers have to consider both formation maintenance and collision avoidance at the same time. Recent works have shown the potentialities of deep reinforcement learning (DRL) to learn collision avoidance policies. However, only the collision factor was considered in the previous works. In this paper, we extend the learning-based policy to the area of formation control by learning a comprehensive task. In particular, a two-stage training scheme is adopted including imitation learning and reinforcement learning. A fusion reward function is proposed to lead the training. Besides, a formation-oriented network architecture is presented for environment perception and long short-term memory (LSTM) is applied to perceive the information of an arbitrary number of obstacles. Various simulations are carried out and the results show the proposed algorithm is able to anticipate the dynamic information of the environment and outperforms traditional methods.
Zezhi Sui, Zhiqiang Pu, Jianqiang Yi, Tianyi Xiong
IJCNN2
2019 Adaptive Neural Network Time-varying Formation Tracking Control for Multi-agent Systems via Minimal Learning Parameter Approach
abstract
This paper investigates the time-varying formation tracking control problem for multi-agent systems with consideration of model uncertainties. For each dimension of an agent, a radial basis function neural network (RBFNN) is first adopted to approximate the model uncertainties online. Taking the square of the norm of the neural network weight vector as a newly developed adaptive parameter, a novel RBFNN-based adaptive control law with minimal learning parameter (MLP) approach is then constructed to tackle the time-varying formation tracking problem. The uniformly ultimately boundedness (UUB) of formation tracking errors is guaranteed through Lyapunov analysis. Compared with other traditional RBFNN-based formation tracking control laws for multi-agent systems, very few parameters need to be updated online in our proposed one, which can greatly lessen the computational burden. Finally, comparative simulation results demonstrate the effectiveness and superiority of the proposed adaptive control law.
Tianyi Xiong, Zhiqiang Pu, Jianqiang Yi, Zezhi Sui
IJCNN2
2018 Desired Compensation Based Adaptive Fuzzy Control for Hypersonic Vehicle with Measurement Noises
abstract
In practice, the measurement noise arising from actual state measurement deteriorates the system performance significantly. To alleviate the noise effect on the hypersonic vehicle's pitch angle and pitch rate channels (i.e. attitude subsystem), this paper designs the fuzzy approximation based adaptive controller. Based on the desired compensation technique, a novel adaptive update law is proposed in the fuzzy approximator, in which the desired values are used to replace the actual states. Hence, the corresponding fuzzy approximator is able to estimate the unknown nonlinearity in the system even when measurement noises exist. In addition, the matching error of the fuzzy approximation and residual uncertainties are handled through an additional robust control term in the control law. Stability analysis is conducted by employing the Lyapunov method. Finally, nominal and comparative simulation results show that the control strategies proposed in this paper can guarantee the bounded tracking performance with consideration of the measurement noise and uncertainty.
Yifan Liu 0011, Jianqiang Yi, Zhiqiang Pu
FUZZ-IEEE3
2018 Path Planning of Multiagent Constrained Formation through Deep Reinforcement Learning
abstract
A parallel deep Q-network (DQN) algorithm is presented for solving multiagent constrained formation path planning, where reaching destination, avoiding obstacles, and maintaining formation are simultaneously considered as independent or interactive tasks. Parallel Q-networks are utilized for each agent to sense different feature information and learn independent behavior policy. Comprehensive reward function is designed in consideration of respective requirements and interaction constraints to correctly guide the training. In order to demonstrate the effectiveness of the algorithm, we build an end-to-end model by designing a pixel game. Both training and testing are carried out in the game with double dueling DQN and the results show that the parallel deep Q-network path planner eventually complete the three tasks very well.
Zezhi Sui, Zhiqiang Pu, Jianqiang Yi, Xiangmin Tan
IJCNN2
2018 Neural Network based Distributed Adaptive Time-varying Formation Control for Multi-UAV Systems with Varying Time Delays
abstract
This paper investigates the time-varying formation control problem for multiple unmanned aerial vehicle (multi- UAV) systems with unknown uncertainties and varying time delays. Firstly, a radial basis function neural network (RBFNN) is adopted to estimate the lumped model uncertainties online for compensation. Then, a novel RBFNN-based fully distributed adaptive control scheme consisting of control law to stabilize the system and adaptive law to adjust RBFNN weights is developed to tackle the time-varying formation tracking problem in the presence of varying time delay. The uniformly ultimately boundedness (UUB) of the formation tracking errors is theoretically analyzed through Lyapunov approach. Comparative simulation results demonstrate the effectiveness of the control and adaptive laws proposed in this paper.
Tianyi Xiong, Zhiqiang Pu, Jianqiang Yi
IJCNN2
2017 Interval type-2 TSK nominal-fuzzy-model-based sliding mode controller design for flexible air-breathing hypersonic vehicles
abstract
This paper presents a novel interval type-2 TSK nominal-fuzzy-model-based sliding mode controller (IT2-TSK-NFMSMC) for flexible air-breathing hypersonic vehicle (FAHV) in order to stress robustness of the control system in dealing with data-driven based fuzzy modelling deviations, system uncertainty and disturbances. We adopt backstepping structure decomposing FAHV model into 5 control subsystems and design controllers, respectively. More specifically, two subsystems are designed with integral sliding mode model controllers. Another three subsystems which directly coupling with flexible mode disturbances are designed with 1T2-TSK-NFMSMCs by the following steps: 1) interval type-2 TSK nominal-fuzzy-models (IT2-TSK-NFM) are generated automatically by using type-2 fuzzy self-organizing methods from experiment datasets; 2) nominal model sliding mode controllers are designed based on the IT2-TSK-NFM, respectively; 3) notch filters are adopted in order to decrease the disturbance effects from the flexible modes; 4) sliding mode compensation controllers are designed through Lyapunov synthesis in order to compensate differences between IT2-TSK-NFM and real models of the FAHV. Several scenarios are studied and the simulation results validate the robustness of the proposed controllers when there exist internal flexible vibration and external system disturbances.
Junlong Gao, Jianqiang Yi, Zhiqiang Pu, Chengdong Li
FUZZ-IEEE3