EDBT 2026 Demo / reviewers in the wild / expert
Zhen Liu 0020
dblp:77/35-20
· DBLP profile ↗
20ranked-venue papers
1as first author
14since 2021 · last 2025
0000-0003-1610-2338ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 13 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TAAE: Team-Aware Attention Extraction for Generalized Agent Coordination in Multi-Agent Reinforcement LearningabstractMulti-agent reinforcement learning(MARL) is a highly monitored field to solving complex decision-making problems, simulating and optimizing the interactions and collaborations of multiple agents in dynamic environments. In this work, we introduce the Team-Aware Attention Extraction framework(TAAE), a novel framework for improving agent coordination in multi-agent reinforcement learning through team-aware attention extraction. The challenges posed by conventional centralized training with decentralized execution (CTDE) approaches are addressed by our method, which often assumes agent homogeneity and overlooks the importance of team-specific information. A team encoder is used by TAAE to categorize agents into specialized teams with distinct roles, and an attention mechanism is exploited to extract relevant information for each team, thus increasing training efficiency. In addition, conditional generative adversarial networks (cGAN) are utilized to facilitate decentralized execution by approximating global information. TAAE is evaluated on the StarCraft II micromanagement benchmark, and its superior performance compared to the baseline algorithms, particularly on complex maps, is demonstrated. Our results confirm the effectiveness of team-based grouping and attention extraction, as well as the robustness of cGAN to maintain high performance during decentralized execution. Zhenfeng Su, Jiaosai Li, Ren Wen, Zhen Liu 0020 |
IJCNN | 5 |
| 2025 | Cognition-Oriented Multiagent Reinforcement LearningabstractInspired by psychological insights into individual behavior, we propose a novel cognition-oriented multiagent reinforcement learning (CORL) framework. CORL equips agents with two distinct types of cognition-situational and self-cognition-derived from local observations. To enhance the informativeness and precision of these cognition types, we introduce two information-theoretical regularizers: one to align situational cognition with the global state and the other to align self-cognition with each agent's identity for improved role differentiation and team coordination. In addition, the centralized training and decentralized execution framework is adopted to train the policy network. Our simulations demonstrate that CORL effectively harnesses local observations for enriched cooperation, leading to pronounced performance improvements, particularly in challenging tasks. Tenghai Qiu, Shiguang Wu 0001, Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi, Yuqian Zhao 0001, Biao Luo 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | RBHAR: Role-Based Heterogeneous Action Representation in Multi-agent Reinforcement Learning
Jiaosai Li, Zhen Liu 0020, Zhenfeng Su |
ICONIP (2) | 2 |
| 2024 | Multiexperience-Assisted Efficient Multiagent Reinforcement LearningabstractRecently, multiagent reinforcement learning (MARL) has shown great potential for learning cooperative policies in multiagent systems (MASs). However, a noticeable drawback of current MARL is the low sample efficiency, which causes a huge amount of interactions with environment. Such amount of interactions greatly hinders the real-world application of MARL. Fortunately, effectively incorporating experience knowledge can assist MARL to quickly find effective solutions, which can significantly alleviate the drawback. In this article, a novel multiexperience-assisted reinforcement learning (MEARL) method is proposed to improve the learning efficiency of MASs. Specifically, monotonicity-constrained reward shaping is innovatively designed using expert experience to provide additional individual rewards to guide multiagent learning efficiently, with the invariance guarantee of the team optimization objective. Furthermore, a reward distribution estimator is specially developed to model an implicated reward distribution of environment by using transition experience from environment, containing collected samples (state-action pair, reward, and next state). This estimator can predict the expectation reward of each agent for the taken action to accurately estimate the state value function and accelerate its convergence. Besides, the performance of MEARL is evaluated on two multiagent environment platforms: our designed unmanned aerial vehicle combat (UAV-C) and StarCraft II Micromanagement (SCII-M). Simulation results demonstrate that the proposed MEARL can greatly improve the learning efficiency and performance of MASs and is superior to the state-of-the-art methods in multiagent tasks. Zhen Liu 0020, Jianqiang Yi, Shiguang Wu 0001, Zhiqiang Pu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Learning Cooperative Policies with Graph Networks in Distributed Swarm SystemsabstractDeriving efficient cooperative policies in uncertain dynamic environments poses huge challenges for a distributed swarm system due to the limited capability of the agents and the complex dynamics of the environment. In this paper, a novel distributed method based on deep reinforcement learning using observation-level and communication-level graph networks is proposed to learn cooperative policies for the distributed swarm system. Specifically, a relational directed graph attention neural network is designed to model observation-level graphs composed of heterogeneous relational graphs among each agent and each type of entities (e.g., obstacles, other teammates, opponents), for extracting different relational representations. Moreover, a relevant directed graph attention network is presented to cut off the ineffective communication among irrelevant agents, and model a relevant communication topology between each agent and relevant homogeneous neighbor agents as an communication-level graph, for promoting efficient inter-agent interactions. Furthermore, a distributed actor-critic algorithm with full parameter sharing is implemented to learn cooperative swarm policies by using distributed critics, which avoids the curse of dimensionality under a centralized critic. Various simulation results validate the effectiveness and generalization of the proposed method, and demonstrate that the proposed method outperforms existing state-of-the-art methods on coverage and pursuit tasks. Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi, Xiaolin Ai, Wanmai Yuan |
IJCNN | 2 |
| 2023 | Peer Incentive Reinforcement Learning for Cooperative Multiagent GamesabstractSocial learning, especially social incentives, is extremely important for humans to achieve a high level of coordination. Inspired by this, we introduce this concept into cooperative multiagent reinforcement learning (MARL), to implicitly address the credit assignment problem and promote the interagent direct interactions for cooperations among agents in cooperative multiagent games. In this article, we propose a novel intrinsic reward method with peer incentives (IRPI) based on actor–critic policy gradient. This method can enable agents to incentivize each other for their cooperations through using causal influence among them. Specifically, a novel intrinsic reward mechanism is innovatively designed to empower each agent the ability to give positive or negative rewards to other peer agents' actions through considering the causal influence of the other agents on it. The mechanism is realized by a feedforward neural network through utilizing causal influence between the agents. The causal influence of one agent on another is inferred via counterfactual reasoning using the joint action-value function in MARL. The quality of the influence is assessed via counterfactual reasoning using the individual value function in MARL. Simulations are carried out on two popular multiagent game testbeds: Starcraft II Micromanagement and Multiagent Particle Environments. Simulation results demonstrate that the proposed IRPI can enhance cooperations among the agents to achieve better performance compared with a number of state-of-the-art MARL methods in a variety of cooperative multiagent games. Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi |
IEEE Trans. Games | 2 |
| 2023 | Attention Enhanced Reinforcement Learning for Multi agent CooperationabstractIn this article, a novel method, called attention enhanced reinforcement learning (AERL), is proposed to address issues including complex interaction, limited communication range, and time-varying communication topology for multi agent cooperation. AERL includes a communication enhanced network (CEN), a graph spatiotemporal long short-term memory network (GST-LSTM), and parameters sharing multi-pseudo critic proximal policy optimization (PS-MPC-PPO). Specifically, CEN based on graph attention mechanism is designed to enlarge the agents' communication range and to deal with complex interaction among the agents. GST-LSTM, which replaces the standard fully connected (FC) operator in LSTM with graph attention operator, is designed to capture the temporal dependence while maintaining the spatial structure learned by CEN. PS-MPC-PPO, which extends proximal policy optimization (PPO) in multi agent systems with parameters' sharing to scale to environments with a large number of agents in training, is designed with multi-pseudo critics to mitigate the bias problem in training and accelerate the convergence process. Simulation results for three groups of representative scenarios including formation control, group containment, and predator-prey games demonstrate the effectiveness and robustness of AERL. Zhiqiang Pu, Huimu Wang, Zhen Liu 0020, Jianqiang Yi, Shiguang Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Learning Continuous 3-DoF Air-to-Air Close-in Combat Strategy using Proximal Policy OptimizationabstractAir-to-air close-in combat is based on many basic fighter maneuvers and can be largely modeled as an algorithmic function of inputs. This paper studies autonomous close-in combat, to learn new strategy that can adapt to different circumstances to fight against an opponent. Current methods for learning close-in combat strategy are largely limited to discrete action sets whether in the form of rules, actions or sub-polices. In contrast, we consider one-on-one air combat game with continuous action space and present a deep reinforcement learning method based on proximal policy optimization (PPO) that learns close-in combat strategy from observations in an end-to-end manner. The state space is designed to promote the learning efficiency of PPO. We also design a minimax strategy for the game. Simulation results show that the learned PPO agent is able to defeat the minimax opponent with about 97% win rate. Luntong Li, Jiajun Chai, Zhen Liu 0020, Yuanheng Zhu, Jianqiang Yi |
CoG | 4 |
| 2022 | Multi-Target Encirclement with Collision Avoidance via Deep Reinforcement Learning using Relational GraphsabstractIn this paper, we propose a novel decentralized method based on deep reinforcement learning using robot-level and target-level relational graphs, to solve the problem of multi-target encirclement with collision avoidance (MECA). Specifically, the robot-level relational graphs, composed of three heterogeneous relational graphs between each robot and other robots, targets and obstacles, are modeled and learned through using graph attention networks (GATs) for extracting different spatial relational representations. Moreover, for each target within the observation of each robot, a target-level relational graph is built with GAT to construct spatial relations from the robot. Furthermore, the movement of each target is modeled by the target-level relational graph and learned through supervised learning for predicting the trajectory of the target. In addition, a knowledge-embedded compound reward function is defined to solve the multi-objective problem in MECA, and guide the policy learning for deriving the behavior of MECA. An actor-critic training algorithm based on the centralized training and decentralized execution framework is adopted to train the policy network. Simulation and real-world experiment results demonstrate the effectiveness and generalization of our method. Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi |
ICRA | 2 |
| 2022 | Intrinsic Reward with Peer Incentives for Cooperative Multi-Agent Reinforcement LearningabstractIn this paper, we propose a novel Intrinsic Reward method with Peer Incentives (IRPI) to promote the inter-agent direct interactions and implicitly address the credit assignment problem in cooperative multi-agent reinforcement learning (MARL). The IRPI method can build mutual incentives between agents by using their causal effect, to realize their advanced cooperation. Specifically, a new intrinsic reward mechanism is conducted, which equips each agent with the ability to reward other agent by using the causal effect between them. Moreover, the mechanism is built through a neural network and learned by using causal effect between the agents. Furthermore, the counterfactual reasoning is used to infer the causal effect between the agents using the joint action-state value function, and then assess the quality of the effect using individual state value function in MARL. Simulational results in Starcraft II Micromanagement demonstrate that the proposed IRPI can enhance cooperation among the RL agents to achieve better performance than some state-of-the-art MARL methods in various cooperative multi-aaent tasks. Zhen Liu 0020, Shiguang Wu 0001, Zhiqiang Pu, Jianqiang Yi |
IJCNN | 2 |
| 2022 | Multi-UAV Cooperative Short-Range Combat via Attention-Based Reinforcement Learning using Individual Reward ShapingabstractIn this paper, we propose a novel distributed method based on attention-based deep reinforcement learning using individual reward shaping, for multiple unmanned aerial vehicles (UAVs) cooperative short-range combat mission. Specifically, a two-level attention distributed policy, composed of observation-level and communication-level attention networks, is designed to enable each UAV to selectively focus on important environmental features and messages, for enhancing the effectiveness of the cooperative policy. Moreover, due to the high complexity and stochasticity of the UAV combat mission, the learning of UAVs is tricky and low efficient. To embed knowledge to accelerate the policy learning, a potential-based individual reward function is constructed by implicitly translating the individual reward into the specific form of dynamic action potentials. In addition, an actor-critic training algorithm based on the centralized training and decentralized execution framework is adopted to train the policy network of UAV maneuver decision. We build a three-dimensional UAV simulation and training platform based on Unity for multi-UAV short-range combat missions. Simulation results demonstrate the effectiveness of the proposed method and the superiority of the attention policy and individual reward shaping. Tenghai Qiu, Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi, Jinying Zhu, Ruiguang Hu |
IROS | 3 |
| 2021 | Semantic Perception Swarm Policy with Deep Reinforcement Learning
Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi |
ICONIP (3) | 2 |
| 2021 | Multi-target Coverage with Connectivity Maintenance using Knowledge-incorporated Policy FrameworkabstractThis paper considers a multi-target coverage problem where a robot team aims to efficiently cover multi-targets while maintaining connectivity in a distributed manner. A novel knowledge-incorporated policy framework is proposed to derive a distributed, efficient, and connectivity guaranteed coverage policy. In particular, a knowledge-guided policy network (KGPnet) is designed, which consists of observation attention representation, interaction attention representation, and knowledge-guided policy learning. Giving credit to the KGPnet, the connectivity guaranteed coverage policy can be applied to different number targets. Moreover, based on the knowledge of the algebraic connectivity and coverage rate, a comprehensive reward is designed to guide the training of the behavior of multi-target coverage with connectivity maintenance. Furthermore, since the policy learned through deep reinforcement learning (DRL) can not guarantee the connectivity of the robot team, a knowledge-nested policy filtering is designed to filter dis-connectivity policies to satisfy the connectivity constraint based on the knowledge model of connectivity maintenance. Various simulations are conducted to verify the effectiveness of the proposed method. Besides, numerous real-world experiments with three-wheel omnidirectional cars and a motion capture system are presented to demonstrate the practicability of the proposed method. Shiguang Wu 0001, Zhiqiang Pu, Zhen Liu 0020, Tenghai Qiu, Jianqiang Yi |
ICRA | 3 |
| 2021 | Multi-Agent Cognition Difference Reinforcement Learning for Multi-Agent CooperationabstractMulti-agent cooperation is one of the most attractive research fields in multi-agent systems. There are many attempts made by researchers in this field to promote the cooperation behavior. However, in partially-observable environments, a large number of agents and complex interactions among the agents cause huge difficulty for policy learning. Moreover, redundant communication contents caused by many agents make effective features hard to be extracted, which prevents the policy from converging. To address the limitations above, a novel method called multi-agent cognition difference reinforcement learning (MACD-RL) is proposed in this paper. The key feature of MACD-RL lies in cognition difference network (CDN) and a soft communication network (SCN). CDN is designed to allow each agent to choose its neighbors (communication targets) adaptively with its environment cognition difference. SCN is designed to handle the complex interactions among the agents with soft attention mechanism. The results of simulations including mixed cooperative and competitive tasks demonstrate that the effectiveness and robustness of the proposed model. Huimu Wang, Tenghai Qiu, Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi, Wanmai Yuan |
IJCNN | 3 |
| 2020 | STGA-LSTM: A Spatial-Temporal Graph Attentional LSTM Scheme for Multi-agent Cooperation
Huimu Wang, Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi |
ICONIP (2) | 2 |
| 2020 | Multi-Robot Cooperative Target Encirclement through Learning Distributed Transferable PolicyabstractMaking efficient motion decisions for a multi-robot system is a challenging problem in target encirclement with collision avoidance. Specifically, each robot with local communication has to consider cooperative target encirclement and collision avoidance simultaneously. In this paper, a distributed transferable policy network framework based on deep reinforcement learning is proposed to solve the problem of multi-robot cooperative target encirclement with collision avoidance. The proposed policy network framework is able to process the information of uncertain number of robots and obstacles, which is a desirable property for multi-robot systems. In particular, graph attention communication mechanism is adopted to model multi-robot interactions as a graph and extract cooperative information from the graph. Long short-term memory is used to accept the states of uncertain number of obstacles. In addition, a compound reward is designed to lead the training of the behavior of target encirclement with collision avoidance. Curriculum learning is implemented to speed up the process of this training. Simulation results validate the effectiveness of the proposed algorithm. Moreover, we further show that the learned policy can directly transfer to different scenarios along with good generalization. Zhen Liu 0020, Shiguang Wu 0001, Zhiqiang Pu, Jianqiang Yi |
IJCNN | 2 |
| 2020 | Adaptive Fuzzy Nonsmooth Backstepping Output-Feedback Control for Hypersonic Vehicles With Finite-Time ConvergenceabstractIt is commonly believed that uncertainties are obstacles to the tracking performances of flexible air-breathing hypersonic vehicles (FAHVs). In addition, the estimation of unmeasured states in an FAHV makes this issue much more complicated. To deal with these difficulties, this article explores a novel adaptive fuzzy nonsmooth backstepping output-feedback control scheme for the FAHV with closed-loop finite-time convergence. First, to approximate the unknown dynamics of the FAHV, some function approximators are constructed by utilizing an interval type-2 (IT2) fuzzy logic system. On this basis, a fixed-time convergent adaptive IT2 fuzzy observer is designed to estimate the unmeasured flight path angle and the unmeasured angle of attack of the FAHV accurately and rapidly. Consequently, the estimation errors of the designed observer are convergent in a fixed time independent of their initial estimation errors. Furthermore, based on the estimated states, velocity and altitude tracking controllers are designed by using a finite-time adaptive IT2 fuzzy nonsmooth backstepping control technique, which avoids the problem of “explosion of complexity” in traditional backstepping methods. Subsequently, rigorous Lyapunov stability analysis is conducted to show the finite-time convergence of the closed-loop signals of the FAHV control system. Finally, the robustness and superiority of the proposed control scheme is validated by several simulations in representative scenarios. Jinlin Sun, Jianqiang Yi, Zhiqiang Pu, Zhen Liu 0020 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2020 | Fixed-Time Control With Uncertainty and Measurement Noise Suppression for Hypersonic Vehicles via Augmented Sliding Mode ObserversabstractUncertainty and measurement noise are main obstacles that limit the tracking control performances of flexible air-breathing hypersonic vehicles (FAHVs). In this article, we propose a novel fixed-time convergent nonsmooth backstepping control scheme for FAHV via augmented sliding mode observers (ASMOs) to overcome these obstacles. The ASMOs are first designed for the FAHV dynamics by employing the measured states corrupted by noises as inputs. On one hand, the ASMOs can simultaneously estimate the uncertainties and filter out the measurement noises. On the other hand, the observation error of each ASMO can be convergent within a fixed time independent of its initial observation error. Then, based on the estimation results, the altitude and velocity tracking controllers are developed by using fixed-time nonsmooth backstepping technique. Afterwards, a Lyapunov-based stability analysis is given to illustrate the fixed-time convergence of the closed-loop signals of the FAHV control system. Finally, comparative simulations are conducted to illustrate the superiority of the proposed control scheme. Jinlin Sun, Zhiqiang Pu, Jianqiang Yi, Zhen Liu 0020 |
IEEE Trans. Ind. Informatics | 4 |
| 2017 | Control of a flexible air-breathing hypersonic vehicle with measurement noises using adaptive interval type-2 fuzzy logic systemabstractIn this paper, a novel control scheme for the flexible air-breathing hypersonic vehicle (FAHV) using adaptive interval type-2 fuzzy logic system (AIT2-FLS) is proposed to reduce the side effects of measurement noises in the velocity channel and altitude channel as well as flexible dynamics in real applications. After input-output linearization of the longitudinal model of FAHV, the dynamic inversion controller is formulated to track the reference commands based on state feedback. The AIT2-FLS is further developed to deal with the model uncertainties and input errors. Besides, the state estimator is applied to estimate the true values of the corrupted outputs. The stability characteristics of both the controller and the state estimator are analyzed. The whole control scheme is finally obtained through combining the controller and the state estimator based on the separation principle. Simulation results demonstrate the robustness of our proposed control scheme against measurement noises and flexibilities. Xinlong Tao, Jianqiang Yi, Ruyi Yuan, Zhen Liu 0020 |
FUZZ-IEEE | 4 |
| 2016 | Immersion and Invariance-Based Output Feedback Control of Air-Breathing Hypersonic VehiclesabstractA new output feedback control design for robust velocity and altitude tracking of an air-breathing hypersonic vehicle (AHSV) is presented in this paper. The control scheme is performed on the assumption that only partial states of AHSV are measurable. The key idea is to employ the immersion and invariance approach to design globally asymptotically stable observers for the unmeasurable states. For controller design, the whole control architecture is constructed using dynamic surface control, based on the decomposition of the longitudinal dynamics of AHSV into velocity and altitude subsystems. Stability analysis is presented using the Lyapunov theory. Representative simulations are carried out on the high-fidelity model, which illustrate the effectiveness and robustness of the proposed scheme. Zhen Liu 0020, Xiang-min Tan, Ruyi Yuan, Jianqiang Yi |
IEEE Trans Autom. Sci. Eng. | 1 |