Zhixiao Sun

dblp:272/8110 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0003-0018-2337ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 9 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 IHGSL: Interpretable Heuristic Graph Structure Learning for Multi-Robot Autonomous Collaborative Systems
abstract
In multi-robot systems, capturing the complex and dynamic interaction relationships is essential for enhancing autonomous collaboration. However, existing learning-based approaches usually overlook the understanding of these relationships, leading to reliability issues and hindering their application to real-world scenarios. This paper proposes a novel approach called Interpretable Heuristic Graph Structure Learning (IHGSL) to better comprehend the complex collaborative relationships in multi-robot systems. We first construct a predicate space to define diverse predicates that express fundamental relationships. Then we employ the variational information bottleneck technique to acquire a latent representation of the current observation by aligning it with the historical trajectory. On this basis, the predicates that the robot should currently focus on the most are learned, and some interaction relationships are established accordingly. Thereby an interpretable relationship graph is generated heuristically to guide the achievement of multi-robot autonomous collaborative decision-making. Through experimental evaluation, we demonstrate the process of relationship inference, thus validating the interpretability of IHGSL. Compared with existing methods, IHGSL also achieves superior collaboration performance, which highlights the effectiveness of the learned heuristic graph structure.
Cuiwei Liu, Zhixiao Sun
IROS5
2025 IB-ToM: Human-AI Coordination for Unseen Partners with Evolving Strategies
Zhen Yang 0011, Yihang Hao, Mingkai Gao, Zhixiao Sun, Haiyin Piao
PRICAI5
2025 Boosting Weak-to-Strong Agents in Multiagent Reinforcement Learning via Balanced PPO
abstract
Multiagent policy gradients (MAPGs), an essential branch of reinforcement learning (RL), have made great progress in both industry and academia. However, existing models do not pay attention to the inadequate training of individual policies, thus limiting the overall performance. We verify the existence of imbalanced training in multiagent tasks and formally define it as an imbalance between policies (IBPs). To address the IBP issue, we propose a dynamic policy balance (DPB) model to balance the learning of each policy by dynamically reweighting the training samples. In addition, current methods for better performance strengthen the exploration of all policies, which leads to disregarding the training differences in the team and reducing learning efficiency. To overcome this drawback, we derive a technique named weighted entropy regularization (WER), a team-level exploration with additional incentives for individuals who exceed the team. DPB and WER are evaluated in homogeneous and heterogeneous tasks, effectively alleviating the imbalanced training problem and improving exploration efficiency. Furthermore, the experimental results show that our models can outperform the state-of-the-art MAPG methods and boast over 12.1% performance gain on average.
Sili Huang, Hechang Chen, Haiyin Piao, Zhixiao Sun, Yi Chang 0001, Lichao Sun 0001, Bo Yang 0002
IEEE Trans. Neural Networks Learn. Syst.4
2024 Discovering Expert-Level Air Combat Knowledge via Deep Excitatory-Inhibitory Factorized Reinforcement Learning
abstract
Artificial Intelligence (AI) has achieved a wide range of successes in autonomous air combat decision-making recently. Previous research demonstrated that AI-enabled air combat approaches could even acquire beyond human-level capabilities. However, there remains a lack of evidence regarding two major difficulties. First, the existing methods with fixed decision intervals are mostly devoted to solving what to act but merely pay attention to when to act, which occasionally misses optimal decision opportunities. Second, the method of an expert-crafted finite maneuver library leads to a lack of tactics diversity, which is vulnerable to an opponent equipped with new tactics. In view of this, we propose a novel Deep Reinforcement Learning (DRL) and prior knowledge hybrid autonomous air combat tactics discovering algorithm, namely deep E xcitatory-i N hibitory f ACT or I zed maneu VE r ( ENACTIVE ) learning. The algorithm consists of two key modules, i.e., ENHANCE and FACTIVE. Specifically, ENHANCE learns to adjust the air combat decision-making intervals and appropriately seize key opportunities. FACTIVE factorizes maneuvers and then jointly optimizes them with significant tactics diversity increments. Extensive experimental results reveal that the proposed method outperforms state-of-the-art algorithms with a 62% winning rate and further obtains a margin of a 2.85-fold increase in terms of global tactic space coverage. It also demonstrates that a variety of discovered air combat tactics are comparable to human experts’ knowledge.
Haiyin Piao, Shengqi Yang, Hechang Chen, Junnan Li 0008, Xuanqi Peng, Xin Yang 0011, Zhen Yang 0011, Zhixiao Sun, Yi Chang 0001
ACM Trans. Intell. Syst. Technol.9
2023 The Sufficiency of Off-Policyness and Soft Clipping: PPO Is Still Insufficient according to an Off-Policy Measure
abstract
The popular Proximal Policy Optimization (PPO) algorithm approximates the solution in a clipped policy space. Does there exist better policies outside of this space? By using a novel surrogate objective that employs the sigmoid function (which provides an interesting way of exploration), we found that the answer is "YES", and the better policies are in fact located very far from the clipped space. We show that PPO is insufficient in "off-policyness", according to an off-policy metric called DEON. Our algorithm explores in a much larger policy space than PPO, and it maximizes the Conservative Policy Iteration (CPI) objective better than PPO during training. To the best of our knowledge, all current PPO methods have the clipping operation and optimize in the clipped policy space. Our method is the first of this kind, which advances the understanding of CPI optimization and policy gradient methods. Code is available at https://github.com/raincchio/P3O.
Xing Chen 0022, Dongcui Diao, Hechang Chen, Hengshuai Yao, Haiyin Piao, Zhixiao Sun, Zhiwei Yang 0005, Randy Goebel, Bei Jiang, Yi Chang 0001
AAAI6
2023 Complex relationship graph abstraction for autonomous air combat collaboration: A learning and expert knowledge hybrid approach
Haiyin Piao, Hechang Chen, Xuanqi Peng, Songyuan Fan, Zhixiao Sun
Expert Syst. Appl.9
2023 Multi-agent air combat with two-stage graph-attention communication
Zhixiao Sun, Huahua Wu, Yandong Shi, Xiangchao Yu, Wenbin Pei, Zhen Yang 0011, Haiyin Piao, Yaqing Hou
Neural Comput. Appl.1
2022 Multi-UAV Disaster Environment Coverage Planning with Limited-Endurance
abstract
Disaster areas involving floods and earthquakes are commonly large, with the rescue time being quite tight, suggesting multi-Unmanned Aerial Vehicles (UAV) exploration rather than employing a single UAV. For such scenarios, current UAV exploration is modeled as a Coverage Path Planning (CPP) problem to achieve full area coverage in the presence of obstacles. However, the UAV's endurance capability is limited, and the rescue time is constrained, prohibiting even multiple UAVs from completing disaster area coverage on time. Therefore, this paper defines a multi-Agent Endurance-limited CPP (MAEl-CPP) problem that is based on an a priori known heatmap of the disaster area, which affords to explore the most valuable areas under UAV limited energy constraints. Furthermore, we propose a path planning algorithm for the MAEl-CPP problem by ranking the possible disaster areas according to their importance through satellite or remote sensing aerial images and completing path planning according to this ranking. Experimental results demonstrate that the search efficiency of the proposed algorithm is 4.2 times that of the existing algorithm.
Hongyu Song, Jiantao Qiu, Zhixiao Sun, Kuijun Lang, Yuan Shen 0001, Yu Wang 0002
ICRA4
2022 Deep Relationship Graph Reinforcement Learning for Multi-Aircraft Air Combat
abstract
Air combat Artificial Intelligence (AI) has attracted increasing attentions from aeronautics engineers and artificial intelligence researchers. However, it is often of great difficulties for the existing methods to solve the collaboration problems in multi-aircraft air combat due to their high complexity incurred by combination explosion. In view of this, we propose a Deep Relationship Graph Reinforcement Learning (DRGRL) algorithm for multi-aircraft collaboration. Specifically, DRGRL significantly simplifies the complex situation space via abstracting the original problem into a symbolic form. Besides, a novel Air Combat Relationship Graph (ACRG) is introduced to represent the learned collaboration pattern, which concentrates on the most important combat relationships for tactic decision making. Consequently, experiments are conducted in an air combat simulation environment named WUKONG. The comprehensive experimental results demonstrate that DRGRL could evidently learn some valuable collaboration patterns and achieve better combat performance than state-of-the-art air combat AI methods.
Haiyin Piao, Yaqing Hou, Zhixiao Sun, Shengqi Yang, Xuanqi Peng, Songyuan Fan
IJCNN5
2022 Learning Smooth Motion Planning for Intelligent Aerial Transportation Vehicles by Stable Auxiliary Gradient
abstract
Deep Reinforcement Learning (DRL) has been widely attempted for solving real-time intelligent aerial transportation vehicle motion planning tasks recently. When interacting with environment, DRL-driven aerial vehicles inevitably switch the steering actions in high frequency during both exploration and execution phase, resulting in the well known flight trajectory oscillation issue, which makes flight dynamics unstable, and even endangers flight safety in serious cases. Unfortunately, there is hardly any literature about achieving flight trajectory smoothness in DRL-based motion planning. In view of this, we originally formalize the practical flight trajectory smoothen problem as a three-level Nested pArameterized Smooth Trajectory Optimization (NASTO) form. On this basis, a novel Stable Auxiliary Gradient (SAG) algorithm is proposed, which significantly smoothens the DRL-generated flight motions by constructing two independent optimization aspects: the major gradient, and the stable auxiliary gradient. Experimental result reveals that the proposed SAG algorithm outperforms baseline DRL-based intelligent aerial transportation vehicle motion planning algorithms in terms of both learning efficiency and flight motion smoothness.
Haiyin Piao, Li Mo 0001, Xin Yang 0011, Zhixiao Sun, Zhen Yang 0011
IEEE Trans. Intell. Transp. Syst.6
2021 Multi-agent hierarchical policy gradient for Air Combat Tactics emergence via self-play
Zhixiao Sun, Haiyin Piao, Zhen Yang 0011, Guang Zhan, Guanglei Meng, Hechang Chen, Xing Chen 0022, Bohao Qu, Yuanjie Lu
Eng. Appl. Artif. Intell.1
2020 Beyond-Visual-Range Air Combat Tactics Auto-Generation by Reinforcement Learning
abstract
For quite a long time, effective Beyond-Visual-Range (BVR) air combat tactics can only be discovered by human pilots in the actual combat process. However, due to the lack of actual combat opportunities, making new air combat tactics innovation was generally considered quite difficult. To address this challenge, we first introduced a solely end-to-end Reinforcement Learning (RL) approach for training competitive air combat agents with adversarial self-play from scratch in a high fidelity air combat simulation environment during training. Furthermore, a Key Air Combat Event Reward Shaping (KAERS) mechanism was proposed to provide sparse but objective shaped rewards beyond episodic win/lose signal to accelerate the initial machine learning process. Experimental results showed that multiple valuable air combat tactical behaviors emerged progressively. We hope this study could be extended to the future of air combat machine intelligence research.
Haiyin Piao, Zhixiao Sun, Guanglei Meng, Hechang Chen, Bohao Qu, Kuijun Lang, Shengqi Yang, Xuanqi Peng
IJCNN2
2020 MA-TREX: Mutli-agent Trajectory-Ranked Reward Extrapolation via Inverse Reinforcement Learning
Sili Huang, Bo Yang 0002, Hechang Chen, Haiyin Piao, Zhixiao Sun, Yi Chang 0001
KSEM (2)5