Haiyin Piao

dblp:269/4228 · DBLP profile ↗
← Back
30ranked-venue papers
5as first author
27since 2021 · last 2026
0000-0002-8519-4750ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 2 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Simple Unified Uncertainty-Guided Framework for Offline-to-Online Reinforcement Learning
abstract
Offline reinforcement learning (RL) provides a promising solution to learning an agent fully relying on a data-driven paradigm. However, constrained by the limited quality of the offline dataset, its performance is often suboptimal. Therefore, it is desired to further finetune the agent via extra online interactions before deployment. Unfortunately, offline-to-online RL can be challenging due to two main challenges: constrained exploratory behavior and state-action distribution shift. In view of this, we propose a simple unified uncertainty-guided (SUNG) framework, which naturally unifies the solution to both challenges with the tool of uncertainty. Specifically, SUNG quantifies uncertainty via a variational autoencoder (VAE)-based state-action visitation density estimator. To facilitate efficient exploration, SUNG presents a practical optimistic exploration strategy to select informative actions with both high value and high uncertainty. Moreover, SUNG develops an adaptive exploitation method by applying conservative offline RL objectives to high-uncertainty samples and standard online RL objectives to low-uncertainty samples to smoothly bridge offline and online stages. SUNG achieves state-of-the-art online finetuning performance when combined with different offline RL methods, across various environments and datasets in the D4RL benchmark. Codes are made publicly available in https://github.com/guosyjlu/ACBR.
Siyuan Guo 0001, Yanchao Sun, Jifeng Hu, Sili Huang, Hechang Chen, Haiyin Piao, Lichao Sun 0001, Yi Chang 0001
IEEE Trans. Neural Networks Learn. Syst.6
2025 Transformer-Based Multi-Agent Reinforcement Learning Method With Credit-Oriented Strategy Differentiation
abstract
The problem of Multi-Agent Reinforcement Learning (MARL) shows a high level of both complexity in the environment and coordination between agents. In order to scale the algorithm to large-scale agent scenarios, neural networks designed for MARL are typically implemented with parameter sharing. These characteristics result in the challenges of partial observability, credit assignment and strategy homogenization. In this paper, a Transformer-Based Multi-Agent Reinforcement Learning Method With Credit-Oriented Strategy Differentiation (TMRC) is presented to address each of these challenges. First, we design a Temporal-Spatial Encoding module and an Attention-Based Value Decomposition module based on the Transformer architecture. The former leverages both temporal and spatial observation information, compensating for the missing environmental perspectives due to partial observability. The latter is designed to identify each agent’s individual contribution in complex interactions, effectively optimizing the credit assignment process. Then, we propose a Credit-Oriented Strategy Differentiation module that differentiates the entity representations of each agent based on their current task differences, allowing agents to have distinct real-time strategies, effectively mitigating the issue of strategy homogenization. We evaluate the proposed method on the SMAC benchmark. It demonstrates better final performance, faster convergence, and greater stability compared to other comparative methods. Additionally, a series of experiments are conducted to validate the effectiveness of the proposed modules. Our code is available at https://github.com/Hkxuan/TMRC.git.
Kaixuan Huang, Bo Jin 0001, Haiyin Piao, Ziqi Wei 0001
IROS4
2025 IB-ToM: Human-AI Coordination for Unseen Partners with Evolving Strategies
Zhen Yang 0011, Yihang Hao, Mingkai Gao, Zhixiao Sun, Haiyin Piao
PRICAI6
2025 Cooperative Multiagent Learning and Exploration With Min-Max Intrinsic Motivation
abstract
In the field of multiagent reinforcement learning (MARL), the ability to effectively explore unknown environments and collect information and experiences that are most beneficial for policy learning represents a critical research area. However, existing work often encounters difficulties in addressing the uncertainties caused by state changes and the inconsistencies between agents' local observations and global information, which presents significant challenges to coordinated exploration among multiple agents. To address this issue, this article proposes a novel MARL exploration method with Min-Max intrinsic motivation (E2M) that promotes the learning of joint policies of agents by introducing surprise minimization and social influence maximization. Since the agent is subject to unstable state changes in the environment, we introduce surprise minimization by computing state entropy to encourage the agents to cope with more stable and familiar situations. This method enables surprise estimation based on the low-dimensional representation of states obtained from random encoders. Furthermore, to prevent surprise minimization from leading to conservative policies, we introduce mutual information between agents' behaviors as social influence. By maximizing social influence, the agents are encouraged to interact to facilitate the emergence of cooperative behavior. The performance of our proposed E2M is testified across a range of popular StarCraft II and Multiagent MuJoCo tasks. Comprehensive results demonstrate its effectiveness in enhancing the cooperative capability of the multiple agents.
Yaqing Hou, Haiyin Piao, Yifeng Zeng, Yew-Soon Ong, Yaochu Jin, Qiang Zhang 0008
IEEE Trans. Cybern.3
2025 Generalizable Causal Reinforcement Learning for Out-of-Distribution Environments
abstract
Out-of-distribution (OOD) generalization is critical for applying reinforcement learning algorithms to real-world applications. To address the OOD problem, recent works focus on learning an OOD adaptation policy by capturing the causal factors affecting the environmental dynamics. However, these works recover the causal factors with only an entangled or binary form, resulting in a limited generalization of the policy that requires extra data from the testing environments. To break this limitation, we propose generalizable causal reinforcement learning (GCRL) to learn a disentangled representation of causal factors, on the basis of which we learn a policy that achieves the OOD generalization without extra training. For capturing the causal factors, GCRL deploys a weakly supervised signal with a two-stage constraint to ensure that all factors can be disentangled. Then, to achieve the OOD generalization through causal factors, we establish the dependence of actions on the learned representation and optimize the policy model across multiple environments. Experimental results show that the established dependence recovers the correct relationship between causal factors and actions when the learned policy could address the target tasks in training environments. Benefiting from the recovered relationship, GCRL achieves the OOD generalization on eight benchmarks from Causal World and Mujoco. Moreover, the policy learned by our model is more explainable and can be controlled to generate semantic actions by intervening in the representation of causal factors.
Sili Huang, Jifeng Hu, Hechang Chen, Peng Cui 0001, Haiyin Piao, Lichao Sun 0001, Bo Yang 0002
IEEE Trans. Ind. Informatics5
2025 Boosting Weak-to-Strong Agents in Multiagent Reinforcement Learning via Balanced PPO
abstract
Multiagent policy gradients (MAPGs), an essential branch of reinforcement learning (RL), have made great progress in both industry and academia. However, existing models do not pay attention to the inadequate training of individual policies, thus limiting the overall performance. We verify the existence of imbalanced training in multiagent tasks and formally define it as an imbalance between policies (IBPs). To address the IBP issue, we propose a dynamic policy balance (DPB) model to balance the learning of each policy by dynamically reweighting the training samples. In addition, current methods for better performance strengthen the exploration of all policies, which leads to disregarding the training differences in the team and reducing learning efficiency. To overcome this drawback, we derive a technique named weighted entropy regularization (WER), a team-level exploration with additional incentives for individuals who exceed the team. DPB and WER are evaluated in homogeneous and heterogeneous tasks, effectively alleviating the imbalanced training problem and improving exploration efficiency. Furthermore, the experimental results show that our models can outperform the state-of-the-art MAPG methods and boast over 12.1% performance gain on average.
Sili Huang, Hechang Chen, Haiyin Piao, Zhixiao Sun, Yi Chang 0001, Lichao Sun 0001, Bo Yang 0002
IEEE Trans. Neural Networks Learn. Syst.3
2024 OVD-Explorer: Optimism Should Not Be the Sole Pursuit of Exploration in Noisy Environments
abstract
In reinforcement learning, the optimism in the face of uncertainty (OFU) is a mainstream principle for directing exploration towards less explored areas, characterized by higher uncertainty. However, in the presence of environmental stochasticity (noise), purely optimistic exploration may lead to excessive probing of high-noise areas, consequently impeding exploration efficiency. Hence, in exploring noisy environments, while optimism-driven exploration serves as a foundation, prudent attention to alleviating unnecessary over-exploration in high-noise areas becomes beneficial. In this work, we propose Optimistic Value Distribution Explorer (OVD-Explorer) to achieve a noise-aware optimistic exploration for continuous control. OVD-Explorer proposes a new measurement of the policy's exploration ability considering noise in optimistic perspectives, and leverages gradient ascent to drive exploration. Practically, OVD-Explorer can be easily integrated with continuous control RL algorithms. Extensive evaluations on the MuJoCo and GridChaos tasks demonstrate the superiority of OVD-Explorer in achieving noise-aware optimistic exploration.
Jinyi Liu 0002, Zhi Wang 0001, Yan Zheng 0002, Jianye Hao, Chenjia Bai, Junjie Ye 0002, Zhen Wang 0004, Haiyin Piao
AAAI8
2024 Phasic Diversity Optimization for Population-Based Reinforcement Learning
abstract
Reviewing the previous work of diversity Reinforcement Learning, diversity is often obtained via an augmented loss function, which requires a balance between reward and diversity. Generally, diversity optimization algorithms use Multi-armed Bandits algorithms to select the coefficient in the pre-defined space. However, the dynamic distribution of reward signals for MABs or the conflict between quality and diversity limits the performance of these methods. We introduce the Phasic Diversity Optimization (PDO) algorithm, a Population-Based Training framework that separates reward and diversity training into distinct phases instead of optimizing a multi-objective function. In the auxiliary phase, agents with poor performance diversified via determinants will not replace the better agents in the archive. The decoupling of reward and diversity allows us to use an aggressive diversity optimization in the auxiliary phase without performance degradation. Furthermore, we construct a dogfight scenario for aerial agents to demonstrate the practicality of the PDO algorithm. We introduce two implementations of PDO archive and conduct tests in the newly proposed adversarial dogfight and MuJoCo simulations. The results show that our proposed algorithm achieves better performance than baselines.
Jingcheng Jiang, Haiyin Piao, Yihang Hao, Chuanlu Jiang, Ziqi Wei 0001, Xin Yang 0011
ICRA2
2024 Deep Ad-hoc Sub-Team Partition Learning for Multi-Agent Air Combat Cooperation
abstract
In the future, unmanned autonomous air combat will encounter large-scale confrontation scenarios, where agents must consider complex time-varying relationships among aircraft when making decisions. Previous works have already introduced Multi-Agent Reinforcement Learning (MARL) into air combat and succeeded in surpassing the human expert level. However, they mainly focus on small-scale air combat with low relationship complexity, e.g., 1-vs-1 or 2-vs-2. As more agents join the confrontation, existing algorithms tend to suffer significant performance degradation due to the increase in problem dimensions. In view of this, this paper proposes Deep Ad-hoc Sub-Team Partition Learning(DASPL) to address large-scale air combat problems. DASPL models multi-agent air combat as a graph to handle the complex relations and introduces an automatic partitioning mechanism to generate dynamic sub-teams, which converts the existing large-scale multi-agent air combat cooperation problem into multiple small-scale equivalence problems. Additionally, DASPL incorporates an efficient message passing method among the participating sub-teams.
Songyuan Fan, Haiyin Piao, Feng Jiang 0001, Roushu Yang
IROS2
2024 Event-intensity Stereo with Cross-modal Fusion and Contrast
abstract
For binocular stereo, traditional cameras excel in capturing fine details and texture information but are limited in terms of dynamic range and their ability to handle rapid motion. On the contrary, event cameras provide pixel-level intensity changes with low latency and a wide dynamic range, albeit at the cost of less detail in their output. It is natural to leverage the strengths of both modalities. We solve this problem by introducing a cross-modal fusion module that learns a visual representation from both sensor inputs. Additionally, we extract and compare dense event-intensity stereo pair features by contrasting “pairs of event-intensity pairs from different views and different modalities and different timestamps”. This provides the flexibility in masking hard negatives and enables networks to effectively combine event-intensity signals within a contrastive learning framework, leading to an improved matching accuracy and facilitating more accurate estimation of disparity. Experimental results validate the effectiveness of our model and the improvement of disparity estimation accuracy.
Shanglai Qu, Tianyu Meng, Haiyin Piao, Xiaopeng Wei, Xin Yang 0011
IROS5
2024 Intelligent recognition method of target tactical behavior intention in air combat based on deep learning
Zhen Yang 0011, Haiyin Piao, Shiyuan Chai, Jichuan Huang
Eng. Appl. Artif. Intell.3
2024 Learning-Based Multi-UAV Flocking Control With Limited Visual Field and Instinctive Repulsion
abstract
This article explores deep reinforcement learning (DRL) for the flocking control of unmanned aerial vehicle (UAV) swarms. The flocking control policy is trained using a centralized-learning-decentralized-execution (CTDE) paradigm, where a centralized critic network augmented with additional information about the entire UAV swarm is utilized to improve learning efficiency. Instead of learning inter-UAV collision avoidance capabilities, a repulsion function is encoded as an inner-UAV "instinct." In addition, the UAVs can obtain the states of other UAVs through onboard sensors in communication-denied environments, and the impact of varying visual fields on flocking control is analyzed. Through extensive simulations, it is shown that the proposed policy with the repulsion function and limited visual field has a success rate of 93.8% in training environments, 85.6% in environments with a high number of UAVs, 91.2% in environments with a high number of obstacles, and 82.2% in environments with dynamic obstacles. Furthermore, the results indicate that the proposed learning-based methods are more suitable than traditional methods in cluttered environments.
Chengchao Bai, Peng Yan 0005, Haiyin Piao, Wei Pan 0004, Jifeng Guo 0004
IEEE Trans. Cybern.3
2024 Discovering Expert-Level Air Combat Knowledge via Deep Excitatory-Inhibitory Factorized Reinforcement Learning
abstract
Artificial Intelligence (AI) has achieved a wide range of successes in autonomous air combat decision-making recently. Previous research demonstrated that AI-enabled air combat approaches could even acquire beyond human-level capabilities. However, there remains a lack of evidence regarding two major difficulties. First, the existing methods with fixed decision intervals are mostly devoted to solving what to act but merely pay attention to when to act, which occasionally misses optimal decision opportunities. Second, the method of an expert-crafted finite maneuver library leads to a lack of tactics diversity, which is vulnerable to an opponent equipped with new tactics. In view of this, we propose a novel Deep Reinforcement Learning (DRL) and prior knowledge hybrid autonomous air combat tactics discovering algorithm, namely deep E xcitatory-i N hibitory f ACT or I zed maneu VE r ( ENACTIVE ) learning. The algorithm consists of two key modules, i.e., ENHANCE and FACTIVE. Specifically, ENHANCE learns to adjust the air combat decision-making intervals and appropriately seize key opportunities. FACTIVE factorizes maneuvers and then jointly optimizes them with significant tactics diversity increments. Extensive experimental results reveal that the proposed method outperforms state-of-the-art algorithms with a 62% winning rate and further obtains a margin of a 2.85-fold increase in terms of global tactic space coverage. It also demonstrates that a variety of discovered air combat tactics are comparable to human experts’ knowledge.
Haiyin Piao, Shengqi Yang, Hechang Chen, Junnan Li 0008, Xuanqi Peng, Xin Yang 0011, Zhen Yang 0011, Zhixiao Sun, Yi Chang 0001
ACM Trans. Intell. Syst. Technol.1
2023 The Sufficiency of Off-Policyness and Soft Clipping: PPO Is Still Insufficient according to an Off-Policy Measure
abstract
The popular Proximal Policy Optimization (PPO) algorithm approximates the solution in a clipped policy space. Does there exist better policies outside of this space? By using a novel surrogate objective that employs the sigmoid function (which provides an interesting way of exploration), we found that the answer is "YES", and the better policies are in fact located very far from the clipped space. We show that PPO is insufficient in "off-policyness", according to an off-policy metric called DEON. Our algorithm explores in a much larger policy space than PPO, and it maximizes the Conservative Policy Iteration (CPI) objective better than PPO during training. To the best of our knowledge, all current PPO methods have the clipping operation and optimize in the clipped policy space. Our method is the first of this kind, which advances the understanding of CPI optimization and policy gradient methods. Code is available at https://github.com/raincchio/P3O.
Xing Chen 0022, Dongcui Diao, Hechang Chen, Hengshuai Yao, Haiyin Piao, Zhixiao Sun, Zhiwei Yang 0005, Randy Goebel, Bei Jiang, Yi Chang 0001
AAAI5
2023 Complex relationship graph abstraction for autonomous air combat collaboration: A learning and expert knowledge hybrid approach
Haiyin Piao, Hechang Chen, Xuanqi Peng, Songyuan Fan, Zhixiao Sun
Expert Syst. Appl.1
2023 Camouflaged Object Segmentation with Omni Perception
Haiyang Mei, Ke Xu 0010, Yunduo Zhou, Yang Wang 0106, Haiyin Piao, Xiaopeng Wei, Xin Yang 0011
Int. J. Comput. Vis.5
2023 Multi-agent air combat with two-stage graph-attention communication
Zhixiao Sun, Huahua Wu, Yandong Shi, Xiangchao Yu, Wenbin Pei, Zhen Yang 0011, Haiyin Piao, Yaqing Hou
Neural Comput. Appl.8
2023 Monocular Camera-Based Complex Obstacle Avoidance via Efficient Deep Reinforcement Learning
abstract
Deep reinforcement learning has achieved great success in laser-based collision avoidance works because the laser can sense accurate depth information without too much redundant data, which can maintain the robustness of the algorithm when it is migrated from the simulation environment to the real world. However, high-cost laser devices are not only difficult to deploy for a large scale of robots but also demonstrate unsatisfactory robustness towards the complex obstacles, including irregular obstacles, e.g., tables, chairs, and shelves, as well as complex ground and special materials. In this paper, we propose a novel monocular camera-based complex obstacle avoidance framework. Particularly, we innovatively transform the captured RGB images to pseudo-laser measurements for efficient deep reinforcement learning. Compared to the traditional laser measurement captured at a certain height that only contains one-dimensional distance information away from the neighboring obstacles, our proposed pseudo-laser measurement fuses the depth and semantic information of the captured RGB image, which makes our method effective for complex obstacles. We also design a feature extraction guidance module to weight the input pseudo-laser measurement, and the agent has more reasonable attention for the current state, which is conducive to improving the accuracy and efficiency of the obstacle avoidance policy. Besides, we adaptively add the synthesized noise to the laser measurement during the training stage to decrease the sim-to-real gap and increase the robustness of our model in the real environment. Finally, the experimental results show that our framework achieves state-of-the-art performance in several virtual and real-world scenarios.
Jianchuan Ding, Lingping Gao, Wenxi Liu, Haiyin Piao, Jia Pan 0001, Zhenjun Du, Xin Yang 0011
IEEE Trans. Circuits Syst. Video Technol.4
2023 A Geometrical Approach to Evaluate the Adversarial Robustness of Deep Neural Networks
abstract
Deep neural networks (DNNs) are widely used for computer vision tasks. However, it has been shown that deep models are vulnerable to adversarial attacks—that is, their performances drop when imperceptible perturbations are made to the original inputs, which may further degrade the following visual tasks or introduce new problems such as data and privacy security. Hence, metrics for evaluating the robustness of deep models against adversarial attacks are desired. However, previous metrics are mainly proposed for evaluating the adversarial robustness of shallow networks on the small-scale datasets. Although the Cross Lipschitz Extreme Value for nEtwork Robustness (CLEVER) metric has been proposed for large-scale datasets (e.g., the ImageNet dataset), it is computationally expensive and its performance relies on a tractable number of samples. In this article, we propose the Adversarial Converging Time Score (ACTS), an attack-dependent metric that quantifies the adversarial robustness of a DNN on a specific input. Our key observation is that local neighborhoods on a DNN’s output surface would have different shapes given different inputs. Hence, given different inputs, it requires different time for converging to an adversarial sample. Based on this geometry meaning, the ACTS measures the converging time as an adversarial robustness metric. We validate the effectiveness and generalization of the proposed ACTS metric against different adversarial attacks on the large-scale ImageNet dataset using state-of-the-art deep networks. Extensive experiments show that our ACTS metric is an efficient and effective adversarial metric over the previous CLEVER metric.
Yang Wang 0106, Bo Dong 0004, Ke Xu 0010, Haiyin Piao, Yufei Ding 0001, Xin Yang 0011
ACM Trans. Multim. Comput. Commun. Appl.4
2022 Deep Relationship Graph Reinforcement Learning for Multi-Aircraft Air Combat
abstract
Air combat Artificial Intelligence (AI) has attracted increasing attentions from aeronautics engineers and artificial intelligence researchers. However, it is often of great difficulties for the existing methods to solve the collaboration problems in multi-aircraft air combat due to their high complexity incurred by combination explosion. In view of this, we propose a Deep Relationship Graph Reinforcement Learning (DRGRL) algorithm for multi-aircraft collaboration. Specifically, DRGRL significantly simplifies the complex situation space via abstracting the original problem into a symbolic form. Besides, a novel Air Combat Relationship Graph (ACRG) is introduced to represent the learned collaboration pattern, which concentrates on the most important combat relationships for tactic decision making. Consequently, experiments are conducted in an air combat simulation environment named WUKONG. The comprehensive experimental results demonstrate that DRGRL could evidently learn some valuable collaboration patterns and achieve better combat performance than state-of-the-art air combat AI methods.
Haiyin Piao, Yaqing Hou, Zhixiao Sun, Shengqi Yang, Xuanqi Peng, Songyuan Fan
IJCNN2
2022 Distributional Reward Estimation for Effective Multi-agent Deep Reinforcement Learning
abstract
Multi-agent reinforcement learning has drawn increasing attention in practice, e.g., robotics and automatic driving, as it can explore optimal policies using samples generated by interacting with the environment. However, high reward uncertainty still remains a problem when we want to train a satisfactory model, because obtaining high-quality reward feedback is usually expensive and even infeasible. To handle this issue, previous methods mainly focus on passive reward correction. At the same time, recent active reward estimation methods have proven to be a recipe for reducing the effect of reward uncertainty. In this paper, we propose a novel Distributional Reward Estimation framework for effective Multi-Agent Reinforcement Learning (DRE-MARL). Our main idea is to design the multi-action-branch reward estimation and policy-weighted reward aggregation for stabilized training. Specifically, we design the multi-action-branch reward estimation to model reward distributions on all action branches. Then we utilize reward aggregation to obtain stable updating signals during training. Our intuition is that consideration of all possible consequences of actions could be useful for learning policies. The superiority of the DRE-MARL is demonstrated using benchmark multi-agent scenarios, compared with the SOTA baselines in terms of both effectiveness and robustness.
Jifeng Hu, Yanchao Sun, Hechang Chen, Sili Huang, Haiyin Piao, Yi Chang 0001, Lichao Sun 0001
NeurIPS5
2022 A Two-Stage Attentive Network for Single Image Super-Resolution
abstract
Recently, deep convolutional neural networks (CNNs) have been widely explored in single image super-resolution (SISR) and contribute remarkable progress. However, most of the existing CNNs-based SISR methods do not adequately explore contextual information in the feature extraction stage and pay little attention to the final high-resolution (HR) image reconstruction step, hence hindering the desired SR performance. To address the above two issues, in this paper, we propose a two-stage attentive network (TSAN) for accurate SISR in a coarse-to-fine manner. Specifically, we design a novel multi-context attentive block (MCAB) to make the network focus on more informative contextual features. Moreover, we present an essential refined attention block (RAB) which could explore useful cues in HR space for reconstructing fine-detailed HR image. Extensive evaluations on four benchmark datasets demonstrate the efficacy of our proposed TSAN in terms of quantitative metrics and visual effects. Code is available athttps://github.com/Jee-King/TSAN.
Jiqing Zhang, Chengjiang Long, Yuxin Wang 0001, Haiyin Piao, Haiyang Mei, Xin Yang 0011
IEEE Trans. Circuits Syst. Video Technol.4
2022 Learning Smooth Motion Planning for Intelligent Aerial Transportation Vehicles by Stable Auxiliary Gradient
abstract
Deep Reinforcement Learning (DRL) has been widely attempted for solving real-time intelligent aerial transportation vehicle motion planning tasks recently. When interacting with environment, DRL-driven aerial vehicles inevitably switch the steering actions in high frequency during both exploration and execution phase, resulting in the well known flight trajectory oscillation issue, which makes flight dynamics unstable, and even endangers flight safety in serious cases. Unfortunately, there is hardly any literature about achieving flight trajectory smoothness in DRL-based motion planning. In view of this, we originally formalize the practical flight trajectory smoothen problem as a three-level Nested pArameterized Smooth Trajectory Optimization (NASTO) form. On this basis, a novel Stable Auxiliary Gradient (SAG) algorithm is proposed, which significantly smoothens the DRL-generated flight motions by constructing two independent optimization aspects: the major gradient, and the stable auxiliary gradient. Experimental result reveals that the proposed SAG algorithm outperforms baseline DRL-based intelligent aerial transportation vehicle motion planning algorithms in terms of both learning efficiency and flight motion smoothness.
Haiyin Piao, Li Mo 0001, Xin Yang 0011, Zhixiao Sun, Zhen Yang 0011
IEEE Trans. Intell. Transp. Syst.1
2021 A Vision-based Irregular Obstacle Avoidance Framework via Deep Reinforcement Learning
abstract
Deep reinforcement learning has achieved great success in laser-based collision avoidance work because the laser can sense accurate depth information without too much redundant data, which can maintain the robustness of the algorithm when it is migrated from the simulation environment to the real world. However, high-cost laser devices are not only difficult to apply on a large scale but also have poor robustness to irregular objects, e.g., tables, chairs, shelves, etc. In this paper, we propose a vision-based collision avoidance framework to solve the challenging problem. Our method attempts to estimate the depth and incorporate the semantic information from RGB data to obtain a new form of data, pseudo-laser data, which combines the advantages of visual information and laser information. Compared to traditional laser data that only contains the one-dimensional distance information captured at a certain height, our proposed pseudo-laser data encodes the depth information and semantic information within the image, which makes our method more effective for irregular obstacles. Besides, we adaptively add noise to the laser data during the training stage to increase the robustness of our model in the real world, due to the estimated depth information is not accurate. Experimental results show that our framework achieves state-of-the-art performance in several unseen virtual and real-world scenarios.
Lingping Gao, Jianchuan Ding, Wenxi Liu, Haiyin Piao, Yuxin Wang 0001, Xin Yang 0011
IROS4
2021 Coordinated Proximal Policy Optimization
abstract
We present Coordinated Proximal Policy Optimization (CoPPO), an algorithm that extends the original Proximal Policy Optimization (PPO) to the multi-agent setting. The key idea lies in the coordinated adaptation of step size during the policy update process among multiple agents. We prove the monotonicity of policy improvement when optimizing a theoretically-grounded joint objective, and derive a simplified optimization objective based on a set of approximations. We then interpret that such an objective in CoPPO can achieve dynamic credit assignment among agents, thereby alleviating the high variance issue during the concurrent update of agent policies. Finally, we demonstrate that CoPPO outperforms several strong baselines and is competitive with the latest multi-agent PPO method (i.e. MAPPO) under typical multi-agent settings, including cooperative matrix games and the StarCraft II micromanagement tasks.
Zifan Wu, Chao Yu 0004, Deheng Ye, Junge Zhang, Haiyin Piao, Hankui Zhuo
NeurIPS5
2021 MBKD: Acceleration structure designed for moving primitives
Haiyin Piao, Pengyuan Du, Letian Yu, Yuxin Wang 0001, Xin Yang 0011
Comput. Graph.1
2021 Multi-agent hierarchical policy gradient for Air Combat Tactics emergence via self-play
Zhixiao Sun, Haiyin Piao, Zhen Yang 0011, Guang Zhan, Guanglei Meng, Hechang Chen, Xing Chen 0022, Bohao Qu, Yuanjie Lu
Eng. Appl. Artif. Intell.2
2020 Beyond-Visual-Range Air Combat Tactics Auto-Generation by Reinforcement Learning
abstract
For quite a long time, effective Beyond-Visual-Range (BVR) air combat tactics can only be discovered by human pilots in the actual combat process. However, due to the lack of actual combat opportunities, making new air combat tactics innovation was generally considered quite difficult. To address this challenge, we first introduced a solely end-to-end Reinforcement Learning (RL) approach for training competitive air combat agents with adversarial self-play from scratch in a high fidelity air combat simulation environment during training. Furthermore, a Key Air Combat Event Reward Shaping (KAERS) mechanism was proposed to provide sparse but objective shaped rewards beyond episodic win/lose signal to accelerate the initial machine learning process. Experimental results showed that multiple valuable air combat tactical behaviors emerged progressively. We hope this study could be extended to the future of air combat machine intelligence research.
Haiyin Piao, Zhixiao Sun, Guanglei Meng, Hechang Chen, Bohao Qu, Kuijun Lang, Shengqi Yang, Xuanqi Peng
IJCNN1
2020 MA-TREX: Mutli-agent Trajectory-Ranked Reward Extrapolation via Inverse Reinforcement Learning
Sili Huang, Bo Yang 0002, Hechang Chen, Haiyin Piao, Zhixiao Sun, Yi Chang 0001
KSEM (2)4
2020 Combining sequence and network information to enhance protein-protein interaction prediction
abstract
BACKGROUND: Protein-protein interactions (PPIs) are of great importance in cellular systems of organisms, since they are the basis of cellular structure and function and many essential cellular processes are related to that. Most proteins perform their functions by interacting with other proteins, so predicting PPIs accurately is crucial for understanding cell physiology. RESULTS: Recently, graph convolutional networks (GCNs) have been proposed to capture the graph structure information and generate representations for nodes in the graph. In our paper, we use GCNs to learn the position information of proteins in the PPIs networks graph, which can reflect the properties of proteins to some extent. Combining amino acid sequence information and position information makes a stronger representation for protein, which improves the accuracy of PPIs prediction. CONCLUSION: In previous research methods, most of them only used protein amino acid sequence as input information to make predictions, without considering the structural information of PPIs networks graph. We first time combine amino acid sequence information and position information to make representations for proteins. The experimental results indicate that our method has strong competitiveness compared with several sequence-based methods.
Leilei Liu, Xianglei Zhu, Yi Ma 0005, Haiyin Piao, Yaodong Yang 0002, Xiaotian Hao, Jiajie Peng
BMC Bioinform.4