Zhenyu Zhang 0013

dblp:01/1844-13 · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0002-9470-7132ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Spatial-temporal clues based interactive message aggregation for multi-agent collaboration
Shaorong Xie, Xiangfeng Luo, Zhenyu Zhang 0013, Hang Yu 0006
Frontiers Comput. Sci.4
2026 TAPPS: Learning diverse policies via task-adaptive partial parameter sharing in heterogeneous multi-agent reinforcement learning
Zhenyu Zhang 0013, Xiangfeng Luo, Shaorong Xie
Knowl. Based Syst.4
2026 Dilated memory in hierarchical reinforcement learning for long-horizontal task
Zhenyu Zhang 0013, Shaorong Xie, Xiangfeng Luo
Neural Networks1
2025 STSR-Nav: Spatial-Temporal Scene Representation for Navigation Policy Learning Based on UAV-UGV Collaborative Perception
abstract
In UAV-UGV collaboration, UGVs enhance their perception capabilities by utilizing shared observation data from UAVs to make effective decisions. However, learning navigation policies directly from multi-perspective observation data is challenging due to spatial inconsistencies and redundancy. Existing approaches either rely on low-level scene representations, which may reduce efficiency, or use mid-level semantic representations that lack essential geometric and temporal details, resulting in sub-optimal decisions. This paper presents STSR-Nav, a novel framework for navigation policy learning based on UAV-UGV collaborative perception. The framework first encodes semantic information and geometric relationships from UAV images using a Bird'View (BEV) scene graph. A learnable graph neural network then integrates the BEV scene graph with local perception to infer spatial features. Temporal learning is applied to these features via a memory network, generating spatial-temporal features that are embedded into the UGV's local perception space using a multi-head attention mechanism. Finally, the fused spatial-temporal features provide a high-level scene representation for navigation policy learning. Unity simulations demonstrate the framework's improved efficiency and generalization capabilities.
Shaorong Xie, Xiangfeng Luo, Xinzhi Wang 0001, Zhenyu Zhang 0013, Tao Wang 0072
CSCWD5
2025 Spatial-temporal intention representation with multi-agent reinforcement learning for unmanned surface vehicles strategies learning in asset guarding task
Yang Li 0151, Shaorong Xie, Hang Yu 0006, Zhenyu Zhang 0013, Xiangfeng Luo
Eng. Appl. Artif. Intell.5
2025 Efficient offline-to-online reinforcement learning with pre-reduced out-of-distribution Q-values
Tao Wang 0072, Xiangfeng Luo, Zhenyu Zhang 0013, Shaorong Xie
Neural Networks3
2024 Knowledge-guided communication preference learning model for multi-agent cooperation
Hang Yu 0006, Zhenyu Zhang 0013, Yang Li 0151, Shaorong Xie, Xiangfeng Luo
Inf. Sci.5
2024 Concept Drift Adaptation by Exploiting Drift Type
abstract
Concept drift is a phenomenon where the distribution of data streams changes over time. When this happens, model predictions become less accurate. Hence, models built in the past need to be re-learned for the current data. Two design questions need to be addressed in designing a strategy to re-learn models: which type of concept drift has occurred, and how to utilize the drift type to improve re-learning performance. Existing drift detection methods are often good at determining when drift has occurred. However, few retrieve information about how the drift came to be present in the stream. Hence, determining the impact of the type of drift on adaptation is difficult. Filling this gap, we designed a framework based on a lazy strategy called Type-Driven Lazy Drift Adaptor (Type-LDA). Type-LDA first retrieves information about both how and when a drift has occurred, then it uses this information to re-learn the new model. To identify the type of drift, a drift type identifier is pre-trained on synthetic data of known drift types. Furthermore, a drift point locator locates the optimal point of drift via a sharing loss. Hence, Type-LDA can select the optimal point, according to the drift type, to re-learn the new model. Experiments validate Type-LDA on both synthetic data and real-world data, and the results show that accurately identifying drift type can improve adaptation accuracy.
Hang Yu 0006, Zhenyu Zhang 0013, Xiangfeng Luo, Shaorong Xie
ACM Trans. Knowl. Discov. Data3
2023 Recurrent prediction model for partially observable MDPs
Shaorong Xie, Zhenyu Zhang 0013, Hang Yu 0006, Xiangfeng Luo
Inf. Sci.2
2022 Learning from Suboptimal Demonstration via Trajectory-Ranked Adversarial Imitation
abstract
Robots trained by Imitation Learning(IL) are used in many tasks(e.g., autonomous vehicle manipulation). Generative Adversarial Imitation Learning (GAIL) assumes that the demonstration set used for training is of high quality. However, such demonstrations are difficult and expensive to obtain. GAIL-related methods fail to learn effective strategies if non-high quality demonstrations are used because the performance of agents trained by this method is limited by the demonstrator's operations. Our idea is to enable the agent to learn strategy with better performance than the demonstrator from a suboptimal demonstration set, which contains non-high quality demonstrations that are easier to obtain. Inspired by this, we propose the Trajectory-Ranked Adversarial Imitation Learning (TRAIL) method. First, for demonstration set processing, we introduce a ranking process and define the concept of Performance Relative Advantage of suboptimal demonstrations to specify the ranking order. Second, for model training, we reconstruct the objective function of GAIL and use an experience replay buffer, enabling the agent to learn implicit features and ranking information from the ranked suboptimal demonstration set and possess the ability to outperform the demonstrator. Experiments show that in Mujoco's tasks, our method can learn from a suboptimal demonstration set and can achieve better performance than baseline methods.
Shaorong Xie, Hang Yu 0006, Xiangfeng Luo, Zhenyu Zhang 0013
ICTAI6
2022 Offline Reinforcement Learning via Policy Regularization and Ensemble Q-Functions
abstract
Offline reinforcement learning aims to learn effective policies from a fixed set of data collected in advance and without further interaction with the environment during learning. This setting will promote the applications of reinforcement learning in the real world, in which interaction is costly or dangerous. However, existing off-policy algorithms can fail in offline settings due to the distributional shift between the learned policy and the policy that collected the dataset. To solve the problem, we develop a lightweight and effective algorithm, policy regularization with behavior model (PRBM). Firstly, PRBM trains a behavior model as the regularization term during policy optimization to avoid choosing out-of-distribution (OOD) actions. Secondly, to avoid the overestimation of OOD actions, PRBM trains multiple Q- functions and uses their min-max mixture to compute Q-values. Our experiments on different datasets from various continuous control tasks demonstrate that PRBM outperforms most baselines (especially on medium quality datasets) and requires only half of their training time.
Tao Wang 0072, Shaorong Xie, Mingke Gao, Zhenyu Zhang 0013, Hang Yu 0006
ICTAI5
2022 USVs-Sim: A general simulation platform for unmanned surface vessels autonomous learning
abstract
Abstract Unmanned surface vessels (USVs) have been fully used in the civilian and military fields in recent years, which dramatically expands protective capability and detection range. However, the marine environment's complexity and variability make that verification of various advanced USVs control algorithms face high costs and high risks. In this article, we present USVs‐Sim, a novel high‐fidelity general simulation platform for USVs autonomous navigation data generation and control strategy testing. USVs‐Sim is a collection of high‐level extensible modules that allows the rapid development and testing of USVs configurations and facilitates the construction of complex ocean scenarios. USVs‐Sim supports the steering or thrusting limits of USVs, as well as unique dynamics profiles. The platform can specify specific USVs sensor systems and change the time of day and weather conditions to generate robust data. USVs‐Sim facilitates training of deep‐learning algorithms by enabling data export from USVs sensors, including vision data, lidar, relative positions of ocean targets. Therefore, USVs‐Sim allows for the rapid prototyping, development, and testing of USVs autonomous control algorithms in a complex marine environment. In this article, we detail the general simulation platform and testing several representative USVs intelligent control algorithms on the platform.
Wei Wang 0296, Yang Li 0151, Zhenyu Zhang 0013, Xiangfeng Luo, Shaorong Xie
Concurr. Comput. Pract. Exp.4
2022 ET-HF: A novel information sharing model to improve multi-agent cooperation
Shaorong Xie, Hang Yu 0006, Yang Li 0151, Zhenyu Zhang 0013, Xiangfeng Luo
Knowl. Based Syst.5
2021 Learning adversarial policy in multiple scenes environment via multi-agent reinforcement learning
abstract
Learning adversarial policy aims to learn behavioural strategies for agents with different goals, is one of the most significant tasks in multi-agent systems. Multi-agent reinforcement learning (MARL), as a state-of-the-art learning-based model, employs centralised or decentralised control methods to learn behavioural strategies by interacting with environments. It suffers from instability and slowness in the training process. Considering that parallel simulation or computation is an effective way to improve training performance, we propose a novel MARL method called Multiple scenes multi-agent proximal Policy Optimisation (MPO) in this paper. In MPO, we first simulate multiple parallel scenes in the training environment. Multiple policies control different agents in the same scene, and each policy also controls several identical agents from multiple scenes. Then, we expand proximal policy optimisation (PPO) with an improved actor-critic network, ensuring the stability of training in multi-agent tasks. The actor network only uses local information for decision making, and the critic network uses global information for training. Finally, effective training trajectories are computed with two criteria from multiple parallel scenes rather than single to accelerate the learning process. We evaluate our approach in two simulated 3D environments, one of which is Unity's official open-source soccer game, and the other is unmanned surface vehicles (USVs) built by Unity. Experiments demonstrate that MPO converges more stable and faster than benchmark methods in model training, and demonstrates excellent adversarial policy compared with benchmark models.
Yang Li 0151, Xinzhi Wang 0001, Wei Wang 0296, Zhenyu Zhang 0013, Jianshu Wang, Xiangfeng Luo, Shaorong Xie
Connect. Sci.4
2020 SEM: Adaptive Staged Experience Access Mechanism for Reinforcement Learning
abstract
Experience memory, which stores and replays experience from the past for off-policy reinforcement learning (RL) agents, has become an essential component of cutting-edge algorithms in recent years. In most previous work, experience memory emphasizes recently acquired experiences while neglects long-ago ones, which causes high variance and catastrophic forgetting in continual learning. To solve the problem, we introduce the Staged Experience Mechanism(SEM) - a novel management mechanism of experience memory. SEM adaptively regulates the proportion of experiences based on the current learning stages, enabling agents not only to learn from new experiences but also from very old ones given limited memory capacity. In the experiment, SEM is coupled with two state-of-the-art off-policy algorithms compared with classic experience replay mechanisms on the suite of MuJoCo tasks. The result shows that RL with SEM is more stable and less variable than its counterparts in most tested continuous control environments, which reduces the training variance by 5% ∼ 104% while ensuring optimal scoring performance.
Jianshu Wang, Xinzhi Wang 0001, Xiangfeng Luo, Zhenyu Zhang 0013, Wei Wang 0296, Yang Li 0151
ICTAI4
2019 Proximal Policy Optimization with Mixed Distributed Training
abstract
Instability and slowness are two main problems in deep reinforcement learning. Even if proximal policy optimization (PPO) is the state of the art, it still suffers from these two problems. We introduce an improved algorithm based on proximal policy optimization, mixed distributed proximal policy optimization (MDPPO), and show that it can accelerate and stabilize the training process. In our algorithm, multiple different policies train simultaneously and each of them controls several identical agents that interact with environments. Actions are sampled by each policy separately as usual, but the trajectories for the training process are collected from all agents, instead of only one policy. We find that if we choose some auxiliary trajectories elaborately to train policies, the algorithm will be more stable and quicker to converge especially in the environments with sparse rewards.
Zhenyu Zhang 0013, Xiangfeng Luo, Tong Liu 0001, Shaorong Xie, Jianshu Wang, Wei Wang 0296, Yang Li 0151, Yan Peng 0001
ICTAI1