Ziyuan Zhou 0005

dblp:193/7974-5 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-2649-8666ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Safe Fuzzy-CBF: A Monitoring Approach for Deep Reinforcement Learning Navigation Agents
abstract
In safety-critical reinforcement learning tasks such as robot navigation, autonomous driving, and industrial control, real-time monitoring of agent behavior is paramount to prevent unsafe actions from compromising system reliability. We introduce Safe Fuzzy Control Barrier Functions (Safe Fuzzy-CBF), a monitoring framework that continuously evaluates agent risk using a smooth control barrier function learned from data via Takagi-Sugeno fuzzy rules, rather than relying on a predefined model-based barrier. In the offline phase, we extract fuzzy rules from trajectory data and construct a continuously differentiable barrier that unifies state, action, and cost into a single risk metric. During online execution, we estimate the barrier’s time derivative using finite differences, apply Lipschitz-based corrections, and when the barrier condition is violated, we solve a small quadratic program in the normalized control space to compute a safe corrective action. Evaluated on three Safety-Gymnasium navigation tasks and the MetaDrive autonomous driving environment, our framework detects emerging safety violations in real time and enforces runtime safety without degrading the underlying RL policy’s performance, achieving up to an 86.8% reduction in accumulated cost.
Guanjun Liu, Ziyuan Zhou 0005, GaiYun Liu
IEEE Trans Autom. Sci. Eng.3
2025 Robust Multi-Agent Reinforcement Learning with Stochastic Adversary
abstract
The performance of models trained by Multi-Agent Reinforcement Learning (MARL) is sensitive to perturbations in observations, lowering their trustworthiness in complex environments. Adversarial training is a valuable approach to enhance their performance robustness. However, existing methods often overfit to adversarial perturbations of observations and fail to incorporate prior information about the policy adopted by their protagonist agent, i.e., the primary one being trained. To address this important issue, this paper introduces Adversarial Training with Stochastic Adversary (ATSA), where the proposed adversary is trained online alongside the protagonist agent. The former consists of Stochastic Director (SDor) and SDor-guided generaTor (STor). SDor performs policy perturbations by minimizing the expected team reward of protagonists and maximizing the entropy of its policy, while STor generates adversarial perturbations of observations by following SDor's guidance. We prove that SDor's soft policy converges to a global optimum according to factorized maximum-entropy MARL and leads to the optimal adversary. This paper also introduces an SDor-STor loss function to quantify the difference between a) perturbations in the agent's policy and b) those advised by SDor. We evaluate our ATSA on StarCraft II tasks and autonomous driving scenarios, demonstrating that a) it is robust against diverse perturbations of observations while maintaining outstanding performance in perturbation-free environments, and b) it outperforms the state-of-the-art methods.
Ziyuan Zhou 0005, Guanjun Liu, MengChu Zhou, Weiran Guo
ICML1
2025 PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning
abstract
Reinforcement Learning (RL) is widely used in tasks where agents interact with an environment to maximize rewards. Building on this foundation, Safe Reinforcement Learning (Safe RL) incorporates a cost metric alongside the reward metric, ensuring that agents adhere to safety constraints during decision-making. In this paper, we identify that Safe RL is vulnerable to backdoor attacks, which can manipulate agents into performing unsafe actions. First, we introduce the relevant concepts and evaluation metrics for backdoor attacks in Safe RL. It is the first attack framework in the Safe RL field that involves both Positive and Negative Action sample (PNAct) is to implant backdoors, where positive action samples provide reference actions and negative action samples indicate actions to be avoided. We theoretically point out the properties of PNAct and design an attack algorithm. Finally, we conduct experiments to evaluate the effectiveness of our proposed backdoor attack framework, evaluating it with the established metrics. This paper highlights the potential risks associated with Safe RL and underscores the feasibility of such attacks. Our code and supplementary material are available at https://github.com/azure-123/PNAct.
Weiran Guo, Guanjun Liu, Ziyuan Zhou 0005
IJCAI3
2025 Robust Training in Multiagent Deep Reinforcement Learning Against Optimal Adversary
abstract
Industry 5.0 enhances manufacturing ability through efficient human-machine interaction, combining human resources and robots to complete tasks more accurately and effectively. Artificial intelligence (AI) plays an essential role in Industry 5.0. As a branch in AI, multiagent deep reinforcement learning (MADRL) attracts vast attention in both academia and industry. However, there is a gap between virtual and physical environments in terms of howcleanan observed state is. In addition, state adversarial attacks can seriously impact the performance of MADRL. Hence, how to improve the robustness of MADRL algorithms is an important research topic. In this article, we propose an optimal policy-based state adversary attack method that would make the MADRL algorithm more robust when it is applied in the training process of agents. Two case studies related to Industry 5.0 and a general case study are presented in which robustness training against the optimal adversarial attack is tested. The MADRL algorithms involved in the experiments include centralized training and decentralized execution (CTDE) framework and shared experience actor-critic (SEAC) to demonstrate the universality of our method.
Weiran Guo, Guanjun Liu, Ziyuan Zhou 0005, Jiacun Wang 0001, Ying Tang 0001
IEEE Trans. Syst. Man Cybern. Syst.3
2024 Enhancing the robustness of QMIX against state-adversarial attacks
Weiran Guo, Guanjun Liu, Ziyuan Zhou 0005, Jiacun Wang 0001
Neurocomputing3
2024 A Robust Mean-Field Actor-Critic Reinforcement Learning Against Adversarial Perturbations on Agent States
abstract
Multiagent deep reinforcement learning (DRL) makes optimal decisions dependent on system states observed by agents, but any uncertainty on the observations may mislead agents to take wrong actions. The mean-field actor-critic (MFAC) reinforcement learning is well-known in the multiagent field since it can effectively handle a scalability problem. However, it is sensitive to state perturbations that can significantly degrade the team rewards. This work proposes a Robust MFAC (RoMFAC) reinforcement learning that has two innovations: 1) a new objective function of training actors, composed of a policy gradient function that is related to the expected cumulative discount reward on sampled clean states and an action loss function that represents the difference between actions taken on clean and adversarial states and 2) a repetitive regularization of the action loss, ensuring the trained actors to obtain excellent performance. Furthermore, this work proposes a game model named a state-adversarial stochastic game (SASG). Despite the Nash equilibrium of SASG may not exist, adversarial perturbations to states in the RoMFAC are proven to be defensible based on SASG. Experimental results show that RoMFAC is robust against adversarial perturbations while maintaining its competitive performance in environments without perturbations.
Ziyuan Zhou 0005, Guanjun Liu, MengChu Zhou
IEEE Trans. Neural Networks Learn. Syst.1
2024 Adversarial Attacks on Multiagent Deep Reinforcement Learning Models in Continuous Action Space
abstract
Multiagent deep reinforcement learning (MADRL) has been recently applied in many fields, including industry 5.0, but it is sensitive to adversarial attacks. Although adversarial attacks can be detrimental, they are crucial for testing and assisting in enhancing the robustness of models. Existing attacks on MADRL-based models are not sufficient since these attacks involve fixed perturbed agents, without taking into account cases where perturbed agents change. In this article, we present a novel adversarial attack framework. In this framework, we define critical agents that change over time, i.e., when they are perturbed a little, the whole multiagent system is perturbed greatly. Then, we identify critical agents through their worst-case joint actions. In this identifying process, we use gradient information, differential evolution, and SARSA to deal with the challenge caused by changes in the perturbed agents and to compute the worst-case joint actions. After identifying them, we use the target attack method to perturb them. We apply our method to attack the models trained by two state-of-the-art MADRL algorithms under three environments, including two industry-related ones. The experimental results demonstrate our method has a stronger perturbing ability than the existing methods.
Ziyuan Zhou 0005, Guanjun Liu, Weiran Guo, MengChu Zhou
IEEE Trans. Syst. Man Cybern. Syst.1
2023 Robustness Testing for Multi-Agent Reinforcement Learning: State Perturbations on Critical Agents
abstract
Multi-agent reinforcement learning (MARL) has been widely applied in many fields, such as smart traffic and unmanned aerial vehicles. However, most MARL algorithms are vulnerable to adversarial perturbations on agent states. Robustness testing for a trained model is an essential step for confirming the trustworthiness of the model against unexpected perturbations. This work proposes a novel Robustness Testing framework for MARL that attacks states of Critical Agents (RTCA). The RTCA has two innovations: 1) a differential evolution (DE) based method to select critical agents as victims and to advise the worst-case joint actions on them, and 2) a team cooperation policy evaluation method employed as the objective function for the optimization of DE. Then, adversarial state perturbations of the critical agents are generated based on the worst-case joint actions. This is the first robustness testing framework with varying victim agents. RTCA demonstrates outstanding performance in terms of the number of victim agents and destroying cooperation policies.
Ziyuan Zhou 0005, Guanjun Liu
ECAI1
2023 MARL Sim2real Transfer: Merging Physical Reality With Digital Virtuality in Metaverse
abstract
Metaverse is an artificial virtual world mapped from and interacting with the real world. In metaverse, digital entities coexist with their physical counterparts. Powered by deep learning, metaverse is inevitably becoming more intelligent in the interactions between reality and virtuality. However, it is confronted with a nontrivial problem known as sim2real transfer when deep learning techniques try to bridge the reality gap between the physical world and simulations. In this article, we use multiagent deep reinforcement learning (MARL) to implement collective intelligence for digital entities as well as their physical counterparts. To model the immersive environments in metaverse, we define a nonstationary variant of Markov games and propose a recurrent MARL solution to it. Based on the solution, MARL sim2real transfer that bridges real and virtual multiple unmanned aerial vehicle (multi-UAV) systems is successfully conducted by employing recurrent multiagent deep deterministic policy gradient (R-MADDPG) with the domain randomization technique. Additionally, we use perception-control modularization to improve the generalization performance of MARL policies and make training more efficient.
Guanjun Liu, Kaiwen Zhang 0010, Ziyuan Zhou 0005, Jiacun Wang 0001
IEEE Trans. Syst. Man Cybern. Syst.4