Yizhou Shen

dblp:318/8454 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0002-3769-7193ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 11 · 2 first-author · 11 since 2021Security and privacy · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Privacy-Aware DRL for Differential Games-Assisted Malware Defense in Edge Intelligence-Enabled Social IoT
abstract
The edge intelligence-enabled Social Internet of Things (SIoT) faces severe security threats from stealthy malware propagation, while existing defenses struggle to model complex behaviors or provide real-time and privacy-aware responses. Herein, we propose a comprehensive malware defense framework integrating a five-state propagation model, continuous-time differential games, and a privacy-aware reinforcement learning algorithm named PP-D3QN (Privacy-Preserving Dueling Double Deep Q Network). The malware propagation model includes susceptible, infectious, patched, quarantined, and removed states, accurately representing centralized and cooperative patching as well as quarantine detection mechanisms. Leveraging differential games, optimal defense strategies are theoretically derived by solving the Hamilton–Jacobi– Bellman equation, dynamically balancing infection risk, patching benefits, and quarantine costs. The PP-D3QN algorithm employs prioritized experience replay with strict control over private data sampling and Gaussian noise perturbation to ensure differential privacy, while learning effective defense strategies through practical interaction with dynamic edge intelligence-enabled SIoT systems. Extensive simulations demonstrate that the proposed method significantly improves malware suppression speed and SIoT nodes recovery rates, showcasing strong theoretical and practical value. This work offers a rigorous and applicable solution for dynamic malware defense under privacypreserving constraints in edge intelligence-enabled SIoT systems.
Shigen Shen, Yizhou Shen, Jingnan Dong, Tian Wang 0001, Ruidong Li 0001
IEEE Trans. Netw. Serv. Manag.3
2025 Deep-Reinforcement-Learning-Based Botnet Propagation Control in the Social Internet of Things
abstract
The rapid development of the social Internet of Things (IoT) enhances interconnectivity but also raises significant network security challenges, particularly from botnet attacks that disrupt system stability. Addressing this issue requires effective strategies to control botnet propagation in social IoT environments. This study develops a social IoT botnet propagation model incorporating social factors to analyze their influences on its propagation dynamics. Based on this, a social IoT botnet propagation control framework is constructed, formulating an optimization problem using Markov games. To solve the optimization problem, we propose SD-DRQN (Social-Dynamics Deep Recurrent Q-Network), a novel deep reinforcement learning algorithm that integrates Long Short-Term Memory (LSTM) layers to improve learning in dynamic social IoT environments. Experimental results validate the performance of the proposed SD-DRQN across various social IoT scenarios, including complex real-world topologies. The algorithm demonstrates faster convergence, superior generalization, and practical applicability, making it an effective solution for botnet propagation control in real-world social IoT deployments.
Shigen Shen, Xuanbin Hao, Yizhou Shen, Huibin Xu, Jingnan Dong, Zhaoxi Fang, Zongda Wu
IEEE Internet Things J.3
2025 Mitigating Malware Propagation in Social Internet of Things Using an Exact Markov-Chain-Based Epidemic Method
abstract
In the Social Internet of Things (SIoT) environment, malware propagation is attracting more and more attention due to increasing damages. Markov chain models have been used to predict epidemic behavior qualitatively and quantitatively, but most of them model random propagation as a basic multiplicative factor. In this article, we propose an epidemic model Susceptible-Infected without command-Infected with command$(SII^{\prime })$, and derive an exact Markov chain for SIoT malware propagation. We also employ a Markov chain for an SIoT malware mitigation system that groups random devices alongside those with detected infections during the malware eradication process. This mitigation mechanism operates at the network scale, addressing the risks associated with large-scale SIoT deployments through a strategic, yet assertive, approach of widespread disconnections. Such a system effectively drives down the basic reproduction number to less than 1, preventing malware from gaining dominance over the network—all accomplished without modifying the recovery rate. We conducted experimental simulations of the proposed model’s dynamic predictions, and the experimental results show that the use of an exact Markov chain model better matches the benchmark results of our proposed model and also verifies the different effects of group-based mitigation in different SIoT contexts.
Hong Zhang 0046, Yizhou Shen, Huibin Xu, Shigen Shen, Ruidong Li 0001
IEEE Internet Things J.3
2025 RT-A3C: Real-time Asynchronous Advantage Actor-Critic for optimally defending malicious attacks in edge-enabled Industrial Internet of Things
Wenyi Zhu, Yizhou Shen, Xiao Zhi Gao 0001, Shigen Shen
J. Inf. Secur. Appl.4
2025 Joint Mean-Field Game and Multiagent Asynchronous Advantage Actor-Critic for Edge Intelligence-Based IoT Malware Propagation Defense
abstract
Defending Edge intelligence-based Internet of Things (EIoT) systems by controlling malware propagation has become a critical issue. Herein, we meet new challenges brought by multiple attackers and defenders for effective mitigation of malware propagation in EIoT. To explore the dynamic changes of malware propagation, a state transition diagram of IoT nodes is proposed to describe the mutual changes of five states: infected, active, dormant, isolated, and hardened, and differential equations for each state are established. We then build a mean-field game model representing the interactions between multiattackers and multidefenders in the EIoT malware propagation defense environment. We further convert the problem of solving the game into an MDP (Markov Decision Process) and propose a distributed algorithm called MFGA3C (Mean-Field Game-based Asynchronous Advantage Actor-Critic) that combines mean-field game and multiagent asynchronous advantage actor-critic to learn the optimal defense policy. Finally, we compare our algorithm MFGA3C with other benchmark algorithms in the EIoT malware propagation countermeasure environment. Experimental results show that our MFGA3C algorithm dominates the average reward, total reward, and the number of successful defenses under three typical EIoT-environmental conditions, which indicates that MFGA3C faster and more robustly learns the optimal malware propagation defense policy, contributing to protecting EIoT systems.
Shigen Shen, Chenpeng Cai, Yizhou Shen, Wenlong Ke, Shui Yu 0001
IEEE Trans. Dependable Secur. Comput.3
2025 RMAAC: Joint Markov Games and Robust Multiagent Actor-Critic for Explainable Malware Defense in Social IoT
abstract
The end-edge-cloud-based Social Internet of Things (SIoT) faces increasing threats from malware. To address these challenges, we propose an explainable novel malware defense framework that integrates Markov games with multi-agent deep reinforcement learning under an end-edge-cloud-based SIoT collaborative architecture. The framework models the interactions between malicious SIoT nodes and edge devices as a multi-agent game problem, incorporating multi-layer defense mechanisms to achieve precise descriptions of attack-defense behaviors. By combining Markov games with the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm, we develop the Robust Multiagent Actor-Critic (RMAAC) algorithm, which enables adaptive strategy optimization. To enhance system interpretability, we introduce SHapley Additive exPlanations (SHAP) value analysis into the defense decision-making process, providing transparent insights into feature contributions and decision rationales. Extensive experimental results demonstrate that the proposed RMAAC algorithm significantly outperforms existing methods, including MADDPG and Minimax Multi-Agent Deep Deterministic Policy Gradient (M3DDPG), in multiple performance metrics such as episode reward, hacking success rate, and cumulative defense number. Through systematic parameter optimization, including batch size and agent interaction speed, the framework provides an effective and sustainable solution for explainable malware defense in SIoT environments.
Shigen Shen, Yizhou Shen, Jingnan Dong, Jie Wu 0001
IEEE Trans. Dependable Secur. Comput.3
2025 Privacy Preservation Strategies for Malware-Infected Edge Intelligence Systems: A Bayesian Stochastic Game-Based Approach
abstract
Malware in the Internet of Things (IoT) is prone to contaminating various IoT end-points through network communication and information transfer, leading to surreptitious privacy leakage and data theft. The existing privacy-preserving approaches including data masking, anonymization, and differential privacy always lack the consideration of strategic interactions among rational agents. Inspired by Bayesian games, we model incomplete stochastic games between IoT end-points and edge nodes in edge intelligence (EI)-enabled IoT systems to conduct probability analysis for predicting and defending privacy leakage caused by malware infection. It is notable that the posterior probability is defined based on the Bayes’ rule to reflect the statistical inference of incomplete privacy leakage information. Such a method can intrinsically characterize the actual situations of IoT end-points. Further, we propose a novel privacy preservation optimization approach named Bayesian advantage actor critic (BA2C) for the practical implementation of optimization decision in EI-enabled IoT privacy-preserving systems. Eventually, we conduct experimental simulations to understand the most effective parameters in decision-making among the successful detection rate, successful infection rate, and false alarm rate. We also compare traditional algorithms and validate the efficacy of the proposed approach.
Yizhou Shen, Carlton Shepherd, Chuadhry Mujeeb Ahmed, Shigen Shen, Shui Yu 0001
IEEE Trans. Mob. Comput.1
2025 Availability Evaluation of Industrial Internet of Things Under Malware Propagation: An Extended Reliability Block Diagram Approach Based on Stochastic Games
abstract
The rise of the industrial Internet of Things (IIoT) has enhanced industrial processes through interconnected devices and data exchange, but it also introduces significant security vulnerabilities, such as malware attacks, which threaten system reliability and availability. To address this challenge, we extend the traditional reliability block diagram (RBD) method by integrating stochastic games to evaluate the security and availability of IIoT systems. Our approach constructs a comprehensive Markov transition matrix using additional node states, enabling detailed simulations of malware spread in IIoT networks. By modeling the interactions between malware and IIoT systems through stochastic games, we propose an innovative reinforcement learning algorithm named evaluation-driven Q-learning (EDQL) to solve these complex scenarios. This novel application of EDQL in the realm of availability evaluation is a significant contribution, providing a rare integration of game theory into this field. We also derive the availability of individual IIoT nodes using reliability theory and integrate these insights into the RBD framework. Experimental results demonstrate that the EDQL algorithm significantly outperforms traditional reinforcement learning methods in malware reward. Furthermore, our method effectively evaluates common IIoT topologies and offers practical deployment recommendations, highlighting its practical impact and significance in enhancing IIoT system security and availability.
Shoujian Yu, Ouwen Jin, Yizhou Shen, Guowen Wu, Shui Yu 0001, Shigen Shen
IEEE Trans. Reliab.3
2025 Integrating Deep Spiking Q-Network Into Hypergame-Theoretic Deceptive Defense for Mitigating Malware Propagation in Edge Intelligence-Enabled IoT Systems
abstract
Internet of Things (IoT) systems are susceptible to compromise due to malware propagation, leading to the data breach and information theft. In this paper, we propose a proactive deception-oriented hypergame-theoretic malware propagation-mitigation (DHMPM) model between IoT nodes and edge devices under asymmetric information in edge intelligence (EI)-enabled IoT systems. We then explore malware-propagated deceptive defense strategies based on deep reinforcement learning. Specifically, IoT nodes and edge devices continually adjust their strategies based on obtained utilities under beliefs perceived by uncertainties from the game environment and system dynamics. Built upon the proposed game DHMPM, we next apply spiking neural networks (SNNs) into deep Q-network to form hypergame-theoretic deep spiking Q-network (HGDSQN), practically converging to the optimal malware-propagated deceptive defense strategy in EI-enabled IoT systems. Such SNNs can simulate biological brains with the pulse communication mechanism and break through the bottleneck of temporal processing in traditional models with deep neural networks, realizing intelligent decision-making and real-time malware defense. We eventually perform experimental simulations that assess the effect of attack arrival probability and learning rate on the optimal learning strategy selection, demonstrating the effectiveness of the proposed HGDSQN algorithm.
Yizhou Shen, Carlton Shepherd, Chuadhry Mujeeb Ahmed, Shigen Shen, Shui Yu 0001
IEEE Trans. Serv. Comput.1
2024 Game-theoretic analytics for privacy preservation in Internet of Things networks: A survey
Yizhou Shen, Carlton Shepherd, Chuadhry Mujeeb Ahmed, Shigen Shen, Wenlong Ke, Shui Yu 0001
Eng. Appl. Artif. Intell.1
2024 MFGD3QN: Enhancing Edge Intelligence Defense Against DDoS With Mean-Field Games and Dueling Double Deep Q-Network
abstract
Distributed Denial-of-Service (DDoS) attacks pose a serious threat to the stability and security of edge intelligence devices. To solve this issue, we first describe the cost of edge intelligence environments in detail and introduce the mean-field term, edge intelligence repair speed, and DDoS attack intensity, which provide a theoretical basis for the subsequent model construction. Secondly, based on Hamilton-Jacobi-Bellman (HJB) backward and Fokker-Planck-Kolmogorov (FPK) forward equations, a mean-field model is proposed. The optimal DDoS defense policy is then solved considering the strongest DDoS attack intensity and the best edge intelligence device repair speed. Further, the focus is turned to the interaction between DDoS attackers and the edge intelligence environment. We give the update rule of the mean-field term and construct a mean-field game (MFG) model with a value function. Finally, we propose a multi-agent deep reinforcement learning algorithm called Mean-Field Games with Dueling Double Deep Q-learning Network (MFGD3QN) to solve the optimal DDoS defense policy problem under the MFG model. In the experiment, we compare MFGD3QN with several benchmark algorithms and verify the superiority of the MFGD3QN algorithm in an edge intelligence environment. We also carry out experiments on different parameters of the MFGD3QN algorithm, lower visibility conditions of the edge intelligence environment, and different repair speeds of edge intelligence devices, which verify the robustness and feasibility of the algorithm.
Shigen Shen, Chenpeng Cai, Yizhou Shen, Wenlong Ke, Shui Yu 0001
IEEE Internet Things J.3
2024 Comparative DQN-Improved Algorithms for Stochastic Games-Based Automated Edge Intelligence-Enabled IoT Malware Spread-Suppression Strategies
abstract
Massive volumes of malware spread incidents continue to occur frequently across the Internet of Things (IoT). Owing to its self-learning and adaptive capability, artificial intelligence (AI) can provide assistance for automatically converging to an optimal strategy. By merging AI into edge computing, we consider an edge intelligence-enabled IoT (EIIoT) environment and provide a stochastic learning strategy for suppressing the spread of IoT malware. In particular, we introduce stochastic game theory to symbolise the whole process of the confrontation between IoT malware and edge nodes. Built upon the theoretical framework to demonstrate the specific spread-suppression architecture, we apply the improved Deep Q-Network algorithms including DDQMS, D2QMS and D3QMS that can deduce the optimal EIIoT malware spread-suppression strategy with better performance. Through experiments, we investigate the influence of related parameters on learning strategy selection, recommending the optimal parameters setting of automated EIIoT malware spread-suppression. We also compare the performance of the proposed three DQN-improved algorithms.
Yizhou Shen, Carlton Shepherd, Chuadhry Mujeeb Ahmed, Shui Yu 0001, Tingting Li 0001
IEEE Internet Things J.1
2024 Combining Lyapunov Optimization With Actor-Critic Networks for Privacy-Aware IIoT Computation Offloading
abstract
Opportunistic computation offloading is an effective way to improve the computing performance of Industrial Internet of Things (IIoT) devices. However, as more and more computing tasks are being offloaded to mobile-edge computing (MEC) servers for processing, it can lead to IIoT privacy and security issues, such as personal usage habits. In this paper, we aim to design a Lyapunov-based privacy-aware framework that defines the amount of IIoT user privacy and designs a “reduced amount of privacy” mechanism. We first define the cumulative privacy amount for each IIoT user and trigger the privacy protection mechanism when the cumulative privacy amount exceeds the set privacy threshold. The offloading data generated by the IIoT user is then transferred to local processing, and finally, the cumulative privacy amount of the IIoT user is reduced. This model ensures that the cumulative privacy of all IIoT users remains stable. We further combine the advantages of Lyapunov optimization and actor-critic networks to address the problem of how to make the model learn the optimal policy and maintain the minimum energy consumption in the long run. Especially, this framework integrates model-based optimization and model-free actor-critic networks to handle the offloading problem with very low computational complexity, and Lyapunov optimization ensures that this framework minimizes energy consumption while stabilizing the data queue. It is demonstrated through experimental simulation results that the proposed scheme can maintain data queue stability and minimize energy consumption under strict security.
Guowen Wu, Xihang Chen, Yizhou Shen, Zhiqi Xu, Hong Zhang 0046, Shigen Shen, Shui Yu 0001
IEEE Internet Things J.3
2024 Novel Intrusion Detection Strategies With Optimal Hyper Parameters for Industrial Internet of Things Based on Stochastic Games and Double Deep Q-Networks
abstract
The Industrial Internet of Things (IIoT) has experienced rapid growth in recent years, with an increasing number of interconnected devices, thereby expanding the attack surface. Effectively detecting intrusions is crucial for safeguarding IIoT systems from malicious attacks. However, due to the dynamic and complex nature of the IIoT environment, designing an intrusion detection strategy that balances accuracy and efficiency remains a significant challenge. In this paper, we propose a novel intrusion detection strategy based on stochastic games and deep reinforcement learning (DRL) for detecting attacks effectively while balancing detection accuracy and efficiency in the IIoT. We model the interaction between attackers and detectors as dynamic adversarial stochastic games with incomplete information, theoretically analyze Nash equilibria, and construct a node-based simulation of interconnected infrastructure within the IIoT. We then propose a novel algorithm DDQN-LP combining Double Deep Q-Networks with “lazy penalty” to determine optimal strategies and encourage agents to promptly conclude the game to reduce overhead. Furthermore, we identify different optimal hyperparameters for training our DRL agents and evaluate their efficacy both theoretically and empirically. We compare our proposed algorithm with other reinforcement learning algorithms, and simulations demonstrate our approach has better performance with a higher detection rate as well as lower consumption.
Shoujian Yu, Yizhou Shen, Guowen Wu, Shui Yu 0001, Shigen Shen
IEEE Internet Things J.3
2024 Deep Q-Network-Based Open-Set Intrusion Detection Solution for Industrial Internet of Things
abstract
Industrial Internet of Things (IIoT) has brought a lot of convenience for the industrial world to digitization, automation and intelligence, but it inevitably introduces inherent cyber security risks, resulting in an issue that traditional intrusion detection techniques are no longer sufficient for IIoT environments. To solve this issue, we propose an open-set solution called DC-IDS for IIoT based on deep reinforcement learning. In this solution, the open-set recognition problem in intrusion detection is modeled as a discrete-time Markov decision process, and Deep Q-Network (DQN) is employed to solve it. Meanwhile, a Conditional Variational Auto-Encoder is introduced to the value network in DQN. Therefore, the open-set recognition problem in intrusion detection is divided into two subproblems, namely known traffic fine-grained classification problem and unknown attacks recognition problem. We use DQN to solve the known traffic fine-grained classification problem. Since the reconstruction error of known traffic is generally smaller than the reconstruction error of unknown attacks, we use reconstruction error to recognize unknown attacks. Experiments on IIoT dataset TON-IoT demonstrate the effectiveness of DC-IDS model, which achieves better performance in terms of the recognition rate of unknown attacks as well as the stability of the model compared to previous proposed methods.
Shoujian Yu, Rong Zhai, Yizhou Shen, Guowen Wu, Hong Zhang 0046, Shui Yu 0001, Shigen Shen
IEEE Internet Things J.3
2024 SGD3QN: Joint Stochastic Games and Dueling Double Deep Q-Networks for Defending Malware Propagation in Edge Intelligence-Enabled Internet of Things
abstract
Malware propagation in IoT (Internet of Things) systems can lead to data leakages, financial losses, and other serious consequences. To solve this issue, we propose a new active IoT malware propagation defence work. Specifically, aided by stochastic games, we express the process of cyber conflicts between IoT system nodes and edge devices considering malware propagation in edge intelligence-enabled IoT. Here, IoT system nodes and edge devices choose their own strategies and receive the corresponding rewards determined by the current state and strategy. After that, the game randomly moves to the next stage according to the distribution of probabilities and the participants’ strategies until reaching the fixed Nash equilibrium point. Following a theoretical analysis, we design and implement SGD3QN (Stochastic Games and Dueling Double Deep Q-networks)—a novel algorithm to receive the optimal strategy for mitigating IoT malware propagataion in practice. Here, the Dueling Double Deep Q-networks are acted as an end-to-end decision control system, in which IoT malware propagataion environment is used as the input to obtain the failure or success experience to update the network parameters, followed by making the optimal decision output. Afterwards, we perform experimental simulations that probe the influence of batch size and replay memory size on the optimal IoT malware propagation defense strategy selection and prove the ascendancy of the proposed SGD3QN-aided decision-making algorithm.
Yizhou Shen, Carlton Shepherd, Chuadhry Mujeeb Ahmed, Shigen Shen, Shui Yu 0001
IEEE Trans. Inf. Forensics Secur.1
2024 Mean-Field Game-Based Task-Offloaded Load Balance for Industrial Mobile Edge Computing Systems Using Software-Defined Networking
abstract
Smart devices (SDs) used in the Industrial Internet of Things can generate computational tasks for processing the data generated during production. However, due to the limited processing power of SDs, it is necessary to transfer these computational tasks to more powerful devices for processing. To this end, we propose a Mobile Edge Computing (MEC) system based on a Software Defined Network (SDN) for SDs to offload their computational tasks. This MEC system includes multiple MEC servers to handle numerous SDs, which leads to load-balancing challenges among these servers. To tackle this problem, we develop a computational offloading model based on mean-field game theory and introduce a mean-field game-based load-balancing algorithm (MFGLB), which reduces processing latency and facilitates task scheduling through Multi-Agent Deep Reinforcement Learning. Each SD in the MEC system is considered a participant in the mean-field game, simplifying the complex stochastic game into a more manageable dual-agent game. We then prove the existence of Nash Equilibrium for this mean-field game. To evaluate the effectiveness of our MFGLB algorithm, we compare its performance with traditional load-balancing algorithms and a stochastic game-based load-balancing algorithm. Our experimental results demonstrate the superiority of MFGLB in reducing processing latency and addressing load imbalances.
Guowen Wu, Hui Wang 0011, Hong Zhang 0046, Yizhou Shen, Shigen Shen, Shui Yu 0001
IEEE Trans. Mob. Comput.4
2024 SAC-PP: Jointly Optimizing Privacy Protection and Computation Offloading for Mobile Edge Computing
abstract
The emergence of mobile edge computing (MEC) imposes an unprecedented pressure on privacy protection, although it helps the improvement of computation performance including energy consumption and computation delay by computation offloading. To this end, we concern about the privacy protection in the MEC system with a curious edge server. We present a deep reinforcement learning (DRL)-driven computation offloading strategy designed to concurrently optimize privacy protection and computation cost. We investigate the potential privacy breaches resulting from offloading patterns, propose an attack model of privacy theft, and correspondingly define an analytical measure to assess privacy protection levels. In pursuit of an ideal computation offloading approach, we propose an algorithm, SAC-PP, which integrates actor-critic, off-policy, and maximum entropy to improve the efficiency of learning processes. We explore the sensitivity of SAC-PP to hyperparameters and the results demonstrate its stability, which facilitates application and deployment in real environments. The relationship between privacy protection and computation cost is analyzed with different reward factors. Compared with benchmarks, the empirical results from simulations illustrate that the proposed computation offloading approach exhibits enhanced learning speed and overall performance.
Shigen Shen, Xuanbin Hao, Zhengjun Gao, Guowen Wu, Yizhou Shen, Hong Zhang 0046, Qiying Cao, Shui Yu 0001
IEEE Trans. Netw. Serv. Manag.5
2022 Signaling game-based availability assessment for edge computing-assisted IoT systems with malware dissemination
Yizhou Shen, Shigen Shen, Zongda Wu, Haiping Zhou, Shui Yu 0001
J. Inf. Secur. Appl.1