Minghui Min

dblp:203/9662 · DBLP profile ↗
← Back
25ranked-venue papers
10as first author
18since 2021 · last 2026
0000-0001-9388-6434ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 21 · 8 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 High Dimensional Distributed Gradient Descent with Arbitrary Number of Byzantine Attackers
abstract
Adversarial attacks pose a major challenge to distributed learning systems, prompting the development of numerous robust learning methods. However, most existing approaches suffer from the curse of dimensionality, i.e. the error increases with the number of model parameters. In this paper, we make a progress towards high dimensional problems, under arbitrary number of Byzantine attackers. The cornerstone of our design is a direct high dimensional semi-verified mean estimation method. The idea is to identify a subspace with large variance. The components of the mean value perpendicular to this subspace are estimated using corrupted gradient vectors uploaded from worker machines, while the components within this subspace are estimated using auxiliary dataset. As a result, a combination of large corrupted dataset and small clean dataset yields significantly better performance than using them separately. We then apply this method as the aggregator for distributed learning problems. The theoretical analysis shows that compared with existing solutions, our method gets rid of sqrt{d} dependence on the dimensionality, and achieves minimax optimal statistical rates. Numerical results validate our theory as well as the effectiveness of the proposed method.
Wenyu Liu 0014, Zong Ke, Minghui Min, Puning Zhao
AAAI5
2026 PowerCloak: Differential Privacy-Based Power Perturbation for Location Privacy in UAV-Enabled Wireless Powered Communication Networks
Zijian Xiang, Peng Zhang 0065, Minghui Min, Shiyin Li, Rui Zhang 0006, Dusit Niyato, Zhu Han 0001
WCNC3
2026 Comparative Performance Analysis of Different Hybrid NOMA Schemes
abstract
Hybrid non-orthogonal multiple access (H-NOMA), which combines the advantages of pure NOMA and conventional OMA, has emerged as a highly promising multiple access technology for future wireless networks. While recent studies have proposed various H-NOMA systems using different successive interference cancellation (SIC) methods, their analyses typically assume a fixed channel gain order between paired users. However, in practice, user pairing is often configured typically based on long-term network deployment requirements or statistical channel characteristics (e.g., geographic layout or average channel gain) rather than instantaneous channel states. This practical pairing strategy leads to random channel gain ordering, where the relative channel gains between paired users are inherently stochastic and time-varying. This aspect is critical and fundamentally affects system performance, yet remains understudied. To address this issue, this paper analyzes the performance of three H-NOMA schemes under such random channel gain ordering: (a) fixed-order SIC (FSIC) aided H-NOMA; (b) hybrid SIC with non-power adaptation (HSIC-NPA) aided H-NOMA; and (c) hybrid SIC with power adaptation (HSIC-PA) aided H-NOMA. For the opportunistic users seeking to maximize data rate, the closed-form expressions for the probability that each H-NOMA scheme underperforms conventional OMA are derived rigorously. Asymptotic analyses in the high SNR regime are also developed. Simulation results validate the theoretical derivations and demonstrate the performance of the H-NOMA schemes across different SNR scenarios, thereby offering foundational insights for deploying robust H-NOMA in next-generation wireless systems.
Ning Wang 0004, Yanshi Sun, Minghui Min, Shiyin Li
IEEE Internet Things J.4
2026 Safe TD3 for Personalized Spatiotemporal Trajectory Privacy Protection
abstract
With the widespread adoption of location-based services (LBS), user-generated trajectory data shows strong spatiotemporal correlation, rendering it highly vulnerable to inference attacks that expose sensitive information. In particular, once semantic locations like “hospital” and “bank” are identified, the risk of trajectory leakage increases substantially. To address this issue, this paper formulates a personalized spatiotemporal trajectory privacy protection framework, which is designed to protect locations with varying semantic sensitivities on the trajectory from the attacker with spatiotemporal correlation information. We model the trajectory privacy protection problem as a Markov Decision Process (MDP) and introduce the reinforcement learning (RL) technique to adjust the privacy parameters dynamically. Specifically, we leverage the twin delayed deep deterministic policy gradient (TD3) algorithm to enhance the stability and accuracy of policy evaluation, enabling efficient learning of optimal policies in continuous action spaces. Furthermore, a safe exploration strategy is incorporated to continuously evaluate and avoid high-risk state-action pairs, thereby enhancing privacy protection. Simulation results demonstrate that the proposed mechanism significantly improves privacy protection while effectively reducing Quality of Service (QoS) loss, exhibiting better convergence and overall system utility.
Minghui Min, Minghui Dai, Shiyin Li, Hongliang Zhang 0001, Miao Pan, Dusit Niyato, Zhu Han 0001
IEEE Trans. Mob. Comput.1
2026 Personalized Location Privacy-Aware Task Offloading: A Dual-Agent DRL Approach
abstract
Multi-access Edge Computing (MEC) enables users to handle resource-intensive and latency-sensitive tasks. However, the offloading behaviors, which are closely correlated with wireless channel conditions, can inadvertently reveal users' location information to untrustworthy MEC servers. Existing location privacy-aware task offloading (LPTO) mechanisms have not fully considered and comprehensively analyzed personalized location privacy protection requirements. To address this gap, this paper proposes a differential privacy (DP)-based personalized LPTO mechanism for MEC environments that jointly optimizes the perturbation region, privacy budget, and offloading rate while maximizing the offloading utility. We quantify personalized privacy requirements by incorporating task sensitivity, user privacy preference, and task priority. Then, we propose a two-timescale (2Ts) optimization framework to solve the complex personalized location privacy-aware task offloading optimization problem. Specifically, we optimize the perturbation region on a long timescale to align with long-term privacy requirements. In contrast, the offloading ratio and privacy budget are dynamically optimized on a short timescale based on instantaneous channel states and offloading workloads. Furthermore, we model the privacy-aware offloading problem as a Markov decision process (MDP) and develop a dual-agent deep reinforcement learning (DRL)-based personalized LPTO mechanism (DDPLM) to optimize strategies under dynamic MEC systems. Simulation results validate that the proposed DDPLM achieves personalized location privacy protection while reducing computational costs.
Minghui Min, Peng Zhang 0065, Yue Zhang 0027, Wenmin Kuang, Hongliang Zhang 0001, Shiyin Li, Dusit Niyato, Zhu Han 0001
IEEE Trans. Mob. Comput.1
2026 Time-Varying Transmission Rates-Aware Dependent Task Offloading and Local Resource Allocation in Multi-Access Edge Computing
abstract
With the increasing diversity and complexity of mobile applications, an application typically needs to execute multiple dependent tasks to achieve its functionality. Computation offloading in multi-access edge computing aims to improve user experience, such as reducing makespan and terminal energy consumption, by offloading some tasks to the designated edge server. Task dependencies impose constraints on the execution order of the tasks, which complicates the offloading decisions. Besides, transmission rates exhibit fluctuations in real-world scenarios due to the mobility of users, and they pose new challenges to the problem of dependent task offloading. To minimize the terminal energy consumption under the given deadline, a joint optimization problem of dependent task offloading and local computing resource allocation with time-varying transmission rates is investigated. Since the proposed problem is NP-hard, we decompose it into two subproblems to reduce the complexity and deal with the coupling of decision variables. Specifically, we first solve the subproblem of dependent task offloading by generating special task sets to implement divide and conquer with a fixed local processing speed. Then, we employ Karush-Kuhn-Tucker (KKT) method to solve the subproblem of local resource allocation with the offloading solution derived from the first subproblem while ensuring that the makespan constraint is met. The proposed scheme makes dynamic decisions on dependent task offloading and resource allocation according to the fluctuations of transmission rates, and dynamically adjusts the solution accordingly to achieve better performance. Experimental results show that our solution outperforms baseline methods, reducing makespan by 27–31%, terminal energy consumption by 70–77%, while achieving 25–65% higher service success ratio on average.
Qiang Zhang 0052, Minghui Min, Zhu Han 0001
IEEE Trans. Mob. Comput.2
2026 Aerial IRS Deployment-Aided Secure Computation Offloading Against DISCO Jamming Attacks
Minghui Min, Peng Zhang 0065, Jiayang Xiao, Shiyin Li, Huan Huang 0001, Hongliang Zhang 0001, Zhu Han 0001
IEEE Trans. Wirel. Commun.1
2026 An Energy Efficient Design of Hybrid NOMA Based on Hybrid SIC With Power Adaptation
abstract
Hybrid non-orthogonal multiple access (H-NOMA) technology, which combines the benefits of non-orthogonal multiple access (NOMA) and orthogonal multiple access (OMA) through flexible resource allocation in a single transmission, has shown great potential for enhancing the performance of wireless communication systems. To further exploit the potential of H-NOMA, this paper proposes a novel design of H-NOMA which jointly incorporates hybrid successive interference cancellation (HSIC) and power adaptation (PA) in the NOMA transmission phase, by introducing a power adaptation factor . For a given power reducing coefficient β, which ensures that the energy consumption of the proposed scheme is lower than that of conventional OMA, the probability that the achievable rate of the proposed HSIC-PA aided H-NOMA scheme fails to outperform its OMA counterpart is derived in closed form. Besides, the impact of user pairing is considered. Furthermore, the asymptotic analysis shows that the aforementioned probability of the proposed H-NOMA scheme can approach zero in the high signal-to-noise ratio (SNR) regime without constraints on either users’ target rates or transmit power. By dynamically adjusting the transmission power of the opportunistic user and the decoding order of HSIC, signal interference between the legacy user and the opportunistic user can be effectively controlled, thereby improving the achievable rate and energy efficiency of the opportunistic user. This represents a significant improvement over conventional H-NOMA schemes, which require specific restrictive conditions to make the probability that their achievable rate underperforms OMA approach zero at high SNR, as shown in existing work. The above observation indicates that, with lower energy consumption, the proposed HSIC-PA aided H-NOMA can achieve a higher data rate than pure OMA with probability 1 at high SNR, leading to improved energy efficiency. Finally, numerical results are provided to verify the accuracy of the analysis and also to demonstrate the superior performance of the proposed H-NOMA scheme.
Ning Wang 0004, Yanshi Sun, Minghui Min, Yuanwei Liu, Shiyin Li
IEEE Trans. Wirel. Commun.4
2025 Personalized Semantic Trajectory Privacy Protection in Location-Based Services: A TD3-Based Approach
abstract
The swift advancement of Location-Based Services (LBSs) raises the danger of trajectory privacy being breached, since the location semantic tags can easily disclose users' sensitive information. Additionally, attackers can exploit temporal correlations to infer sensitive personal information. This paper formulates a personalized semantic trajectory privacy protection framework designed to protect locations with varying sensitivities on the trajectory from the attacker with temporal correlation information. We model the trajectory privacy protection problem as a Markov Decision Process (MDP) and introduce the Reinforcement Learning (RL) technique to dynamically adjust the privacy parameters. Specifically, we leverage the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm to enhance the stability and accuracy of policy evaluation, enabling efficient learning of optimal policies in continuous action spaces. Simulation results indicate that the TD3-based personalized semantic trajectory privacy protection mechanism effectively balances the Quality of Service and semantic trajectory privacy while realizing personalized trajectory privacy protection.
Minghui Dai, Minghui Min, Jinling Song, Hongliang Zhang 0001, Zhu Han 0001
WCNC3
2025 Gradient descent based polarization channel estimation in extremely largescale MIMO systems
Jinling Song, Shiyin Li, Faguang Wang, Minghui Min
Wirel. Networks6
2024 Protecting Personalized Trajectory with Differential Privacy under Temporal Correlations
abstract
Location-based services (LBSs) in vehicular ad hoc networks (VANETs) offer users numerous conveniences. However, the extensive use of LBSs raises concerns about the privacy of users' trajectories, as adversaries can exploit temporal correlations between different locations to extract personal information. Additionally, users have varying privacy requirements depending on the time and location. To address these issues, this paper proposes a personalized trajectory privacy protection mechanism (PTPPM). This mechanism first uses the temporal correlation between trajectory locations to determine the possible location set for each time instant. We identify a protection location set (PLS) for each location by employing the Hilbert curve-based minimum distance search algorithm. This approach incor-porates the complementary features of geo-indistinguishability and distortion privacy. We put forth a novel Permute-and-Flip mechanism for location perturbation, which maps its initial application in data publishing privacy protection to a location perturbation mechanism. This mechanism generates fake locations with smaller perturbation distances while improving the balance between privacy and quality of service (QoS). Simulation results show that our mechanism outperforms the benchmark by providing enhanced privacy protection while meeting user's QoS requirements.
Mingge Cao, Haopeng Zhu, Minghui Min, Yulu Li, Shiyin Li, Hongliang Zhang 0001, Zhu Han 0001
WCNC3
2024 Geo-Perturbation for Task Allocation in 3-D Mobile Crowdsourcing: An A3C-Based Approach
abstract
Location privacy protection (LPP) has become a key concern during mobile crowdsourcing (MCS) task allocation. Existing LPP mechanisms for MCS applications mainly focus on two-dimensional (2D) plane scenarios or directly apply 2D techniques into three-dimensional (3D) space scenarios, leaving the height dimension of 3D geolocation vulnerable to privacy breaches. To facilitate the LPP in 3D MCS, we propose a learning-based geo-perturbation mechanism using 3D geo-indistinguishability (3D-GI). In this mechanism, we first define an optimization objective to balance location privacy and MCS server profit, making it adaptable to different types of MCS applications. Then, we adopt the Asynchronous Advantage Actor-Critic (A3C) algorithm to design a reinforcement learning (RL)-based approach without knowing the accurate system and attack models. This approach enables us to derive the optimal perturbation policy in continuous policy space and accelerates the learning speed using asynchronous multi-thread training. Simulation results demonstrate that the proposed mechanism can better balance location privacy and server profit in 3D MCS applications compared to existing benchmarks.
Minghui Min, Haopeng Zhu, Junhuai Xu, Jingwen Tong, Shiyin Li, Jiangang Shu
IEEE Internet Things J.1
2024 Multi-System Fusion Positioning Method Based on Factor Graph
abstract
Ultra-wideband (UWB) positioning system offers high-precision location capabilities. However, it introduces positive biases in complex environments. Pedestrian Dead Reckoning (PDR) algorithm based on Inertial Measurement Unit (IMU) can maintain robust tracking even in cases of abrupt changes in pedestrian trajectories but suffers from cumulative errors. Therefore, in this study, the strengths of both systems are combined. Hence, a factor graph model is established to enhance the multi-system fusion localization method based on factor graphs. Experimental verification in both straight-line trajectories and scenarios involving state mutations demonstrates an integrated average positioning accuracy within 0.1m. When compared to traditional system fusion localization methods, the accuracy is enhanced by more than 50%.
Sheng Xing, Minghui Min, Shiyin Li
IEEE Signal Process. Lett.4
2024 Personalized 3D Location Privacy Protection With Differential and Distortion Geo-Perturbation
abstract
The rapid development of indoor location-based services (LBS) has raised concerns about location privacy protection in the 3-dimensional (3D) space. The existing 2-dimensional (2D) location privacy protection mechanisms (LPPMs) cannot effectively resist attacks in 3D environments. Furthermore, users may have various sensitive attributes at different locations and times. In this paper, we first formally study the relationship between two complementary notions of geo-indistinguishability and distortion privacy (i.e., expected inference error) in the 3D space and develop a two-phase personalized 3D LPPM (P3DLPPM). In Phase I, we search for neighboring locations to formulate a protection location set (PLS) for hiding the actual location based on the above-mentioned relationship. To realize this, we develop a 3D Hilbert curve-based minimum distance searching algorithm to find the PLS with minimum diameter for each location while guaranteeing differential privacy. In Phase II, we put forth a novel Permute-and-Flip mechanism for location perturbation, which maps its initial application in data publishing privacy protection to a location perturbation mechanism. It generates fake locations with smaller perturbation distances while improving the balance between privacy and quality of service (QoS). Simulation results show that the proposed P3DLPPM can significantly improve personalized privacy protection while meeting the user's QoS needs.
Minghui Min, Haopeng Zhu, Jiahao Ding, Shiyin Li, Liang Xiao 0003, Miao Pan, Zhu Han 0001
IEEE Trans. Dependable Secur. Comput.1
2023 Privacy-Aware Laser Wireless Power Transfer for Aerial Multi-Access Edge Computing: A Colonel Blotto Game Approach
abstract
This article studies the integration of laser-beamed wireless power transfer (WPT) into high-altitude platform (HAP)-aided multiaccess edge computing (MEC) systems for the HAP-connected aerial user equipments (AUEs). By discretizing the 3-D coverage space of the HAP, we present a multitier tile grid-based spatial structure to provide the aerial locations in the form of tile grids to AUEs for laser charging. We identify a new privacy vulnerability caused by the openness during the WPT signaling transfer in the presence of a terrestrial adversary, which is able to launch the attacks by distributing the false tile grids to the AUEs. To address this vulnerability and enhance the location privacy of AUEs, we then propose a Colonel Blotto (CB) game framework to formulate the competitive tile grid allocation problem for the HAP and the adversary. The attack-defense interaction between the adversary and the HAP as a defender in their tile grid allocations to the AUEs is formulated as a CB game, which models the competition of two players for limited resources over multiple battlefields. Moreover, we derive the mixed-strategy Nash equilibria of the game for both symmetric and asymmetric tile grids between two players. Simulation results show that the proposed framework significantly outperforms the design baselines with a given privacy protection level in terms of system-wide expected total utilities.
Long Zhang 0003, Yao Wang 0001, Minghui Min, Chao Guo 0002, Vishal Sharma 0001, Zhu Han 0001
IEEE Internet Things J.3
2022 Multiagent DDPG-Based Joint Task Partitioning and Power Control in Fog Computing Networks
abstract
Fog computing is an energy-efficient and cost-effective paradigm to help alleviate the pressure of resource-constrained mobile devices (MDs) running computation-intensive applications. In this article, we investigate the joint task partitioning and power control problem in a fog computing network with multiple MDs and fog devices (FDs), where each MD has to complete a periodic computation task under the constraints of delay and energy consumption. Each task can be partitioned into multiple subtasks and offloaded to the FDs according to the task partition strategy and transmission power strategy to reduce task execution delay and energy consumption. To this end, we present a multiagent deep deterministic policy gradient (MADDPG)-based task offloading algorithm for MDs to maximize the long-term system utility including the execution delay and energy consumption. Each MD inputs the local information, e.g., the task requirements, the available communication, and computation resources of the FDs, the computation resources, and the battery level of the MD into a distributed actor network to generate a task offloading policy, while a centralized critic network is used to update the weights of the actor networks to improve offloading performance. Numerical simulation results demonstrate the effectiveness of the proposed scheme in improving the system utility, reducing the average execution delay as well as the average energy consumption.
Zhipeng Cheng, Minghui Min, Minghui LiWang, Lianfen Huang, Zhibin Gao
IEEE Internet Things J.2
2022 3D Geo-Indistinguishability for Indoor Location-Based Services
abstract
Indoor location-based services (LBS) are widely used in large-scale indoor buildings, such as high-rise hospitals and multi-story shopping malls. At the same time, location privacy protection in such three-dimensional (3D) space has recently attracted considerable attention. Currently, most existing location privacy protection schemes focus on two-dimensional (2D) location protection and fail to prevent location inference attacks when the user’s location data include height dimension, i.e., 3D geolocation. Enlightened by the concept of differential privacy, in this paper we first study the impact factors of the degree of indistinguishability of 3D geolocations. Then, we quantify location privacy for LBS applications in the 3D space with geo-indistinguishability (3D-GI) rigorously and provably. We develop a mechanism of three-variates Laplacian to generate perturbed locations considering the locations’ X, Y, and Z-coordinates simultaneously, guaranteeing geo-indistinguishability. Furthermore, the discretization noise-adding mechanism is studied to satisfy geo-indistinguishability in the 3D space under the finite precision of hardware/devices. Considering the discretized mechanism can only satisfy geo-indistinguishability in finite 3D space and users visit the limited regions, we further study the truncation of the Laplacian mechanism to limit the generated perturbed locations within a specific region. Simulation results demonstrate that the proposed 3D-GI outperforms the benchmarks while guaranteeing privacy regardless of the adversary’s prior knowledge.
Minghui Min, Liang Xiao 0003, Jiahao Ding, Hongliang Zhang 0001, Shiyin Li, Miao Pan, Zhu Han 0001
IEEE Trans. Wirel. Commun.1
2021 Joint Client Selection and Task Assignment for Multi-Task Federated Learning in MEC Networks
abstract
In this paper, we investigate the multi-task federated learning in mobile edge computing (MEC) networks where a central server assigns different federated learning tasks to different MEC servers and select feasible clients to participate in the federated learning training process. The problem is formulated as a joint client selection and task assignment problem to maximize the total utility of all tasks, subject to the trained model quality and total training latency. Since the above-mentioned problem is NP-Hard, it poses challenges to obtain the optimal solution within polynomial time, the problem is transformed into a many-to-one-to-one 3D matching problem. To further reduce the computation while ensuring the matching stability, we first adopt the spectral clustering algorithm to cluster the clients into multiple client clusters. Then we reformulate the problem as a 3-Partite weighted hypergraph total weight maximization problem. Finally, we propose a greedy and local search (GLS) based algorithm to resolve the problem. Simulation results demonstrate the effectiveness of the proposed algorithm as compared with baseline schemes.
Zhipeng Cheng, Minghui Min, Minghui LiWang, Zhibin Gao, Lianfen Huang
GLOBECOM2
2020 Joint Task Offloading and Resource Allocation for Mobile Edge Computing in Ultra-Dense Network
abstract
Mobile edge computing (MEC) enabled user-centric ultra-dense network (UDN) is a promising solution to the energy constrained mobile users with delay-sensitive and computation intensive applications. Due to the high density of access points and MEC servers in UDN, both task offloading decision, power control, communication and computation resource allocation need to be addressed. To this end, we consider the joint problem of task offloading, uplink transmission power control, communication and computation resource allocation in a UDN, where the task of each user can be partitioned into several subtasks and offloaded to different access points. To handle the continuous action space of task partitioning and power control, we propose a multi-agent deep deterministic policy gradient (MADDPG) approach to solve this problem. Simulation results reveal the effectiveness of the proposed method.
Zhipeng Cheng, Minghui Min, Zhibin Gao, Lianfen Huang
GLOBECOM2
2019 Protecting Semantic Trajectory Privacy for VANET with Reinforcement Learning
abstract
Location-based services in vehicular ad hoc networks (VANETs) have to protect user privacy and address the challenge due to the disclosure of the vehicle movement trajectory. In this paper, we propose an reinforcement learning (RL) based differential privacy mechanism that randomizes the released vehicle locations to protect the semantic trajectory of the vehicle and uses RL to select the obfuscation policy. Based on the semantic location of the vehicle and the attack history, this scheme enables a vehicle to optimize the obfuscation policy in terms of the privacy gain and the quality of service loss without being aware of the current attack model in a dynamic privacy protection game. Simulation results show that this scheme can increase the privacy gain, decrease the quality of service loss, and thus improve the utility of the vehicle in comparison with a benchmark scheme.
Weihang Wang 0001, Minghui Min, Liang Xiao 0003, Ye Chen 0011, Huaiyu Dai
ICC2
2019 Learning-Based Privacy-Aware Offloading for Healthcare IoT With Energy Harvesting
abstract
Mobile edge computing helps healthcare Internet of Things (IoT) devices with energy harvesting provide satisfactory quality of experiences for computation intensive applications. We propose a reinforcement learning (RL)-based privacy-aware offloading scheme to help healthcare IoT devices protect both the user location privacy and the usage pattern privacy. More specifically, this scheme enables a healthcare IoT device to choose the offloading rate that improves the computation performance, protects user privacy, and saves the energy of the IoT device without being aware of the privacy leakage, IoT energy consumption, and edge computation model. This scheme uses transfer learning to reduce the random exploration at the initial learning process and applies a Dyna architecture that provides simulated offloading experiences to accelerate the learning process. A post-decision state learning method uses the known channel state model to further improve the offloading performance. We provide the performance bound of this scheme regarding the privacy level, the energy consumption, and the computation latency for three typical healthcare IoT offloading scenarios. Simulation results show that this scheme can reduce the computation latency, save the energy consumption, and improve the privacy level of the healthcare IoT device compared with the benchmark scheme.
Minghui Min, Xiaoyue Wan, Liang Xiao 0003, Ye Chen 0011, Minghua Xia, Di Wu 0001, Huaiyu Dai
IEEE Internet Things J.1
2018 Reinforcement Learning-Based Interference Control for Ultra-Dense Small Cells
abstract
The densification deployment of small cells emerging into 5G cellular networks can achieve high capacity, but is faced with the challenge of how to manage energy consumption and inter-cell interference well in time-varying channels. In this paper, we propose a reinforcement learning based downlink power control algorithm to manage interference for the ultra-dense small cell networks. More specifically, base stations of the small cells use Q-learning to select the downlink transmit powers. A transfer learning method called hotbooting is applied to further accelerate the learning speed and save the energy consumption based on the estimated user density without being aware of the network and channel model of the other small cells. Simulation results demonstrate this scheme significantly improves the network throughput and saves the energy consumption compared with the benchmark, a data-driven based transmission power adaptation scheme.
Hailu Zhang, Minghui Min, Liang Xiao 0003, Sicong Liu 0002, Peng Cheng 0001, Mugen Peng
GLOBECOM2
2018 Learning-Based Defense against Malicious Unmanned Aerial Vehicles
abstract
Adversary unmanned aerial vehicles (UAVs) seriously threaten public security and user privacy. In this paper, we propose a reinforcement learning (RL) based defense framework to address malicious UAVs close to a target estate such as a company or an institute. This framework uses Q-learning to choose the defense policy such as jamming the global positioning system signals (GPS) and hacking, and laser shooting. According to the defense history and the current security status of the target estate, this scheme can improve the UAV defense performance in the dynamic game without being aware of the UAV attack policy and environment model in the area of interests. Simulation results show that this scheme can reduce the risk rate of the estate and improve the utility compared with the benchmark scheme against malicious UAVs.
Minghui Min, Liang Xiao 0003, Dongjin Xu, Lianfen Huang, Mugen Peng
VTC Spring1
2018 Defense Against Advanced Persistent Threats in Dynamic Cloud Storage: A Colonel Blotto Game Approach
abstract
Advanced persistent threat (APT) attackers apply multiple sophisticated methods to continuously and stealthily steal information from the targeted cloud storage systems and can even induce the storage system to apply a specific defense strategy and attack it accordingly. In this paper, the interactions between an APT attacker and a defender allocating their central processing units (CPUs) over multiple storage devices in a cloud storage system are formulated as a Colonel Blotto game. The Nash equilibria of the CPU allocation game are derived for both symmetric and asymmetric CPUs between the APT attacker and the defender to evaluate how the limited CPU resources, the data storage size and the number of storage devices impact the expected data protection level and the utility of the cloud storage system. A CPU allocation scheme based on “hotbooting” policy hill-climbing that exploits the experiences in similar scenarios to initialize the quality values to accelerate the learning speed is proposed for the defender to achieve the optimal APT defense performance in the dynamic game without being aware of the APT attack model and the data storage model. A hotbooting deep${Q}$-network-based CPU allocation scheme further improves the APT detection performance for the case with a large number of CPUs and storage devices. Simulation results show that our proposed reinforcement learning-based CPU allocation can improve both the data protection level and the utility of the cloud storage system compared with the${Q}$-learning-based CPU allocation against APTs.
Minghui Min, Liang Xiao 0003, Caixia Xie, Mohammad Hajimirsadeghi, Narayan B. Mandayam
IEEE Internet Things J.1
2017 Defense against advanced persistent threats: A Colonel Blotto game approach
abstract
An Advanced Persistent Threat (APT) attacker applies multiple sophisticated methods to continuously and stealthily attack targeted cyber systems. In this paper, the interactions between an APT attacker and a cloud system defender in their allocation of the Central Processing Units (CPUs) over multiple devices are formulated as a Colonel Blotto game (CBG), which models the competition of two players under given resource constraints over multiple battlefields. The Nash equilibria (NEs) of the CBG-based APT defense game are derived for the case with symmetric players and the case with asymmetric players each with different total number of CPUs. The expected data protection level and the utility of the defender are provided for each game at the NE. An APT defense strategy based on the policy hill-climbing (PHC) algorithm is proposed for the defender to achieve the optimal CPU allocation distribution over the devices in the dynamic defense game without being aware of the APT attack model. Simulation results have verified the efficacy of our proposed algorithm, showing that both the data protection level and the utility of the defender are improved compared with the benchmark greedy allocation algorithm.
Minghui Min, Liang Xiao 0003, Caixia Xie, Mohammad Hajimirsadeghi, Narayan B. Mandayam
ICC1