VLDB 2026 Research / reviewers in the wild / expert
Liang Xiao 0003
dblp:x/LiangXiao3
· DBLP profile ↗
153ranked-venue papers
41as first author
79since 2021 · last 2026
0000-0003-2402-611XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 111 · 31 first-author · 61 since 2021Security and privacy · 14 · 7 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 since 2021Systems, architecture and hardware · 5 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM-Aided UAV Routing Against Jamming Attacks
Jieling Li, Liang Xiao 0003, Qiaoxin Chen, Chengyao Wang |
ICC | 3 |
| 2026 | LLM-Aided UAV Anti-jamming Communications Based on Reinforcement Learning
Pengli Zhang, Mingyang Fang, Liang Xiao 0003, Qiaoxin Chen, Jieling Li |
ICC | 3 |
| 2026 | RLCVP: Collaborative Vehicular Perception Against Data Fabrication Attacks
Zhiping Lin 0002, Liang Xiao 0003, Zefang Lv, Chen Chen 0037, Liqing Ye |
WCNC | 3 |
| 2026 | UAV-Based Jamming Detection for Large Language Model-Enabled Wireless NetworksabstractReinforcement learning (RL) has been used for uncrewed aerial vehicle (UAV)-based jamming detection by selecting hypothesis-test thresholds on received signal strength (RSS) to safeguard large language model (LLM) inference over wireless links. However, existing methods struggle against smart jammers with dynamic jamming probability. In this paper, we propose a UAV-based jamming detection scheme for LLM-enabled wireless networks to detect smart jamming. The proposed scheme augments traditional wireless features with an LLM-inferred communication environment indicator derived from multimodal sensing (e.g., images/videos from mobile devices), improving accuracy beyond RSS alone. Our approach jointly optimizes the test threshold and the UAV’s sniffer channel assignment, even with a limited number of antennas. We achieve this via a RL policy that leverages UAV imagery, e.g., suspicious radio devices and measuring relative distances, to accelerate detection speed. We implement the proposed scheme on a UAV equipped with a universal software radio peripheral to alert on jamming attacks in a wireless network integrated with a 7-billion-parameter vision-language multimodal LLM. Experiments with a UAV at 3 m detecting a smart jammer during image–text LLM inference across three mobile devices and an edge server demonstrate improvements in both detection accuracy and speed. Qiaoxin Chen, Liang Xiao 0003, Jieling Li, Haoyu Chen 0005, Hongbin Jin, Ying-Jun Angela Zhang |
IEEE Trans. Commun. | 2 |
| 2026 | Reinforcement Learning Based Model Adaptation and Partition for Dynamic DNN Empowered Mobile VisionabstractMost mobile vision systems adopt background subtraction or task offloading, but inference performance remains constrained by background dynamics. Dynamic deep neural networks (DNNs) scale computation without relying on historical information. However, a dynamic DNN tuned to a specific spatial granularity cannot excel across all images. For example, resolution level models perform well on images with sparse and large objects but poorly on images with dense and small objects. Changing system overheads make the choice of an optimal edge-assisted scheme for dynamic DNNs difficult. In this paper, we propose an efficient reinforcement-learning-based model adaptation and partition scheme for mobile vision empowered by dynamic DNNs, which jointly adapts the backbone network and the partition point of the inference model to accelerate inference without degrading inference accuracy. The scheme selects hyperparameters for dynamic DNNs, such as spatial granularity and complexity level, based on image characteristics and measured performance without relying on previous frames. A model partition algorithm is proposed to split the changing inference model in polynomial complexity using the Ford–Fulkerson algorithm. Experimental results based on object detection and image classification show that the proposed scheme reduces inference latency and energy consumption by over 71.2% and 31.8%, without degrading inference accuracy compared to the benchmarks. Liang Xiao 0003, Tuhao Li, Zefang Lv, Ziyue Qiao, Hui Xiong 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2026 | Reinforcement Learning-Based Edge-Assisted Inference With Multimodal DataabstractDeep neural networks (DNNs) extract embeddings from multimodal data such as audio, images, and LiDAR, each with heterogeneous computational demands and representation abilities to support multimodal services such as audio-visual speech recognition. Reinforcement learning (RL)-based model selection and splitting schemes determine the model variants from DNNs and the components to offload in unimodal models to reduce inference latency, aiming for a three-fold trade-off among inference accuracy, computation cost, and communication overhead but ignore the heterogeneity of different modalities within multimodal DNNs. In this paper, we propose an RL-based edge-assisted multimodal inference scheme that optimizes model selection at modality level and edge-assisted policies, including collaborative servers and partition points for each feature extractor to perform multimodal DNNs on mobile devices. Based on the information complexity and historical influence of each modality, as well as real-time observations such as channel gain, and previous inference performance, the policy distributions are designed to maximize the utility, as a weighted sum of inference latency and energy consumption, and inference accuracy. Safe policy exploration further mitigates risks such as low inference accuracy, intolerable inference latency, and improper allocation of computational resources to specific modalities. We analyze the computational complexity affected by the number of model variants, edge servers, and partition points and derive performance bounds for inference latency, energy consumption, and utility under specific sample sizes and data rates. Experimental results show that the proposed schemes improve inference performance compared to benchmark schemes. Liang Xiao 0003, Chuxuan Wang, Zefang Lv, Yiwen Zhan 0002, Yilin Xiao 0001, Helin Yang |
IEEE Trans. Mob. Comput. | 2 |
| 2026 | RL-Based Anti-Jamming Maritime Communications for LLM InferenceabstractReinforcement learning (RL)-based anti-jamming communication schemes select the transmit power and channel to send images or point clouds, but have low transmission reliability due to lack of maritime environmental features and the delayed feedback under harsh channel conditions caused by sea surface reflections and wave fluctuations. In this paper, we propose an RL-based anti-jamming maritime communication scheme for LLM inference that enables user equipment (UE) to optimize uplink power, channel and data compression ratio for transmitting multi-modal data based on data size of each modality, number of received packets, environmental features such as weather conditions and jamming features such as the received jamming power with quality-of-service guarantee. The historical anti-jamming experiences are used to recover the delayed feedback or loss and reduce communication outage under harsh maritime channel. The upper bound in terms of latency, UE energy consumption and inference accuracy is provided based on a maritime anti-jamming game to show the impact of the number of modalities, bandwidth and channel gains. The proposed scheme is implemented based on UEs equipped with Jetson Orin, camera and BME280 sensors for transmitting 200-KB temperature, humidity and pressure as well as images to the control center for LLaVA-based environmental feature extraction and marine object detection. Experimental results based on 4 UEs at Harbor show the performance gain with 40.5% less latency, 46.4% lower UE energy consumption and 10.9% higher inference accuracy against the smart jammer. Liang Xiao 0003, Pengli Zhang, Haoyu Chen 0005, Zefang Lv, Manhao Jiang |
IEEE Trans. Wirel. Commun. | 2 |
| 2026 | Learning-Based Anti-Jamming Energy-Efficient Wide-Area CommunicationsabstractReinforcement learning (RL)-based low-power wide-area networks (LPWANs) such as Long Range (LoRa) enable the access point (AP) to optimize the uplink power and channel of user equipment (UE), but the energy efficiency and quality of service (QoS) are low due to frequent downlink feedback against smart jamming. In this paper, we propose an RL-based anti-jamming energy-efficient wide-area communication scheme to optimize the transmit power, channel and downlink feedback policy with QoS guarantee. The similarities of channel states and UE communication tasks are evaluated to update learning parameters and accelerate policy selection under low duty cycles in LPWANs. We further propose a deep RL-based scheme that exploits neural networks to address the quantization error of the received power at AP and UE battery level and update weights according to prioritized experience replay to improve learning efficiency. The computational complexity and the performance bound in terms of signal-to-noise ratio and UE energy consumption are provided based on the Stackelberg equilibrium of the anti-jamming game. Our proposed schemes are implemented based on Raspberry Pi and BME280 sensors to support the LoRa-based transmission of temperature, humidity and pressure data. Experimental results on campus show the performance gain with less packet loss rate and UE energy consumption over the benchmark against smart jamming. Liang Xiao 0003, Shuohua Wang, Zefang Lv, Yiwen Zhan 0002, Haoyu Chen 0005 |
IEEE Trans. Wirel. Commun. | 2 |
| 2025 | WiFi CSI Based Energy-Efficient Drone DetectionabstractNetwork interface card (NIC) enabled active drone detection that analyzes the WiFi channel state information (CSI) variations due to drone vibrations and propeller rotations, which depends on frequent channel estimation to obtain sufficient CSI samples to distinguish drone motion-related information such as Doppler effects. However, the detection is inaccurate against interference due to the degraded channel estimation performance, such as pilot symbol error, and the excessive CSI sampling rate resulting in high energy consumption. In this paper, we propose a reinforcement learning assisted WiFi CSI based energy-efficient drone detection against interference, which optimizes the signal sampling across NICs regarding the transmit channel and the CSI sampling rate as well as the detection threshold. Based on the received signal strength indicator and the sampled CSI data in terms of amplitude and phase, the detection policy is chosen to enhance the utility as a weighted sum of the drone detection accuracy and energy consumption. The filtering mechanism quantifies policy similarity through the Kullback-Leibler divergences of both sampled data distributions and detection performance metrics to eliminate redundant policies, thereby improving policy exploration efficiency. Experiment results based on two NICs controlled by a Raspberry Pi and a laptop to detect a DJI Mavic 3 in the outdoor environment against interference show the performance gain over the benchmark. Qiaoxin Chen, Changjin Yu, Liang Xiao 0003, Jieling Li, Yunjun Zhu, Liqing Ye |
GLOBECOM | 3 |
| 2025 | Reinforcement Learning Based Edge-Assisted Dynamic Inference for Mobile VisionabstractMobile device can deploy high-complexity vision tasks by reducing spatial redundancy in the inference of deep neural networks (DNNs) and offloading computations to remote servers via wireless channels. However, a trained DNN model with fixed network architecture faces challenges in dynamic scenarios with less overlap between previous and current frames. In this paper, we propose a RL-based edge-assisted mobile vision scheme based on dynamic DNN models to balance inference latency and capacity without relying on historical information for dynamic visual scenarios. This scheme optimizes the hyperparameter settings and complexity level of dynamic neural networks as well as the collaborative server and the partition points of the DAG structured inference model. Based on the image embeddings, channel gain, and previous inference performance, the inference policy is selected to optimize the utility function, which is a weighted sum of inference latency, local energy consumption, and inference accuracy. Experimental results demonstrate the performance improvements of our proposed scheme over the benchmarks. Liang Xiao 0003, Tuhao Li, Zefang Lv, Ziyue Qiao, Hui Xiong 0001 |
GLOBECOM | 2 |
| 2025 | RF Distillation Diffusion Model: An Efficient RFF Data Augmentation MethodabstractRadio Frequency Fingerprint (RFF) based physical layer authentication technology provides enhanced security for wireless communications. However, the spatiotemporal overlap of wireless signals makes it challenging to label wireless device samples. Moreover, generative networks, such as Generative Adversarial Networks (GANs) struggle to retain the subtle signal features. Utilizing existing unlabeled samples poses a significant challenge for RFF data augmentation. This paper proposes the RF Distillation Diffusion (RFDD) model, an efficient method for RFF data augmentation. RFDD employs a conditional diffusion model to generate high-quality RF signals, which fully utilizes existing unlabeled samples to learn the data distribution of signals, and labeled samples are used to enhance the capability of RFF feature extraction. Additionally, knowledge distillation is utilized to improve sampling efficiency. Experimental results show that the RFDD can fully use 20% unlabeled samples to generate high-quality synthetic RF signals within only 0.45 seconds and improve RFF identification accuracy by 25.63% with 40% synthetic signals under SNRs from -5 to 5 dB. The code is available at https://github.com/XMU-Kai/RFDD Caidan Zhao, Jingqian Chen, Liang Xiao 0003 |
ICASSP | 4 |
| 2025 | Reinforcement Learning Based UAV Swarm Enabled 3-D Multimodal Jamming DetectionabstractReinforcement learning based jamming detection that chooses the test threshold to evaluate the received signal strength indicator (RSSI) and the packet loss rate is inaccurate for unmanned aerial vehicles (UAVs) lacking target indication against smart jammers. In this paper, we propose a UAV swarm-enabled multimodal jamming detection scheme to optimize the test thresholds based on vision information such as object classes and relative distances, along with the RSSI, the channel gain and the communication performance, including packet delivery ratio, bit error rate and packet delivery delay. A machine learning classifier is used to assess the joint variation among RSSI, channel gain and communication performance, with the output outlier score compared with the test threshold to detect the jammer. The detection results shared from neighboring UAVs are exploited in the update of policy distribution to refine jamming signal resolution and thus reduce the miss detection rate. The utility bound is derived based on the Nash equilibrium of the jamming detection game between the UAV swarm and the jammer. Experimental results based on 5 UAVs to detect a smart jammer show that our proposed scheme enhances the detection accuracy compared with the benchmarks. Qiaoxin Chen, Liang Xiao 0003, Jieling Li, Yuxiao Ren, Zefang Lv, Hongbin Jin |
ICC | 3 |
| 2025 | Reinforcement Learning Based Anti-Jamming FANET Routing with QoS GuaranteeabstractReinforcement learning (RL) based flying ad-hoc network (FANET) routing enables unmanned aerial vehicles (UAVs) to choose the next-hop, but the quality of service (QoS) and energy efficiency have to be enhanced against jamming due to the inaccurate path quality estimation. In this paper, we propose an RL based anti-jamming FANET routing with QoS guarantee to optimize the transmit power and the originator message broadcast interval for the route discovery to estimate the path quality to select the next-hop from the routing table. Based on the path availability history, the channel conditions, the received jamming power and the transmission quality regarding the number of the received originator messages and the bit error rate during route discovery, the routing policy is selected to enhance the throughput and the energy efficiency. The backward estimation of the throughput in the utility evaluation addresses the delayed feedback from the destination. The performance bound is derived in terms of network topology and channel gain based on the Nash equilibrium of the cooperative game among the UAVs. Simulation results provide the performance gain of the throughput and energy consumption over the benchmarks. Jieling Li, Chuxuan Wang, Liang Xiao 0003, Zefang Lv, Pengli Zhang, Helin Yang |
ICC | 3 |
| 2025 | RL-based Anti-Jamming Maritime Communications for Multi-Modal PerceptionabstractReinforcement learning (RL)-based communication scheme selects the transmit power and channel to improve communication reliability, but energy consumption and transmission delay degrades due to multi-modal data transmission such as images and point clouds for object detection under harsh maritime channel against jamming. In this paper, we propose a RL-based anti-jamming maritime communication scheme for multi-modal perception that optimizes the uplink transmit power, channel and compression ratios of each modality to improve perception accuracy and reduce communication energy consumption and transmission delay. The transmission policy of multi-modal data is determined by the probability distribution based on the data size and maritime environmental features such as rainfall and wind speed. The upper bound in terms of transmission delay and communication energy consumption is provided based on the Nash equilibrium of anti-jamming perception game to show the impact of the number of modalities and channel gain. Simulation results for ship detection based on the images and radar point clouds show the performance gain over benchmark against the smart jammer. Liang Xiao 0003, Pengli Zhang, Haoyu Chen 0005, Zefang Lv |
VTC2025-Fall | 3 |
| 2025 | Learning-Based Low-Latency Collaborative Inference for Multi-Branch Models in D2D-Assisted MECabstractDeploying high-complexity deep neural network (DNN) on mobile devices presents significant challenges, stemming from the conflict between their computationally intensive demands and the constrained computational resources. Device-to-device ($D$2$D$) assisted mobile edge computing (MEC) exploits sharing resources among devices to perform DNN inference collaboratively for higher efficiency. However, most existing collaborative inference schemes ignore the structure of DNN models, thus suffering from high latency under dynamic network conditions and computational loads. In this paper, we propose a learning-based low-latency collaborative inference scheme for multi-branch models, which schedules the computations of each DNN layer to optimize the alignment between the DNN structure's characteristics and the D2D network environments. Specifically, the computational task of each branch is assigned to nearby mobile devices or the edge server for efficient collaboration. In addition, we design proximal policy optimization with task-specific architecture to choose the computational scheduling policy, which contains a two-stream NN to extract features from the D2D network and the tasks at each layer, as well as a multi-head output to obtain the collaborative devices for all tasks in the layer. Simulation results show that the proposed scheme reduces the inference latency compared with benchmarks. Yilin Xiao 0001, Xiaozhen Lu, Liang Xiao 0003 |
WCNC | 4 |
| 2025 | Reinforcement-Learning-Based APT Defense for Large-Scale Smart GridsabstractReinforcement-learning-(RL)-based advanced persistent threats (APTs) defense schemes choose the scan interval to enhance the detection accuracy, and the data protection level can be further improved by optimizing the repair rate corresponding to methods, such as malicious files removal, security software version update and passwords reset, each requiring different computational resources, such as CPUs. In this article, we propose an RL-based APT defense scheme to optimize both the continuous scan interval of metering data and repair rate to mitigate the potential loss for meter data management systems (MDMSs) in smart grids with large-scale meters. Based on the size of the metering data stored at MDMS and the compromised data, the data tag granularity and the number of CPUs for repairing, neural networks extract the state feature, address the quantization error of defense policy, and update the weights based on the APT defense experiences and the shared weights of neighboring MDMSs to improve the defense performance as a weighted sum of the defense duration, detection accuracy and data protection level. The computational complexity of the proposed scheme and the performance bounds according to the Nash Equilibrium of the game between the MDMS and the APT attacker are provided. Simulation results based on three MDMSs that receive the metering data from 50–150 smart meters and an APT attacker with the selected attack interval up to 5 s show the performance gain over the benchmark based on the optimal control theory and Q-learning in large-scale smart grids. Liang Xiao 0003, Zefang Lv, Zhiping Lin 0002, Yousong Du |
IEEE Internet Things J. | 1 |
| 2025 | Reinforcement Learning-Based Accurate Worm Detection for Smart GridsabstractReinforcement learning (RL) based worm detection chooses the test threshold to evaluate the network traffic features such as the spectral flatness measure (SFM), but the detection of the evasive worm that modifies the scan rate and the worm propagation speed to manipulate the network traffic features is inaccurate due to the estimation and quantization error in the test threshold. In this paper, we propose an RL based accurate worm detection for smart grids that enables the control center to optimize the test threshold based on the number of meters, the infection time series and the number of connections to new destination IP addresses, besides the traffic log size received from each data concentrator and the number of the previously infected meters. A constraint on the maximum missed detection rate required by the smart grids is exploited in the detection policy distribution to support the reliable data transmission. A deep RL version addresses the quantization error in terms of the test threshold in the SFM evaluation and the infection time series in the state formulation, and compress the state space of the traffic log size received from a large number of data concentrators. Based on a worm detection game, the performance bound is provided under the specified detection window size and the propagation speed of evasive worm. Simulation results for 3600 meters show that the performance gain of the detection accuracy and latency against evasive worm over the benchmarks. Liang Xiao 0003, Jieling Li, Yilin Xiao 0001, Zefang Lv, Chuxuan Wang, Pengmin Li |
IEEE Internet Things J. | 1 |
| 2025 | Learning-Based Energy-Efficient Anti-Jamming FANET Routing With QoS GuaranteeabstractReinforcement learning (RL) based flying ad-hoc network (FANET) routing enables unmanned aerial vehicles (UAVs) to choose the next-hop to forward the packets, but the quality of service (QoS) and energy efficiency have to be enhanced due to the inaccurate path quality estimation under jamming attacks. In this paper, we propose an RL based energy-efficient anti-jamming FANET routing scheme with QoS guarantee to optimize both the transmit power and the originator message broadcast interval for the route discovery to estimate the path quality to select the next-hop from the routing table. Based on the path availability history, the channel conditions and the received jamming power, as well as the transmission quality regarding the number of the received originator messages and the bit error rate during route discovery, the routing policy is selected to enhance the throughput and the energy efficiency under jamming attacks with changing power. The backward estimation of the throughput in the utility evaluation addresses the delayed feedback from the destination under large-scale networks. The deep neural networks are further designed to address the quantization error of the transmission quality and the channel gain for UAVs with high mobility to enhance the path exploration efficiency. In addition, the upper bound in terms of network topology and channel gain is derived based on the Nash equilibrium of the anti-jamming routing game. The proposed routing scheme is implemented to improve the image transmission quality against jamming in outdoor environments. Experimental results based on UAVs equipped with Raspberry Pi show the performance gain of the throughput and the energy consumption. Jieling Li, Liang Xiao 0003, Chuxuan Wang, Zefang Lv, Pengli Zhang, Helin Yang |
IEEE Trans. Commun. | 2 |
| 2025 | Blockchain-Enabled Secure Offloading for VEC: A Multi-Agent Reinforcement Learning ApproachabstractVehicular edge computing (VEC) helps improve the task computational performance of vehicles on roads but has difficulty in defending against eavesdropping and selfish attacks simultaneously. In this paper, we design a reputation-based smart contract with blockchain and propose a multi-agent reinforcement learning (RL) based secure offloading scheme for VEC against both eavesdropping and selfish attacks. This scheme has a three-level hierarchical structure for each vehicle and uses the reputations obtained from the blockchain as the basis to optimize the edge node selection, offloading ratio, and power allocation, which aims to reduce the task computational latency, the vehicle energy consumption and eavesdropping rate. By using a punishment function based on the constraints, this scheme avoids exploring dangerous policies that can cause task failure or severe data leakage. A multi-agent deep RL-based secure offloading scheme is proposed for vehicles with sufficient resources, which evaluates the long-term risk rather than the punishment function to further improve the secure offloading performance. The regret bound is analyzedand the cumulative reward upper bound is provided. Simulation results verify the effectiveness of our schemes as compared with the benchmark. Xiaozhen Lu, Liang Xiao 0003, Yilin Xiao 0001, Zehui Xiong, Zhe Liu 0001, Yanyong Zhang, Weihua Zhuang |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | Reinforcement Learning-Based Personalized Differentially Private Federated LearningabstractDue to the different privacy and local model quality requirements for each participant, federated learning (FL) is vulnerable to membership inference attacks. To solve this issue, we propose a risk-aware reinforcement learning (RL)-based personalized differentially private FL framework. This framework uses local model accuracy and privacy loss as the constraints to satisfy the user’s personalized requirements. By designing a multi-agent RL, this framework optimizes perturbation policy including perturbation mechanisms and parameters (such as privacy budget and probabilistic relaxation). The goal of each participant is to improve global accuracy and reduce privacy loss, attack success rate, and short-term risk value. Firstly, the framework designs a two-level hierarchical policy selection module to choose the perturbation policy to accelerate learning speed. Secondly, our proposed framework designs a punishment function to evaluate short-term risk and an R-network to estimate long-term risk, which guarantees safe exploration. Thirdly, this framework formulates an improved Boltzmann policy distribution to increase the impact of risk, thus avoiding risky policies that may cause severe privacy leakage or local task failure. We also analyze the convergence performance and provide privacy analysis for both Gaussian and Laplace mechanisms. Experimental results based on the MNIST dataset demonstrate the effectiveness of our framework compared with benchmarks. Xiaozhen Lu, Liang Xiao 0003, Huaiyu Dai |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Reinforcement Learning-Based False Data Injection Attacks in Smart GridsabstractFalse data injection (FDI) attacks construct attack vectors to inject false data into tampered meters with the goal of falsifying state estimation, but resulting in low successful attack rate with high attack costs in terms of the number of tampered meters in large-scale smart grids, because the bad data detection at the control center chooses the dynamic detection thresholds to identify the modified meter measurements. In this article, we propose a reinforcement learning-based FDI attack scheme that optimizes both the tampered meters and the false data to enhance the success attack rate and injected errors while reducing attack costs. Based on meter measurements and previous performance, the attack vector is constructed to induce more errors in state estimation and bypass bad data detection. The performance bounds regarding the successful attack rate and the injected error are derived in terms of the number of bus phase angles, the susceptance of the transmission line, and the maximum false data based on the Nash equilibrium of the FDI game. Simulations performed on both the IEEE 14-bus and IEEE 118-bus systems demonstrate the performance gain over the benchmarks. Liang Xiao 0003, Haoyu Chen 0005, Zefang Lv, Chuxuan Wang, Yilin Xiao 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | Collaborative Perception Against Data Fabrication Attacks in Vehicular NetworksabstractCollaborative perception in vehicular networks enables the connected autonomous vehicle (CAV) to gather sensing data, such as feature maps of light detection and ranging (LiDAR) point clouds, from neighboring CAVs to achieve higher perception accuracy, which has performance degradation against data fabrication attacks that share falsified sensing data with random probability. In this paper, we exploit the spatial consistency check to detect the potentially manipulated regions in LiDAR point clouds and measure the inconsistency degree of the received sensing data based on the number of conflict regions, which is the basis for determining the falsified sensing data if the inconsistency degree exceeds the threshold of the hypothesis test. The reinforcement learning (RL)-based collaborative vehicular perception scheme against data fabrication attacks is further proposed to choose CAVs based on the inconsistency degrees, the data quality measured by the confidence scores, the channel gains and the CAV reputations, which enhances the utility as the weighted sum of perception accuracy, speed and minimum latency requirement for data transmission. In addition, the multi-layer perceptron-based neural networks extract the perception features of sensing data from historical experiences, such as the data quality of received feature maps, as well as compress the RL state that linearly increases with the network scales and the spatial granularity of LiDAR point clouds for faster learning. Experimental results based on 10 CAVs equipped with LiDAR sensors and NVIDIA computational units to detect 20 vehicles against data fabrication attacks show that our proposed scheme outperforms the benchmarks in terms of perception accuracy and speed. Zhiping Lin 0002, Liang Xiao 0003, Zefang Lv |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | Multi-UAV-Assisted MEC in Internet of Vehicles With Combined Multi-Modal Semantic Communication Under Jamming AttacksabstractSemantic communication technology, which transmits only relevant semantic information, can significantly conserve communication resources and reduce service time. This technology is particularly promising for unpilotedaerial vehicle (UAV)-assisted mobile edge computing (MEC) in the internet of vehicles (IoV). However, integrating semantic communication with UAV-assisted vehicle MEC is susceptible to malicious jamming. This paper introduces a reliable communication method that combines multi-modal semantic communication with UAV-assisted vehicle MEC to minimize delays in communication and computation while maintaining semantic accuracy during jamming attacks. Our approach optimizes UAV trajectories, user associations, and channel selections, enabling the UAV to select optimal positions when associating with different modal users and reducing the impact of jammers during multi-modal task reception. Due to the non-convex nature of the optimization problem and the highly dynamic environment, we employ the semantic communication combined with the multi-agent twin delayed deep deterministic policy gradient (SC-MA-TD3) approach, a multi-agent deep reinforcement learning (DRL) strategy that fosters UAV cooperation for efficient resource allocation. Simulation results show that our approach outperforms existing approaches in reducing delays and enhancing semantic accuracy. Shuai Liu 0019, Helin Yang, Mengting Zheng, Liang Xiao 0003 |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Blockchain-Based Intelligent Trusted Computational Resource Allocation for Low-Altitude NetworksabstractIn low-altitude networks, unmanned aerial vehicles (UAVs) can offer services such as logistics, intelligence surveillance, and environmental monitoring, aided by base stations (BSs) with substantial computational resources. However, BSs must defend against malicious UAVs that may overload resources or launch denial-of-service attacks. In this paper, we formulate a blockchain-enabled access control model, which uses the UAV identities (IDs) and trajectories, positive and negative interactions with the BS to evaluate the reputations of UAVs. In the blockchain, the elected miner generates blocks containing UAV IDs, coordinates, interactions, and reputation values. To defend against malicious UAVs, this paper formulates a trusted computational resource allocation optimization problem, solved by safe reinforcement learning (RL) with a three-level hierarchical structure. Specifically, this method uses the designed structure to optimize the BS access control, resource allocations, and block size. In particular, we design an E-network to evaluate the long-term risk resulting from the chosen policy, which is used to refine the policy distribution for safe exploration. A modified reward function accounts for immediate risks, preventing short-term dangerous explorations that could lead to illegal access or computational failures. We prove the Lyapunov asymptotic stability of the proposed system and derive the reward upper bound. Simulation results show that our scheme can converge to the upper bound, outperform the benchmark, and validate the effectiveness via ablation experiments. Xiaozhen Lu, Qihui Wu 0001, Liang Xiao 0003 |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | Edge-Assisted Collaborative Perception Against Jamming and Interference in Vehicular NetworksabstractCollaborative perception of connected autonomous vehicles (CAVs) that offload the sensing data, such as the feature map extracted from light detection and ranging (LiDAR) point clouds, to an edge device such as the roadside unit (RSU) to detect traffic objects has severe performance degradation due to the offloading latency and packet loss rate (PLR) under jamming and interference. In this paper, we propose an edge-assisted reinforcement learning (RL)-based collaborative perception scheme for CAVs to enhance the accuracy and speed against jamming and interference in LiDAR-based object detection. Based on the spatial confidence score of the feature map, the data size, the channel gains, the received jamming power and interference level, this scheme chooses the critical regions of the feature map, radio channel and transmit power with the hierarchical structure to enhance the learning efficiency. The risk level of the selected policy evaluates the time asynchronization and information loss of the shared feature map using the multi-level risk function based on multiple thresholds of the offloading latency and PLR, with assigning different penalties to mitigate the selection of high-risk policies that degrade perception performance. The upper performance bound in terms of the perception accuracy, latency and utility is provided based on the Stackelberg equilibrium of the game between the jammer and CAVs. Experimental results based on the Robosense RS-LiDAR-16 sensors and the Raspberry Pi to detect 10 vehicles in an$8.5\times 4\times 3.5$m3area show the performance gain with 22.4% higher perception accuracy and 41.3% less latency compared with the benchmark against a smart jammer. Zhiping Lin 0002, Liang Xiao 0003, Zefang Lv, Yunjun Zhu, Yanyong Zhang, Yong-Jin Liu 0001 |
IEEE Trans. Wirel. Commun. | 2 |
| 2024 | Reinforcement Learning Based Collaborative Perception for Vehicular NetworksabstractReinforcement learning (RL)-based collaborative perception in vehicular networks chooses the sub-frame of radio channel resources for connected autonomous vehicles (CAVs) to exchange sensing data to enhance the perception performance, but leads to inaccurate detection in the light detection and ranging (LiDAR)-based object detection due to the asynchronous scan period of the LiDAR point clouds. This paper proposes a RL-based collaborative perception scheme to choose the transmit power and sub-frame to share the feature maps extracted from point clouds. Based on the estimated packet timestamp, the network topology and the channel gains among CAVs, this scheme enhances the perception accuracy and latency against path-loss and interference. The collaborative risk in the policy distribution is formulated as a weighted sum of the perception latency and packet loss rate to avoid the time asynchronization and information loss of the feature map exchange. The performance bound of the perception accuracy and latency is provided based on a Nash equilibrium of the cooperative game among CAVs. Simulation results based on five CAVs show the performance gain of the perception accuracy and latency over the benchmarks. Zhiping Lin 0002, Yunjun Zhu, Jieling Li, Liang Xiao 0003, Yuliang Tang, Yanyong Zhang |
GLOBECOM | 5 |
| 2024 | Joint Channel Selection and Power Control for Multi-UAV-Enabled Anti-Jamming Communications Based on Game Guided Reinforcement LearningabstractUnmanned aerial vehicles (UAVs) have been widely employed as airborne base stations to enhance terrestrial communications. However, the growing demand for communications, spectrum and energy resources are increasingly in short supply, while malicious jamming from jammers threatens the communications reliability. To address these challenges, we propose a joint reliable channel selection and power control approach for multi-UAV-enabled communications networks under malicious jamming attacks, with the goal of maximizing user communication capacity under limited energy constraint and avoiding malicious jamming from jammers. Due to the dynamic and time-varying nature of communication environments, we propose an intelligent resource optimization algorithm based on game theory guided reinforcement learning. To be specific, we employ a hierarchical learning algorithm based on the Stackelberg game to help users in the follower layer cooperatively select channels to against co-channel interference and jamming, and develop a deep reinforcement learning-based algorithm for dynamic power control to maintain communication efficiency. Simulation results demonstrate that our proposed approach can significantly improve the user communication rate and achieves faster convergence compared with existing algorithms. Helin Yang, Changyuan Xu, Ziling Shao 0001, Liang Xiao 0003, Yifu Jiang, Zehui Xiong |
GLOBECOM | 5 |
| 2024 | Reinforcement Learning Based Energy-Efficient Anti-Jamming NB-IoT CommunicationsabstractNarrowband Internet of Things (NB-IoTs) with limited power supply and long range requirements can apply reinforcement learning (RL) based anti-jamming techniques to choose the transmit power and channel to enhance the communication reliability, but have low energy efficiency for the energy-constrained user equipment (UE). In this paper, we propose a RL-based energy-efficient anti-jamming NB-IoT communication scheme to choose the transmission policy including the UE transmit channel and power as well as whether to send them to UE over the downlink channel. Based on the received power on each channel, UE battery level and message size, the transmission policy distribution is formulated according to the risk value that indicates communication reliability exceeding the narrowband bit error rate (BER) requirement and the long-term utility as the weighted sum of UE energy consumption, BER level and frequency hopping overhead for higher energy efficiency and reliability. The similarity of the state-action pairs is evaluated using the risk value to update the long-term risk levels of feasible policies via the knowledge reuse technique for faster learning. The computational complexity and the upper bounds of our proposed scheme in terms of BER and energy consumption are provided. Simulation results show the performance gain over the benchmark against jamming. Shuohua Wang, Zefang Lv, Jieling Li, Liang Xiao 0003 |
GLOBECOM | 5 |
| 2024 | Reinforcement Learning-Based Secure Video Transmission For IOV SystemsabstractThe rapid growth in the number of vehicles and types of services such as video transmission improves the quality-of-service (QoS) requirements and increases the difficulty in resisting eavesdropping attacks on the Internet of Vehicles (IoV). Existing video transmission schemes that either ignore the impact of eavesdropping attacks or have the full knowledge of the attack model have performance degradation in highly dynamic IoV systems. In this paper, we propose a reinforcement learning-based secure video transmission scheme for IoV systems, which jointly optimizes the access control policy for each vehicle (i.e., the selection of access nodes such as the base stations or unmanned aerial vehicles) and the corresponding transmit power level against active eavesdropping. This scheme uses the QoS and eavesdropping rate as the criteria to evaluate the long-term risk of each state-action pair, which is estimated by a designed deep Q-network to avoid the risky access control policies that cause severe data leakage or video transmission failure. Simulation results show that our scheme reduces the energy consumption, transmission latency, and eavesdropping rate compared with the benchmark. Xiaozhen Lu, Yanling Bu, Liang Xiao 0003 |
ICIP | 6 |
| 2024 | Reinforcement Learning based Edge-Assisted Inference for Maritime UAV NetworksabstractEdge-assisted inference that enables each unmanned aerial vehicle (UAV) to offload marine tasks for maritime applications such as target tracking and data collection, but the inference speed and energy efficiency are affected by sea surface movement and wave occlusions under broad area maritime environment with instability channel. In this paper, we propose an RL-based edge-assisted inference scheme for maritime UAV networks to optimize the deep neural networks partition point, the transmit power and the collaborative edge server to enhance the utility as the weighted sum of the data-related energy consumption and the inference latency. Based on the average sea wave height, the channel gain, the number of marine tasks and the battery level, the self-correcting mechanism makes a trade-off between the overestimation and the underestimation in collaborative inference policy without additional computational cost. The bounds of the data-related energy consumption and the inference latency are derived under the specific data rate and inference computation amounts. Simulation results show the effectiveness of the proposed edge-assisted inference scheme. Chuxuan Wang, Jieling Li, Liqing Ye, Liang Xiao 0003 |
ISPA | 5 |
| 2024 | Intelligent Energy-Efficient and Fair Resource Scheduling for UAV-Assisted Space-Air-Ground Integrated Networks Under Jamming AttacksabstractThe space-air-ground integrated network (SAGIN) is a crucial technology for sixth-generation (6G) wireless communication networks to achieve seamless coverage and high throughput. In this paper, we propose an unmanned aerial vehicle (UAV)-assisted SAGIN structure, where the UAV is responsible for collecting data from ground users (GUs) and transmitting it to low-earth orbit (LEO) satellites. This paper also formulates a joint energy-efficient and fair resource scheduling optimization problem under jamming attacks and limited energy constraints, where the line-of-sight (LoS) links between the UAV and GUs are susceptible to being jammed. Due to the non-convex problem and dynamic environments, a deep reinforcement learning (DRL)-based twin delayed deep deterministic policy gradient (TD3) is developed to search optimal UAV trajectory to maximize energy efficiency (EE) and fairness against jamming. Simulation results verify that the proposed intelligent resource scheduling algorithm outperforms the baseline algorithms in terms of EE and fairness index in different settings. Shihao Chen, Helin Yang, Liang Xiao 0003, Changyuan Xu, Xianzhong Xie, Zehui Xiong |
VTC Spring | 3 |
| 2024 | Reinforcement Learning Based Interference Coordination for Port CommunicationsabstractReliable port communications support maritime applications such as vessel navigation and cargo tracking for a large number of mobile users on ships, but the quality of services (QoS) such as data rate and energy consumption is severely degraded by inter-cell interference. In this paper, we propose a deep reinforcement learning (RL)-based interference coordination scheme for port communications to reduce the transmission latency and energy consumption, and improve the data rate. Based on the signal-to-interference plus noise ratio, the channel gains, the estimated interference levels and the transmission latency, the base station chooses the transmit power and downlink bandwidth constraint to avoid choosing risk policies that cause the communication performance degradation. In addition, a two-level hierarchical structure with two convolution networks and four fully connected layers is designed to reduce the algorithm complexity and enhance the convergence speed. Simulation results verify the performance gain of the proposed scheme in terms of the data rate, the transmission latency, and the energy consumption compared with the benchmark. Siyao Li, Chuhuan Liu, Liang Xiao 0003, Helin Yang |
VTC Spring | 4 |
| 2024 | Energy-Efficient Resource Management for Multi-UAV NOMA Networks Based on Deep Reinforcement LearningabstractCellular-connected unmanned aerial vehicles (UAVs) play an essential role in cellular networks. Combined with non-orthogonal multiple access (NOMA) technique, UAVs can provide better performance in various communication scenarios. In this paper, we investigate a NOMA-enhanced UAV-assisted cellular network where multiple UAVs are deployed as aerial base stations to provide communication services for mobile ground users in the presence of a malicious jammer. We propose a two-step learning-based resource scheduling approach. First, an algorithm based on K-means clustering is proposed to partition ground users (GUs) to reduce mutual interference. Moreover, a cooperative multi-agent twin delayed deep deterministic algorithm is proposed to jointly optimize UAVs' trajectories, power allocation and GU association to maximize the system energy efficiency (EE) while guaranteeing minimum quality-of-service (QoS) requirements. Extensive results demonstrate that the proposed solution can efficiently improve EE and QoS performances under jamming attacks compared with existing popular approaches. Xiangda Lin, Helin Yang, Kailong Lin, Liang Xiao 0003, Zhaoyuan Shi, Zehui Xiong |
VTC Spring | 4 |
| 2024 | Reinforcement Learning Based Jamming Detection for Reliable Wireless CommunicationsabstractBoth the radio spectrum features such as power spectral density (PSD) and the communication performance such as packet loss rate (PLR) can be exploited to detect jamming attacks, with the resulting detection results used to enhance the reliability of wireless communications. In this paper, we propose a reinforcement learning (RL)-based jamming detection scheme based on the channel energy, the received signal strength indi-cator of each packet, the channel gains, PLR and transmission latency of mobile devices, in which the test threshold and the number of PSD bins are optimized by access point to enhance the utility as a weighted function of the detection speed and accuracy. The detection results are exploited for mobile devices to choose the transmit power and channel to reduce the PLR and transmission latency. Experimental results based on the universal software radio peripheral and Raspberry Pi to detect four jamming types including constant, sweeping, random and smart jamming show that our proposed schemes improve the detection accuracy and speed, as well as the communication performance. Zhiping Lin 0002, Qiaoxin Chen, Liang Xiao 0003 |
VTC Spring | 5 |
| 2024 | Efficient Normalizing Flow-Based Radio Frequency Fingerprinting Identification for Network SecurityabstractDevice authentication plays a key role in securing Internet of Things (IoT), where radio frequency fingerprinting (RFF) identification is an emerging physical layer security technique by exploiting intrinsic and unique hardware impairments of wireless devices. However, recent works mainly focus on the identification of authorized devices, while neglecting the harm misidentification of illegal devices. Thus, this paper proposes a normalizing flow-based RFF method (NFRFF) for illegal device identification. The proposed NFRFF designs parallel flows and a fusion flow to handle the distributions of radio frequency signal samples and employs the dual attention mechanism to enhance the efficiency of information fusion and the ability to recognize illegal wireless devices. Ultimately, NFRFF is capable of assigning lower likelihoods to signal samples from illegal devices and higher ones to those from authorized devices. Simulation results verify that after extracting multi-scale radio frequency feature maps using VGG-16, NFRFF achieves an AUROC of 0.991 under 120 authorized devices and 30 illegal devices, achieving excellent identification performance. Weiwei Zeng, Helin Yang, Kailong Lin, Liang Xiao 0003 |
VTC Fall | 4 |
| 2024 | Risk-Aware Reinforcement Learning Based Federated Learning Framework for Io VabstractFederated learning helps protect data privacy for Internet of vehicles (Io V) by selecting a number of participated nodes but suffers from performance degradation such as low model training accuracy in the highly dynamic and large-scale Io V systems under selfish attacks. In this paper, we propose a risk-aware reinforcement learning based federated learning framework against selfish attacks for Io V,which jointly optimizes the training policy (i.e., the selection of participated vehicles and the corresponding local training data size) based on the state including the global model training accuracy, local model quality, training latency, data rate, and participation rate. By designing a punishment function to evaluate the immediate risk of each choosing training policy, this scheme avoids risky policies that result in extremely low training accuracy and high training latency to satisfy the requirements of local tasks such as the quality of service requirements. An evaluated neural network involved fully connected layers is designed to fast extract the global and local training features and thus accelerate the convergence speed. Experimental results based on both the MNIST and CIFAR-10 datasets verify that our scheme outperforms the benchmarks with higher training accuracy and less training latency. Xiaozhen Lu, Liang Xiao 0003 |
WCNC | 4 |
| 2024 | Reinforcement Learning Based QoS-Aware Anti-Jamming Underwater Video TransmissionabstractUnderwater video transmission has to ensure quality-of-service (QoS) against jamming with severe multipath effect and narrow bandwidth limitation that degrade the communication performance under variable channel state. In this paper, we propose a reinforcement learning (RL)-based QoS-aware underwater video transmission scheme to optimize the video compression ratio, modulation format and transmit power based on the state consisting of the channel gain and previous transmission performance. This scheme evaluates the risk level that indicates the probability of failing the QoS and the long-term expected utility of each transmission policy under the current state to improve the anti-jamming communication performance. We derive the performance bound of the utility and analyze its relationship with transmission policy. Simulation results illustrate that our scheme improves the QoS by reducing the frame loss rate (FLR), transmission delay and increasing spatial-spectral entropy-based quality (SSEQ) compared with the benchmark. Shaoxuan Li, Zefang Lv, Liang Xiao 0003, Wei Su 0002 |
WCNC | 5 |
| 2024 | Reinforcement Learning Based Energy-Efficient Fast Routing for FANETsabstractReinforcement learning (RL) based flying ad-hoc network (FANET) routing enables unmanned aerial vehicles (UAVs) to choose the next-hop to increase the packet delivery ratio, but the routing latency and energy consumption have to be further reduced over inaccurate feedback for large-scale networks. In this paper, we propose an RL based energy-efficient fast routing for each UAV to choose the forwarding decision and the power. Based on the state consisting of the battery level, channel conditions and forwarding decisions of the one-hop neighbors, the routing policy is chosen to enhance the utility as the weighted sum of the delivery success indicator, the latency and the energy consumption. The number of the latency violations and the learning parameters shared among the one-hop neighbors are exploited in the update of the routing policy distribution following the latency constraint with the reduced energy consumption. The deep neural networks address the state quantization error of the latency and the channel gain for UAVs with high mobility under large-scale networks. The performance bound regarding the end-to-end latency and the energy consumption is derived in terms of network topology and channel gain based on the packet forwarding game. The performance gain over the benchmark is provided via both simulation and experimental results. Jieling Li, Liang Xiao 0003, Xuchen Qi, Zefang Lv, Qiaoxin Chen, Yong-Jin Liu 0001 |
IEEE Trans. Commun. | 2 |
| 2024 | Learning-Based Resource Management Optimization for UAV-Assisted MEC Against JammingabstractIn recent years, jointly optimizing unmanned aerial vehicle (UAV) hover point selection and resource management for UAV-assisted mobile edge computing (MEC) is a hot research topic. Unlike previous studies, this paper investigates the optimization problem of hover point selection and resource management under dynamic jamming attacks, where the objective is to maximize overall communication and computing efficiency while taking into account constraints on total UAV power and the availability of channels. Due to the non-convex problem and highly dynamic environments, we then propose an advanced deep reinforcement learning (DRL) algorithm to jointly optimize UAV hover point selection, task collection time ratio, transmission power, channel selection, and task offloading ratio to improve the efficiency of UAV-assisted MEC. Specifically, the algorithm optimizes UAV hover point selection to minimize the negative effect of jamming attacks, and then manages resources to improve UAV task processing capacity and reduce energy consumption while mitigating jamming. Simulation results demonstrate that our proposed learning-based algorithm significantly enhances the computing and offloading efficiency in complex and dynamic UAV-assisted MEC environments against jamming compared to other existing algorithms. Shuai Liu 0019, Helin Yang, Liang Xiao 0003, Mengting Zheng, Huabing Lu, Zehui Xiong |
IEEE Trans. Commun. | 3 |
| 2024 | Personalized 3D Location Privacy Protection With Differential and Distortion Geo-PerturbationabstractThe rapid development of indoor location-based services (LBS) has raised concerns about location privacy protection in the 3-dimensional (3D) space. The existing 2-dimensional (2D) location privacy protection mechanisms (LPPMs) cannot effectively resist attacks in 3D environments. Furthermore, users may have various sensitive attributes at different locations and times. In this paper, we first formally study the relationship between two complementary notions of geo-indistinguishability and distortion privacy (i.e., expected inference error) in the 3D space and develop a two-phase personalized 3D LPPM (P3DLPPM). In Phase I, we search for neighboring locations to formulate a protection location set (PLS) for hiding the actual location based on the above-mentioned relationship. To realize this, we develop a 3D Hilbert curve-based minimum distance searching algorithm to find the PLS with minimum diameter for each location while guaranteeing differential privacy. In Phase II, we put forth a novel Permute-and-Flip mechanism for location perturbation, which maps its initial application in data publishing privacy protection to a location perturbation mechanism. It generates fake locations with smaller perturbation distances while improving the balance between privacy and quality of service (QoS). Simulation results show that the proposed P3DLPPM can significantly improve personalized privacy protection while meeting the user's QoS needs. Minghui Min, Haopeng Zhu, Jiahao Ding, Shiyin Li, Liang Xiao 0003, Miao Pan, Zhu Han 0001 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2024 | Safe Multi-Agent Reinforcement Learning for Wireless Applications Against Adversarial CommunicationsabstractBased on the network observations and learning parameters shared by the neighboring learning agents, multi-agent reinforcement learning (RL) has to enhance the performance over adversarial communications, in which spoofing attackers send fake learning messages to fool the learning agent and thus degrade the performance of wireless applications. In this paper, we propose a safe multi-agent RL algorithm for wireless applications against adversarial communications, in which each learning agent chooses the cooperative agents to share the learning information and authenticates the received learning messages before integrating them into the RL state formulation and the learning parameter update. The communication policy distribution for the cooperative agent selection is formulated based on the long-term discounted reward and the sharing reputation for each neighboring agent, which is updated based on the authentication results to indicate the probability as a spoofing attacker. Neural networks are designed to estimate the long-term discounted reward and the sharing reputation for the learning agent with sufficient computational resources in large-scale wireless networks to enhance the agent selection security. As a case study, our proposed algorithm is implemented in the unmanned aerial vehicle swarm anti-jamming video transmission against spoofing attackers that send fake received jamming power as well as Q-values and neural network weights in the anti-jamming transmission policy learning. Both simulation and experimental results are provided to verify the performance gain over the benchmark. Zefang Lv, Liang Xiao 0003, Haoyu Chen 0005, Xiangyang Ji |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Risk-Aware Reinforcement Learning-Based Federated Learning for IoV SystemsabstractFederated learning (FL) that improves data privacy reduces the computational overhead for Internet of Vehicles (IoV) systems but has difficulty in defending against selfish attacks due to the restricted quality of service requirements and the high mobility of vehicles. In this paper, we design a risk-aware hierarchical reinforcement learning-based FL framework for IoV to resist selfish attacks. By designing a two-level hierarchical policy selection module that consists of two deep neural networks, this framework divides the training policy into two sub-policies, i.e., the selection of FL participants and the corresponding local training data size, which are chosen based on the previous training performance and vehicle participation performance. This framework designs a risk-aware safety guide to avoid dangerous states such as local task failure resulting from risky training policies. Specifically, the guide uses a warning signal to evaluate the short-term risk of each state-action pair, applies an R-network to estimate the long-term risks for modifying the chosen training policy, and designs a punishment function for the modified training policy to revise the immediate reward to further enhance the safe exploration. We analyze the convergence performance and computational complexity of our scheme. Experimental results on MNIST, CIFAR-10, and Stanford Cars datasets verify the effectiveness of our scheme, including the global model accuracy, training latency, detection success rate, and convergence speed compared with the benchmarks FedAvg, MFL, DQNPS, and SHRL. Xiaozhen Lu, Liang Xiao 0003, Wei Wang 0100, Qihui Wu 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Risk-Aware Federated Reinforcement Learning-Based Secure IoV CommunicationsabstractWith the rapid growth in the number of high-mobility vehicles and booming enhanced applications with restricted latency requirements, downlink communication in Internet of Vehicles (IoV) systems has become increasingly vulnerable to active eavesdropping attacks. This paper proposes a federated learning-enabled secure communication framework for IoV against active eavesdropping, in which the roadside units (RSUs) apply reinforcement learning (RL) model to optimize their downlink transmit power levels, and the server helps update the RL models of the RSUs. First, we design a multi-agent deep RL algorithm for each RSU, which designs a punishment and a blacklist mechanism to mitigate risky explorations related to severe data leakage or communication outages. Second, this framework designs a risk-aware RL for the server, which uses a two-level hierarchical structure to choose the number of participated RSUs and the corresponding local training data size for higher optimization speed. This framework considers both the reward and risk in the selection of policies to reduce the probability of exploring the risky training policies that cause defense failure of the RSUs against active eavesdropping. Third, we analyze the convergence performance, computational complexity, and reward upper bound, which reveals how the power constraint, radio bandwidth and data size affect the secure communication performance. Simulation and experimental results validate the effectiveness of our schemes, such as the reductions of the eavesdropping rate, training latency, and the loss of local models compared to the benchmarks. Xiaozhen Lu, Liang Xiao 0003, Yilin Xiao 0001, Wei Wang 0100, Nan Qi 0001, Qian Wang 0002 |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | Deep Reinforcement Learning-Based Resource Management for UAV-Assisted Mobile Edge Computing Against JammingabstractIn mobile edge computing (MEC) systems, multiple unmanned aerial vehicles (UAVs) can be utilized as aerial servers to provide computing, communication, and storage services for edge users, called UAV-assisted MEC, which has emerged as a promising technology to improve both the computing and communication performances. Unlike existing works without considering jamming attacks, we investigate a multi-UAV-assisted-MEC scenario under multiple malicious jammers and then propose a resource management approach with the objective of minimizing both the system energy consumption and latency. Due to the time-varying nature of communication environments, we design a multi-agent deep reinforcement learning (MADRL)-based resource management approach to dynamically adjust the CPU frequency, communication bandwidth, and channel access selection of UAVs to enhance the system performance against jamming attacks. On this basis, in order to enhance the algorithm learning efficiency, we propose a multi-agent twin-delayed deep deterministic policy algorithm in combination with the prioritized experience replay mechanism (PER-MATD3) to effectively search for the joint resource management strategy under high-dimensional state and action spaces, where the time-varying channel state information and imperfect attack behavior information are also effectively trained to improve the learning capacity and convergence speed. Simulation and experimental results verify that the proposed approach can significantly decrease the overall system latency (i.e., computing and communication latency) and energy consumption compared to other benchmark algorithms under different real-world settings. Ziling Shao 0001, Helin Yang, Liang Xiao 0003, Wei Su 0002, Zehui Xiong |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | UAV-Enabled Semantic Communication in Mobile Edge Computing Under Jamming Attacks: An Intelligent Resource Management ApproachabstractThe integration of semantic communication with mobile edge computing (MEC) has emerged as a prominent research area. In this paper, we explore a novel scenario where semantic communication is integrated with unmanned aerial vehicles (UAVs) to enhance MEC, particularly in the face of jamming attacks. Our research focuses on addressing the resource management challenge to minimize task completion time and maximize semantic spectral efficiency (SSE) while adhering to quality of service requirements and resource constraints. Given the non-convexity of this problem and the dynamic behavior of jamming attacks, this paper proposes a deep reinforcement learning (DRL) algorithm by jointly optimizing UAV trajectories, user associations, and channel selections against jamming. In detail, the proposed anti-jamming DRL-based resource management approach can effectively capture the jammer’s behavior, and learn to adjust semantic task and resource scheduling strategies with the objective to minimize the negative effect of jamming attacks on task offloading and semantic communication. Simulation results demonstrate that the proposed approach outperforms baseline algorithms in terms of task completion time and total SSE under different real-world settings. Shuai Liu 0019, Helin Yang, Mengting Zheng, Liang Xiao 0003, Zehui Xiong, Dusit Niyato |
IEEE Trans. Wirel. Commun. | 4 |
| 2024 | Efficient Communications in Multi-Agent Reinforcement Learning for Mobile ApplicationsabstractThe environment observations and learning experiences shared by the cooperative learning agents accelerate multi-agent reinforcement learning (MARL) with partial observations for mobile applications but the performance degrades due to the redundant and outdated observations under severe channel fading in wireless networks. In this paper, we propose an efficient communication scheme in MARL for mobile applications that enables each learning agent to optimize the cooperative agents and the learning parameters to integrate the shared information. The cooperative agents are chosen according to the learning environment observations, the channel states, and the task similarity with neighboring agents. The learning parameters are chosen based on the attention mechanism that exploits the correlation with the local observation to enhance the agent receptive field for efficient policy exploration. Neural networks with weights updated based on the learning factors determined by the task similarity are designed to further improve the learning efficiency. The performance bounds including the information gain from the learning agent cooperation, the communication cost and the utility are provided based on the Nash equilibrium of the cooperative MARL communication game. The proposed scheme is implemented in the anti-jamming video transmission of the unmanned aerial vehicle swarms to optimize the transmit channel and power and experimental results verify the performance gain over the benchmark. Zefang Lv, Liang Xiao 0003, Yousong Du, Yunjun Zhu, Shuai Han 0002, Yong-Jin Liu 0001 |
IEEE Trans. Wirel. Commun. | 2 |
| 2024 | Energy Harvesting UAV-RIS-Assisted Maritime Communications Based on Deep Reinforcement Learning Against JammingabstractWith the rapid development of maritime activities, efficient and reliable maritime communications have attracted ever-increasing attention, and mounting reconfigurable intelligent surface (RIS) on unmanned aerial vehicle (UAV), called UAV-RIS, can provide flexible and adaptable services for maritime communications. In this paper, we investigate a UAV-RIS-assisted maritime communication system under a malicious jammer, where a UAV-RIS is deployed to jointly adjust its placement and RIS surface elements to maximize the system energy efficiency (EE) and guarantee quality of service requirements against jamming attacks. In addition, an adaptive energy harvesting scheme is developed for information transmission (IT) and energy harvesting (EH) simultaneously to enhance the endurance of the UAV by deploying different IT times for each RIS element. Considering the non-convex optimization problem and highly complex maritime environments, an intelligent resource management approach based on deep reinforcement learning is proposed to jointly optimize the base station’s transmit power, placement of UAV-RIS, and RISs reflecting beamforming. Furthermore, hindsight experience replay is adopted to improve the learning efficiency and performance. The simulation results demonstrate that the proposed approach achieves the better EE and EH performances under different real-world settings compared with existing popular approaches. Helin Yang, Kailong Lin, Liang Xiao 0003, Zehui Xiong, Zhu Han 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2024 | Learning-Based Reliable and Secure Transmission for UAV-RIS-Assisted Communication SystemsabstractMounting reconfigurable intelligent surface (RIS) on unmanned aerial vehicle (UAV), called UAV-RIS, combines the benefits of these two techniques, which can further improve the communication performance. However, high-quality air-ground channel links are more vulnerable to both the adversarial eavesdropping and the malicious jamming. Therefore, this paper proposes a reliable and secure communication approach assisted by the UAV-RIS to maximize the secrecy rate, while ensuring the quality of service (QoS) requirement of the legitimate user against both the eavesdroppers and the jammer. Specifically, with the imperfect channel state information and behaviors of mixed attacks, we try to maximize the achievable worst-case secrecy rate by jointly designing the transmit beamforming, artificial noise, UAV-RIS placement, and RIS’s passive beamforming. As the optimization problem is non-convex and the environment is highly dynamic, a post-decision state deep Q-network combined with Fourier feature mapping algorithm (called PDS-DQN-FFM) is further designed to effectively achieve the robust anti-attack transmission strategy. Simulation results demonstrate that our proposed learning based reliable and secure transmission approach significantly enhances both the secrecy rate and QoS satisfaction level as compared with existing approaches. Helin Yang, Shuai Liu 0019, Liang Xiao 0003, Yi Zhang 0035, Zehui Xiong, Weihua Zhuang |
IEEE Trans. Wirel. Commun. | 3 |
| 2023 | Multi-Agent Reinforcement Learning for Wireless Networks Against Adversarial CommunicationsabstractBased on the efficient and reliable exchange of learning messages containing both the policy selection experiences such as the learning parameters and observations among the learning agents, multi -agent reinforcement learning (RL) has to address adversarial communications that send fake learning messages to learning agents with the goal of decreasing the RL rewards or even failing the learning tasks. In this paper, we propose a multi-agent RL (MARL) communication framework for wireless networks against adversarial communications, in which each learning agent chooses the cooperative agents to share learning messages based on the agent reputation that indicates the probability to send fake learning messages. By comparing with the learning history, each learning agent authenticates the received learning messages before integrating them in the RL task state formulation and the learning parameter update for robust task learning. The learning factor that increases with the correlation between the local and the shared observation is calculated to update the task Q-values and neural network weights based on the shared learning parameters. As a case study, the multi-agent deep Q-network within our proposed MARL communication framework is implemented in the UAV swarm video transmission system and the performance gain over the benchmark is provided in the simulation results based on 5-UAV swarm against an attacker that sends fake observations and neural network weights. Zefang Lv, Liang Xiao 0003, Helin Yang, Xiangyang Ji |
GLOBECOM | 3 |
| 2023 | Efficient Communications for Multi-Agent Reinforcement Learning in Wireless NetworksabstractMulti-agent reinforcement learning (RL) utilizes the observations and learning experiences shared among the agents to accelerate learning speed under partial observations and the resulting learning efficiency depends on the cooperative agent selection and the RL task state formulation. In this paper, we propose an efficient communication scheme for multi-agent RL that enables each learning agent to optimize the cooperative agent selection and the task state formulation to improve the learning performance and the quality of service for RL-based applications in wireless networks. Based on the local observation, the radio channel states, the similarity of RL task with neighboring agents and previous communication cost, this scheme formulates a communication state, which is input to a neural network to estimate the communication policy distribution. The RL task state of the learning agent, which consists of the local observation such as channel states and previous task performance, as well as the correlation between the shared and the local observation extracted based on the attention mechanism, is formulated to enhance the agent receptive field. In addition, the shared learning information is also exploited to update the local learning parameters such as the task Q-values and neural network weights and further improve the RL task policy exploration. As a case study, the proposed communication scheme is implemented in the multi-agent deep Q-network based anti-jamming unmanned aerial vehicle swarm communications and the performance gain over the benchmark is verified via simulation results. Zefang Lv, Yousong Du, Liang Xiao 0003, Shuai Han 0002, Xiangyang Ji |
GLOBECOM | 4 |
| 2023 | Reinforcement Learning Based Energy-Efficient Routing with Latency Constraints for FANETsabstractReinforcement learning (RL) enables flying ad-hoc networks (FANETs) to choose the next hop unmanned aerial vehicles (UAV s) with shorter routing path, but may raise the retransmission rate and fails to guarantee the quality of service (QoS) under the high mobility and fast fading channels. In this paper, we propose an RL based routing scheme that optimizes both the routing and the power allocation to protect the latency QoS and save routing energy consumption of the FANET. Based on the routing history, the channel conditions, the battery level and the shared knowledge from the neighbors, this scheme formulates the routing policy distribution with safe exploration to select the stable path and thus reduce the retransmission rate. Specifically, the risk value with respect to end-to-end latency constraint is designed to evaluate the routing policy and reduce the exploration probability of the high-latency routing. Based on the distributed value function approach, the learning parameter such as the state value functions shared among neighbors is exploited to accelerate the routing process and enhance the routing stability under the dynamic network topology. Simulation results verify the routing performance gain of our proposed scheme over the benchmark. Xuchen Qi, Jieling Li, Zefang Lv, Liang Xiao 0003 |
GLOBECOM | 4 |
| 2023 | Energy and Latency-Aware Resource Management for UAV-Assisted Mobile Edge Computing Against JammingabstractUnmanned aerial vehicles (UAVs) have been increasingly employed as aerial servers in mobile edge computing (MEC) systems, providing essential computing, communication, and storage services for edge users. This UAV-assisted MEC paradigm shows great promise in enhancing both computing and communication performances. However, the presence of malicious jammers poses significant challenges to the system's reliability and efficiency. In this study, we explore the resource management problem in a multi-UAV-assisted MEC scenario under the influence of multiple malicious jammers. To mitigate the impact of jamming attacks, we propose a resource management approach with the primary objective of minimizing system energy consumption and latency while adhering to UAV energy constraints. Due to the dynamic and time-varying nature of the communication environment, we present a deep reinforcement learning (DRL)-based algorithm that dynamically adjusts the CPU frequency and communication bandwidth of the UAV to optimize the system performance even under jamming attacks. Through simulations, we demonstrate the effectiveness of the proposed algorithm in significantly reducing the overall system latency (both computational and communication latency) as well as minimizing energy consumption. Ziling Shao 0001, Helin Yang, Liang Xiao 0003, Wei Su 0002, Zehui Xiong |
GLOBECOM | 3 |
| 2023 | Joint Trajectory Optimization and Power Control for Cognitive UAV-Assisted Secure CommunicationsabstractCognitive unmanned aerial vehicle (UAV) communication systems combine benefits of both the cognitive radio and UAV, which improves the spectral efficiency and communication coverage area. However, high-quality air-ground channel links maybe more vulnerable to potential eavesdropping or jamming attacks. Thus, this paper proposes a secure transmission approach assisted by deploying a cooperative UAV to transmit artificial noise to jam an active eavesdropper, in order to maximize the system secrecy rate under the quality of service (QoS) requirement of a primary device. Specifically, we jointly optimize the flight trajectory and transmission power of the cooperative jammer to maximize the system's secrecy rate under strict constraints. To achieve this, we convert the non-convex problem into an approximately convex problem using the block coordinate descent algorithm and successive convex approximation method. Simulation results show that compared to existing algorithms, the proposed algorithm in this study can significantly improve the system's secrecy rate. Helin Yang, Liang Xiao 0003, Huabing Lu, Zehui Xiong |
GLOBECOM | 3 |
| 2023 | Reliable Communications for Hypersonic Vehicles: A Reinforcement Learning ApproachabstractThe ultra-high speed (e.g., typically moving with 10–20 Mach) of the hypersonic vehicle (HSV) causes a plasma sheath, which severely degrades the communication performance and results in communication blackouts. In this paper, we propose a deep reinforcement learning (RL)-based HSV reliable communications scheme against jamming, which enables the HSV to select the carrier frequency and transmit power according to the signal quality, the estimated voltage standing wave ratio, flight altitude, flight speed, and angle of attack. Specifically, we design a deep two-level hierarchical structure to compress the high-dimensional state and action space, with the added advantage of leveraging transfer learning to reduce initial exploration and expedite the optimization process. To optimize the learning speed, the dueling architecture is implemented in the deep network to measure the state value and the advantage function of the policies. In contrast to the benchmark, the simulation results indicate that the proposed scheme yields a significant reduction in both bit error rate and transmit power. Jingchen Xu, Zhiping Lin 0002, Yousong Du, Helin Yang, Liang Xiao 0003 |
GLOBECOM | 6 |
| 2023 | Reinforcement Learning Based UAV Swarm Communications Against JammingabstractReinforcement learning based unmanned aerial vehicle (UAV) swarm communications have to address the challenges raised by the large-scale dynamic network and strong jamming and interference. In this paper, we propose a multiagent reinforcement learning based UAV swarm anti-jamming communication scheme to optimize the UAV relay selection and power allocation based on the network topology, channel states, previous performance and the network states shared by neighboring UAVs. This scheme formulates the policy distribution to improve the policy space exploration and designs a soft learning mechanism to guide the policy update and stabilize the learning process. According to transfer learning, the shared swarm experiences are exploited to accelerate the initial policy learning. We investigate the computational complexity of the proposed scheme and derive the performance bound regarding the message bit error rate, the swarm energy consumption and the utility. Simulation results show that the proposed scheme improves the swarm communication performance and saves energy consumption compared with the benchmark scheme. Zefang Lv, Guohang Niu, Liang Xiao 0003, Chengwen Xing, Wenyuan Xu 0001 |
ICC | 3 |
| 2023 | Reinforcement Learning Based Friendly Jamming for Digital Twins Against Active EavesdroppingabstractDigital twin systems (DTs) are susceptible to active eavesdroppers engaging in wiretapping and jamming activities, aimed at increasing the physical layer's transmit power to steal additional virtual information. In this paper, we propose a deep reinforcement learning-based friendly jamming method for intratwin communications in DTs that enable the friendly jammer to optimize jamming frequency, power and the jamming duration against active eavesdropping. A safe and hierarchical architecture is designed that utilizes information such as the channel state of the device-server and the hostile jamming strength or wiretap channel of the active eavesdropper to improve anti-eavesdropping performance and secrecy rate. We apply the proposed friendly jamming method using universal software radio peripherals and assess its performance through experimentation. The experimental results illustrate that the proposed strategies significantly enhance the DTs secrecy rate in cross-layer transmission, and reduce the eavesdropping data rate and the physical layer energy consumption compared to existing friendly jamming methods. Kunze Li, Yuxiao Ren, Zhiping Lin 0002, Liang Xiao 0003 |
MSN | 4 |
| 2023 | Redistillation of Radio Frequency Knowledge for RFF Imbalanced Sample RecognitionabstractRadio Frequency Fingerprint (RFF) technology is an effective means to defend against cheating and counterfeiting attacks in wireless communication. However, to move from a theoretical algorithm to a practical application, the challenges of imbalanced data samples and environmental noise must be addressed for Radio Frequency (RF) identification technology. Although noise reduction can restore the signal to some extent, the recognition performance of RFF technology is affected when the dataset is imbalanced. While many RF identification algorithms focus on identification performance under a low Signal-to-Noise Ratio (SNR), performance degradation caused by data imbalance is a pressing problem that requires attention. Directly applying re-sampling algorithms in imbalanced dataset processing can lead to data overlap and neural network over-fitting. To address these issues, this paper proposes a “Redistillation of Radio Frequency Knowledge” (RRFK) algorithm combined with knowledge distillation (KD). The experimental results show that the proposed algorithm can achieve good recognition performance in both stepped and long-tail imbalanced data sets. Caidan Zhao, Liang Xiao 0003 |
SMC | 3 |
| 2023 | Random Railings Enhancement For RFF Imbalanced Data AugmentationabstractRadio Frequency Fingerprint (RFF) technology is an effective means to defend against cheating and counterfeit attacks in wireless communication. A deep learning-based RFF recognition algorithm can achieve well recognition performance, but it needs many balanced samples to train the model. However, the problem of sample imbalance is widespread in RFF identification tasks, and the number of signal samples of illegal devices is minimal. Neural network models usually can't learn these minority representations well, which seriously affects the performance of RFF recognition. Many advanced algorithms proposed to alleviate the problem of data imbalance don't perform well in the task of RFF recognition because they ignore the characteristics of RFF signals. Therefore, an algorithm based on Random Railings Enhancement (RRE) is proposed in this paper, which fills the data set with random masks according to the signal values front and rear. RRE protects the original signal's information, effectively expands the rare dataset, and has the effect of data enhancement. The experimental results show that the RRE can improve the performance of Radio Frequency (RF) identification technology tasks in the case of imbalanced data sets. Caidan Zhao, Liang Xiao 0003, Xiangyu Huang |
WCNC | 3 |
| 2023 | Real-Time DDoS Defense in 5G-Enabled IoT: A Multidomain Collaboration PerspectiveabstractWhile 5G networks have accelerated the development of the Internet of Things (IoT), they have also introduced a large number of vulnerable IoT devices into the network, which would lead to severe Distributed Denial-of-Service (DDoS) attacks. The newly emerging DDoS attack methods generally have a shorter duration, which imposes higher requirements for the response time of DDoS mitigation technologies. Existing DDoS defense methods cannot achieve real-time detection due to the difficulty of reducing the delay of feature extraction and large-scale data processing. In this article, we focus on the timeliness of DDoS detection and mitigation. We hope that deploying effective defense countermeasures at the source side will block the majority of DDoS attack traffic in real time before it enters the data network (DN). To this end, we propose a real-time DDoS defense framework based on multidomain collaboration that combines multisource information to detect attack sessions with high accuracy in 5G networks. To operate the framework at line rate, we propose an optimal packet sampling strategy based on the accurate session size estimation, which can greatly reduce the detection overhead while ensuring good accuracy. In a typical scenario with an attack session size larger than 10, this method can achieve a 99% detection rate while reducing the packet inspection rate (PIR) to less than 37%. Xu Chen 0004, Yunfei Chen 0001, Wei Feng 0001, Liang Xiao 0003, Xiangling Li, Jie Zhang 0003, Ning Ge 0001 |
IEEE Internet Things J. | 4 |
| 2023 | Reinforcement Learning Based Energy-Efficient Collaborative Inference for Mobile Edge ComputingabstractCollaborative inference in mobile edge computing (MEC) enables mobile devices to offload the computation tasks for the computation-intensive perception services, and the inference policy determines the inference latency and energy consumption. The optimal inference policy depends on the inference performance model of deep learning, the data generation model and the network model that are rarely known by mobile devices in time. In this paper, we propose a multi-agent reinforcement learning (RL) based energy-efficient MEC collaborative inference scheme, which enables each mobile device to choose both the partition point of deep learning and the collaborative edge of each mobile device based on the image quantity, the channel conditions and the previous inference performance. A learning experience exchange mechanism exploits the Q-values of the neighboring mobile devices to accelerate the inference policy optimization with less energy consumption. We also provide a deep multi-agent RL based inference scheme to accelerate learning for large-scale MEC networks, in which an actor network yields the collaborative inference policy probability distribution and a critic network guides the weight update of the actor network to enhance sample efficiency. We provide the inference performance bound and analyze the computational complexity. Both simulation and experimental results show that our proposed schemes reduce the inference latency and save the MEC energy consumption. Yilin Xiao 0001, Liang Xiao 0003, Kunpeng Wan, Helin Yang, Yi Zhang 0035, Yi Wu 0010, Yanyong Zhang |
IEEE Trans. Commun. | 2 |
| 2023 | An Advanced Integrated Visible Light Communication and Localization SystemabstractVisible light communication (VLC) is an emerging wireless technology to support high transmission rate for indoor devices by using existing lighting infrastructure, and VLC-based indoor localization is capable of providing high-accuracy localization. However, current VLC-based localization systems suffer from several key challenges such as sensitivity to random tilting of the receiver, which limits its full potential in real-world applications. In this paper, we design an integrated visible light communication and localization (VLCL) system to simultaneously support accurate real-time localization and communication services for indoor devices. To achieve this, an advanced differential phase difference of arrival (A-DPDOA) localization design is developed to simplify hardware and improve tracking robustness. In addition, a joint adaptive modulation, subcarrier and power allocation scheme is also proposed, which aims to improve the communication data rate and localization accuracy. Extensive experiments are performed to demonstrate that the proposed integrated VLCL system achieves higher localization accuracy and transmission data rate, compared to existing systems and schemes. Experiments also illustrate that the localization algorithm is more robust against the random tilting of the receiver under device movement in two-dimensional and three-dimensional scenarios. Helin Yang, Sheng Zhang 0023, Arokiaswami Alphones, Chen Chen 0037, Kwok-Yan Lam, Zehui Xiong, Liang Xiao 0003, Yi Zhang 0035 |
IEEE Trans. Commun. | 7 |
| 2023 | Broadcast Secrecy Rate Maximization in UAV-Empowered IRS Backscatter CommunicationsabstractThe backscatter communications (BackCom) and physical layer security are respected to realize extremely low-power secure communications in the imminent sixth generation (6G). In a BackCom system, the backscatter device without radio frequency components sends messages to users by reflecting the external signals. However, the double-fading effect limits BackCom’s performance and the commonly used broadcast mode is vulnerable to eavesdropping. Two promising technologies, intelligent reflecting surface (IRS) and unmanned aerial vehicle (UAV), show excellent potential in handling these problems. In this paper, we propose a UAV-empowered IRS-BackCom network, where an IRS acts as the backscatter device and uses the received signals from a UAV for BackCom. We aim to guarantee secure transmission and maximize the broadcast secrecy rate by jointly optimizing the UAV’s beamformer and trajectory and the IRS’s reflection coefficient. To tackle the non-convex problem, we leverage the block coordinate descent method to decompose it into three subproblems. Specifically, the UAV’s beamformer and trajectory and the IRS’s reflection matrix are optimized alternatively. Further, we adopt reinforcement learning to facilitate the intractable UAV’s trajectory optimization. Simulation results verify the feasibility and effectiveness of the proposed system model and the optimization scheme. Shuai Han 0002, Liang Xiao 0003, Cheng Li 0005 |
IEEE Trans. Wirel. Commun. | 3 |
| 2023 | Multi-Agent Reinforcement Learning Based UAV Swarm Communications Against JammingabstractThe swarm relay and power allocation policy determines the bit error rate and the energy consumption of unmanned aerial vehicles (UAVs) and can be optimized based on the network and jamming model, which is rarely known by UAVs. In this paper, we propose a multi-agent reinforcement learning (RL)-based UAV swarm communication scheme to optimize the relay selection and power allocation against jamming. Based on the network topology, channel states, previous performance and observations shared by the neighboring UAVs, this scheme formulates the policy distribution to improve the policy exploration and applies a policy learning mechanism to stabilize the learning process. Based on transfer learning, the shared swarm experiences are exploited to accelerate the initial learning and improve policy optimization. A deep RL-based scheme is proposed to mitigate the state quantization error for the rapidly changing channel states under high swarm moving speed and thus further improve the anti-jamming performance. This scheme designs a policy network with four fully connected layers to approximate the policy distribution and uses another two neural networks to estimate the average policy distribution and the expected long-term utility, respectively, to update the policy network for stabilized deep learning. We investigate the computational complexity and derive the performance bound regarding the bit error rate, the energy consumption and the utility. Simulation and experimental results verify the performance gain of our proposed schemes over related works. Zefang Lv, Liang Xiao 0003, Yousong Du, Guohang Niu, Chengwen Xing, Wenyuan Xu 0001 |
IEEE Trans. Wirel. Commun. | 2 |
| 2022 | Environment-Aware Reinforcement Learning Based VANET Communications Against Jamming and InterferenceabstractThe jamming and interference in vehicular ad hoc networks (VANETs) depend on the channel states of vehicles from the ambient radio transmitter, which in turn result from the topologies and radio features. In this paper, we propose an environment-aware reinforcement learning (RL)-based VANET communication scheme against jamming and interference that applies the post decision state algorithm to optimize the power allocation and channel selection without relying on the jamming attack model. This scheme exploits the environment information in the state formulation due to the traffic density and their locations reflect the interference level, as well as the location of transmission vehicle combined with building structure and heights indicate the channel gain and shadowing. The proposed post decision state-based RL method employs the estimated future communication distances of the moving vehicles to accelerate the learning process. We provide the performance bounds of the energy consumption, bit error rate (BER), and utility based on a Nash equilibrium. Simulation results show that the proposed scheme significantly reduces the BER with less energy consumption compared with the benchmark. Zhiping Lin 0002, Xiaohao Yan, Liang Xiao 0003, Yan Shi 0002, Yuliang Tang, Jun Liu 0006 |
GLOBECOM | 3 |
| 2022 | Reinforcement Learning Based Vulnerability Analysis for Smart Grids Against False Data Injection Attacks
Liang Xiao 0003, Zefang Lv |
WASA (1) | 3 |
| 2022 | Reinforcement learning based energy efficient robot relay for unmanned aerial vehicles against smart jamming
Xiaozhen Lu, Jingfang Jie, Liang Xiao 0003, Jin Li 0002, Yanyong Zhang |
Sci. China Inf. Sci. | 4 |
| 2022 | DDoS Defense for IoT: A Stackelberg Game Model-Enabled Collaborative FrameworkabstractThe proliferation of Distributed Denial of Service (DDoS) attacks in Internet of Things (IoT) not only threatens the security of digital devices and infrastructure but also severely degrades IoT system performance due to the overly consumed network resources. With the knowledge of identity information of devices and signaling data, Internet service providers (ISPs) can detect and block DDoS traffic by monitoring the upstream IoT packets, and thereby, improve network efficiency. However, inspecting all data packets online for DDoS detection will significantly increase both the network delay and the computational overhead. Therefore, the packet sampling strategy is crucial for the defenders to detect DDoS attacks. To this end, this article formulates a Stackelberg game model to analyze the collaborative IoT packet sampling against DDoS attacks. Through the equilibrium analysis of the DDoS game, we derive the lower bound of packet sampling rate (PSR) that can effectively deter potential attackers. Unlike traditional offline detection, our proposed packet sampling strategy can support both the online detection and proactive prevention of DDoS traffic. As a use case, a multipoint DDoS defense framework is developed to address the IP spoofing in 5G networks based on the proposed packet sampling strategy, which deters DDoS attacks and reduces the packet sampling cost, and thereby, maximizes the IoT utility, compared with existing methods. In typical reflection attacks (in which no more than five packets of response are triggered by a request packet), our proposed scheme not only reduces more than 70% of the sampling rate but also demonstrates superior robustness against boundary condition variation. Xu Chen 0004, Liang Xiao 0003, Wei Feng 0001, Ning Ge 0001, Xianbin Wang 0001 |
IEEE Internet Things J. | 2 |
| 2022 | IRS-Aided Energy-Efficient Secure WBAN Transmission Based on Deep Reinforcement LearningabstractWireless body area networks (WBANs) are vulnerable to active eavesdropping that simultaneously perform sniffing and jamming to raise the sensor transmit power, and thus steal more healthcare data. In this paper, we propose an intelligent reflecting surface (IRS)-aided reinforcement learning (RL) based secure WBAN transmission scheme that enables the coordinator to jointly optimize the sensor encryption key and transmit power, as well as the IRS phase shifts against active eavesdropping. A Dyna architecture is designed to improve the learning efficiency with the simulated transmission experiences and safe exploration is applied to avoid the risky policies that result in severe data leakage. A deep RL based WBAN transmission scheme is proposed to further improve the secure transmission with lower eavesdropping rate, intercept probability, sensor energy consumption and transmission latency for the coordinators that support deep learning. We analyze the computational complexity and investigate the equilibrium of the secure transmission game between the coordinator and the eavesdropper to provide the performance bounds, which is verified via the simulation results, showing the efficacy of our proposed schemes. Liang Xiao 0003, Siyuan Hong, Helin Yang, Xiangyang Ji |
IEEE Trans. Commun. | 1 |
| 2022 | Reinforcement Learning Based Network Coding for Drone-Aided Secure Wireless CommunicationsabstractActive eavesdropper sends jamming signals to raise the transmit power of base stations and steal more information from cellular systems. Network coding resists the active eavesdroppers that cannot obtain all the data flows, but highly relies on the wiretap channel states that are rarely known in wireless networks. In this paper, we present a reinforcement learning (RL) based random linear network coding scheme for drone-aided cellular systems to address eavesdropping. In this scheme, the network coding policy, including the encoded packet number, the packet and power allocation, is chosen based on the measured jamming power, previous transmission performance and BS channel states. A virtual model generates simulated experiences to update Q-values besides real experiences for faster policy optimization. We also propose a deep RL version and design a hierarchical architecture to further accelerate the policy exploration and improve the anti-eavesdropping performance, in terms of the intercept probability, the latency, the outage probability and the energy consumption. We analyze the computational complexity, drone deployment, secure coverage area and the performance bound of the proposed schemes, which are verified via simulation results. Liang Xiao 0003, Yi Zhang 0035, Li-Chun Wang 0001, Shaodan Ma |
IEEE Trans. Commun. | 1 |
| 2022 | Safe Exploration in Wireless Security: A Safe Reinforcement Learning Algorithm With Hierarchical StructureabstractMost safe reinforcement learning (RL) algorithms depend on the accurate reward that is rarely available in wireless security applications and suffer from severe performance degradation for the learning agents that have to choose the policy from a large action set. In this paper, we propose a safe RL algorithm, which uses a policy priority-based hierarchical structure to divide each policy into sub-policies with different selection priorities and thus compresses the action set. By applying inter-agent transfer learning to initialize the learning parameters, this algorithm accelerates the initial exploration of the optimal policy. Based on a security criterion that evaluates the risk value, the sub-policy distribution formulation avoids the dangerous sub-policies that cause learning failure such as severe network security problems in wireless security applications, e.g., Internet services interruption. We also propose a deep safe RL and design four deep neural networks in each sub-policy selection to further improve the learning efficiency for the learning agents that support four convolutional neural networks (CNNs): The Q-network evaluates the long-term expected reward of each sub-policy under the current state, and the E-network evaluates the long-term risk value. The target Q and E-networks update the learning parameters of the corresponding CNN to improve the policy exploration stability. As a case study, our proposed safe RL algorithms are implemented in the anti-jamming communication of unmanned aerial vehicles (UAVs) to select the frequency channel and transmit power to the ground node. Experimental results show that our proposed schemes significantly improve the UAV communication performance, save the UAV energy and increase the reward compared with the benchmark against jamming. Xiaozhen Lu, Liang Xiao 0003, Guohang Niu, Xiangyang Ji, Qian Wang 0002 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | 3D Geo-Indistinguishability for Indoor Location-Based ServicesabstractIndoor location-based services (LBS) are widely used in large-scale indoor buildings, such as high-rise hospitals and multi-story shopping malls. At the same time, location privacy protection in such three-dimensional (3D) space has recently attracted considerable attention. Currently, most existing location privacy protection schemes focus on two-dimensional (2D) location protection and fail to prevent location inference attacks when the user’s location data include height dimension, i.e., 3D geolocation. Enlightened by the concept of differential privacy, in this paper we first study the impact factors of the degree of indistinguishability of 3D geolocations. Then, we quantify location privacy for LBS applications in the 3D space with geo-indistinguishability (3D-GI) rigorously and provably. We develop a mechanism of three-variates Laplacian to generate perturbed locations considering the locations’ X, Y, and Z-coordinates simultaneously, guaranteeing geo-indistinguishability. Furthermore, the discretization noise-adding mechanism is studied to satisfy geo-indistinguishability in the 3D space under the finite precision of hardware/devices. Considering the discretized mechanism can only satisfy geo-indistinguishability in finite 3D space and users visit the limited regions, we further study the truncation of the Laplacian mechanism to limit the generated perturbed locations within a specific region. Simulation results demonstrate that the proposed 3D-GI outperforms the benchmarks while guaranteeing privacy regardless of the adversary’s prior knowledge. Minghui Min, Liang Xiao 0003, Jiahao Ding, Hongliang Zhang 0001, Shiyin Li, Miao Pan, Zhu Han 0001 |
IEEE Trans. Wirel. Commun. | 2 |
| 2022 | Distributed Deep Reinforcement Learning-Based Spectrum and Power Allocation for Heterogeneous NetworksabstractThis paper investigates the problem of distributed resource management in two-tier heterogeneous networks, where each cell selects its joint device association, spectrum allocation, and power allocation strategy based only on locally-observed information without any central controller. As the optimization problem with devices’ quality-of-service (QoS) constraints is non-convex and NP-hard, we model it as a Markov decision process (MDP). Considering the fact that the network is highly complex with large state and action spaces, a multi-agent dueling deep-Q network-based algorithm combined with distributed coordinated learning is proposed to effectively learn the optimized intelligent resource management policy, where the algorithm adopts dueling deep network to learn the action-value distribution by estimating both the state-value and action advantage functions. Under the distributed coordinated learning manner and dueling architecture, the learning algorithm can rapidly converge to the optimized policy. Simulation results demonstrate that the proposed distributed coordinated learning algorithm outperforms other existing learning algorithms in terms of learning efficiency, network data rate, and QoS satisfaction probability. Helin Yang, Jun Zhao 0007, Kwok-Yan Lam, Zehui Xiong, Qingqing Wu 0001, Liang Xiao 0003 |
IEEE Trans. Wirel. Commun. | 6 |
| 2021 | Drone-Aided Network Coding for Secure Wireless Communications: A Reinforcement Learning ApproachabstractThis study investigates how base stations (BSs) apply network coding to protect the downlink data and how drones relay the coded packets to resist active eavesdropping that performs jamming to induce the BS to raise the transmit power and thus steal more data. We present a drone-aided network coding framework for secure downlink transmission, which incorporates a random linear network coding algorithm to encode the BS messages against active eavesdropping. This framework designs a model-based reinforcement learning to choose the BS network coding and transmission policy based on the jamming power sent by the active eavesdropper, the previous transmission performance, and the BS channel states without the prior knowledge of the drone-eavesdropper channel states. The learning parameters such as the Q-values are updated by the real experiences in the downlink transmission process besides the simulated experiences that are generated from the virtual model in the designed Dyna architecture. Simulation results show that our proposed scheme outperforms the benchmarks in terms of the intercept probability, the transmission performance, and the BS energy consumption. Xiaozhen Lu, Liang Xiao 0003, Li-Chun Wang 0001 |
GLOBECOM | 4 |
| 2021 | Reinforcement Learning Based Sensor Encryption and Power Control for Low-Latency WBANs
Siyuan Hong, Xiaozhen Lu, Liang Xiao 0003, Guohang Niu, Helin Yang |
WASA (2) | 3 |
| 2021 | Deep-Reinforcement-Learning-Based User Profile Perturbation for Privacy-Aware RecommendationabstractUser profile perturbation protects privacy in the release of user profiles to receive recommendation services, in which the privacy budget as a privacy parameter can be controlled to effect a tradeoff between the recommendation quality and privacy protection against inference attacks. In this article, we propose a deep reinforcement learning (RL)-based user profile perturbation scheme for recommendation systems. This scheme applies differential privacy to protect user privacy and uses deep RL to choose the privacy budget against inference attackers. Based on an evaluated neural network (NN) and a target NN, this scheme enables a user device to optimize the privacy budget over time based on the sensitivity level of the clicked item, the similarities among the recommended items, and the estimated privacy loss. We provide an upper bound on the privacy protection performance of this scheme in the recommendation game and evaluate its computational complexity. Simulation results for a movie recommendation system show that this scheme increases the user privacy protection level for a given recommendation quality compared with benchmark schemes. Yilin Xiao 0001, Liang Xiao 0003, Xiaozhen Lu, Hailu Zhang, Shui Yu 0001, H. Vincent Poor |
IEEE Internet Things J. | 2 |
| 2021 | Privacy-Preserving Federated Learning for UAV-Enabled Networks: Learning-Based Joint Scheduling and Resource ManagementabstractUnmanned aerial vehicles (UAVs) are capable of serving as flying base stations (BSs) for supporting data collection, machine learning (ML) model training, and wireless communications. However, due to the privacy concerns of devices and limited computation or communication resource of UAVs, it is impractical to send raw data of devices to UAV servers for model training. Moreover, due to the dynamic channel condition and heterogeneous computing capacity of devices in UAV-enabled networks, the reliability and efficiency of data sharing require to be further improved. In this paper, we develop an asynchronous federated learning (AFL) framework for multi-UAV-enabled networks, which can provide asynchronous distributed computing by enabling model training locally without transmitting raw sensitive data to UAV servers. The device selection strategy is also introduced into the AFL framework to keep the low-quality devices from affecting the learning efficiency and accuracy. Moreover, we propose an asynchronous advantage actor-critic (A3C) based joint device selection, UAVs placement, and resource management algorithm to enhance the federated convergence speed and accuracy. Simulation results demonstrate that our proposed framework and algorithm achieve higher learning accuracy and faster federated execution time compared to other existing solutions. Helin Yang, Jun Zhao 0007, Zehui Xiong, Kwok-Yan Lam, Sumei Sun, Liang Xiao 0003 |
IEEE J. Sel. Areas Commun. | 6 |
| 2021 | UAV Anti-Jamming Video Transmissions With QoE Guarantee: A Reinforcement Learning-Based ApproachabstractUnmanned aerial vehicles (UAVs) that are widely utilized for video capturing, processing and transmission have to address jamming attacks with dynamic topology and limited energy. In this paper, we propose a reinforcement learning (RL)-based UAV anti-jamming video transmission scheme to choose the video compression quantization parameter, the channel coding rate, the modulation and power control strategies against jamming attacks. More specifically, this scheme applies RL to choose the UAV video compression and transmission policy based on the observed video task priority, the UAV-controller channel state and the received jamming power. This scheme enables the UAV to guarantee the video quality-of-experience (QoE) and reduce the energy consumption without relying on the jamming model or the video service model. A safe RL-based approach is further proposed, which uses deep learning to accelerate the UAV learning process and reduce the video transmission outage probability. The computational complexity is provided and the optimal utility of the UAV is derived and verified via simulations. Simulation results show that the proposed schemes significantly improve the video quality and reduce the transmission latency and energy consumption of the UAV compared with existing schemes. Liang Xiao 0003, Yuzhen Ding, Jinhao Huang, Sicong Liu 0002, Yuliang Tang, Huaiyu Dai |
IEEE Trans. Commun. | 1 |
| 2021 | Reinforcement Learning-Based Physical-Layer Authentication for Controller Area NetworksabstractIn controller area networks (CANs), electronic control units (ECUs) such as telematics ECUs and on-board diagnostic ports must protect the message exchange from spoofing attacks. In this paper, we propose a CAN bus authentication framework that exploits physical layer features of the messages, including message arrival intervals and signal voltages, and applies reinforcement learning to choose the authentication mode and parameter. By applying the Dyna architecture and using a double estimator, this scheme improves the utility in terms of authentication accuracy without changing the CAN bus protocol or the ECU components and requiring knowledge of the spoofing model. We also propose a deep learning version to further improve the authentication efficiency for the CAN bus. The learning scheme applies a hierarchical structure to reduce the exploration time, and uses two deep neural networks to compress the high-dimensional state space and to fully exploit the physical authentication experiences. We provide the computational complexity and the performance analysis. Experimental results verify the theoretical analysis and show that our proposed schemes significantly improve the authentication accuracy as compared with benchmark schemes. Liang Xiao 0003, Xiaozhen Lu, Tangwei Xu, Weihua Zhuang, Huaiyu Dai |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2021 | Deep Reinforcement Learning-Based Intelligent Reflecting Surface for Secure Wireless CommunicationsabstractIn this paper, we study an intelligent reflecting surface (IRS)-aided wireless secure communication system, where an IRS is deployed to adjust its reflecting elements to secure the communication of multiple legitimate users in the presence of multiple eavesdroppers. Aiming to improve the system secrecy rate, a design problem for jointly optimizing the base station (BS)'s beamforming and the IRS's reflecting beamforming is formulated considering different quality of service (QoS) requirements and time-varying channel conditions. As the system is highly dynamic and complex, and it is challenging to address the non-convex optimization problem, a novel deep reinforcement learning (DRL)-based secure beamforming approach is firstly proposed to achieve the optimal beamforming policy against eavesdroppers in dynamic environments. Furthermore, post-decision state (PDS) and prioritized experience replay (PER) schemes are utilized to enhance the learning efficiency and secrecy performance. Specifically, a modified PDS scheme is presented to trace the channel dynamic and adjust the beamforming policy against channel uncertainty accordingly. Simulation results demonstrate that the proposed deep PDS-PER learning based secure beamforming approach can significantly improve the system secrecy rate and QoS satisfaction probability in IRS-aided secure communication systems. Helin Yang, Zehui Xiong, Jun Zhao 0007, Dusit Niyato, Liang Xiao 0003, Qingqing Wu 0001 |
IEEE Trans. Wirel. Commun. | 5 |
| 2020 | Learning Based Energy Efficient Radar Power Control Against Deceptive JammingabstractMultiple-input and multiple-output (MIMO) radars are vulnerable to deceptive jamming launched by false target generators that send jamming signals with the goal of pretending that the radar echo signals are reflected by faked targets. In this paper, we present a reinforcement learning based energy efficient power control scheme to detect deceptive jamming for frequency diverse array MIMO radars. This scheme enables a radar to choose the transmit power over the antennas without relying on the known deceptive jamming model. Instead, based on the emergency level, the battery level, the echo signal quality, the antenna phase differences, the received jamming power and the previous detection error rate, this scheme improves the detection accuracy and the energy efficiency, and uses a Dyna architecture to train the learning parameters with the simulated jamming detection experiences for faster optimization in the dynamic game against deceptive jamming. Simulation results show that this scheme effectively improves the deceptive jamming detection accuracy and saves the radar energy. Index Terms-MIMO, radar, reinforcement learning, deceptive jamming. Xiaozhen Lu, Liang Xiao 0003, Li Xu 0002 |
GLOBECOM | 3 |
| 2020 | (τ, ϵ)-Greedy Reinforcement Learning For Anti-Jamming Wireless CommunicationsabstractIn this article, we propose a(τ, ε)-greedy reinforcement learning algorithm for anti-jamming wireless communications, which chooses previous action with probability τ and applies ε-greedy with probability 1-τ. The key idea of our algorithm is that the more valuable the previous action is, the higher probability of directly performing it at the current time slot without learning. For this purpose, the average utility of several previous actions is first calculated as a threshold for the valuable action judgment. Then, probability τ is formulated as a Gaussian-like function with respect to the difference between the threshold and the utility of the previous action, which makes the wireless devices find the optimal action at a faster speed in the early stage, and eventually ensures the convergence. As a concrete example, the proposed algorithm is implemented in a wireless communication system against multiple jammers. Simulation results show that compared with ε-greedy, the (τ, ε)-greedy obtains faster convergence rate and slightly higher signalto-interference-plus-noise ratio when being applied to Qlearning, deep Q-networks (DQN), double DQN (DDQN), and prioritized experience reply based DDQN (PDDQN). The source code is available at https://github.com/GZHUDVL/tau-epsilon-greedy-RL. Yuan-Gen Wang, Jin Li 0002, Liang Xiao 0003, Guopu Zhu |
GLOBECOM | 4 |
| 2020 | Auto-Generating Neural Networks with Reinforcement Learning for Multi-Purpose Image ForensicsabstractDesigning a forensic convolutional neural network (CNN) is usually based on some ad-hoc intuition and domain knowledge. Many methods to automate neural network design have been proposed for computer vision tasks, but they may not be directly applied to image forensic problems, which tend to detect weak traces signals left by image operations rather than strong image content signals. In this paper, we propose an approach to learn an optimal forensic CNN structure with reinforcement learning for detecting multiple image tampering operations. A learning agent is introduced to select CNN layers sequentially in a limited state-action space using Q-learning with an $\epsilon$-greedy strategy and experience replay. The experiments demonstrate that the auto-generated network performs better than other classic image forensic methods and shows more robustness against JPEG compression. To our knowledge, this is the first attempt to design forensic deep neural networks automatically with reinforcement learning. Yujun Wei, Yifang Chen 0002, Xiangui Kang, Z. Jane Wang 0001, Liang Xiao 0003 |
ICME | 5 |
| 2020 | Electric Network Frequency Based Audio Forensics Using Convolutional Neural Networks
Maoyu Mao, Zhongcheng Xiao, Xiangui Kang, Liang Xiao 0003 |
IFIP Int. Conf. Digital Forensics | 5 |
| 2020 | Reinforcement Learning Aided Network Architecture Generation for JPEG Image SteganalysisabstractThe architectures of convolutional neural networks used in steganalysis have been designed heuristically. In this paper, an automatic Network Architecture Generation algorithm based on reinforcement learning for JPEG image Steganalysis (JS-NAG) has been proposed. Different from the automatic neural network generation methods in computer vision which are based on the strong content signals, steganalysis is based on the weak embedded signals, thus needs specific design. In the proposed method, the agent is trained to sequentially select some high-performing blocks using Q-learning to generate networks. An early stop strategy and a well-designed performance prediction function have been utilized to reduce the search time. To generate the optimal networks, hundreds of networks have been searched and trained on 3 GPUs for 15 days. To further improve the detection accuracy, we make an ensemble classifier out of the generated convolutional neural networks. The experimental results have shown that the proposed method significantly outperforms the current state-of-the-art CNN based methods. Beiling Lu, Liang Xiao 0003, Xiangui Kang, Yun Q. Shi 0001 |
IH&MMSec | 3 |
| 2020 | Energy Efficient Relay in UAV Networks Against Jamming: A Reinforcement Learning Based ApproachabstractUnmanned aerial vehicle (UAV) networks are vulnerable to jamming attacks because of the high mobility, limited battery and scarce spectrum resources of UAVs. In this paper, we propose a reinforcement learning based UAV relay scheme to improve the anti-jamming capability and save energy consumption of the UAV network. Based on the real-time channel conditions and the historical relay experiences, the proposed scheme enables UAVs to improve the policy of relay power and strategies without knowing the UAV network and channel model. Simulation results show that the proposed UAV relay scheme reduces the bit error rate of the messages and reduces the energy consumption of the UAV network compared with the state-of-the-art benchmark. Weihang Wang 0004, Xiaozhen Lu, Sicong Liu 0002, Liang Xiao 0003 |
VTC Spring | 4 |
| 2020 | Mobile Edge Computing Against Smart Attacks with Deep Reinforcement Learning in Cognitive MIMO IoT Systems
Songyang Ge, Beiling Lu, Liang Xiao 0003, Jie Gong 0003, Xiang Chen 0007, Yun Liu 0016 |
Mob. Networks Appl. | 3 |
| 2020 | DeepFusion: predicting movie popularity via cross-platform feature fusion
Wen Bai, Yipeng Zhou, Di Wu 0001, Gang Liu 0028, Liang Xiao 0003 |
Multim. Tools Appl. | 7 |
| 2020 | A Reinforcement Learning and Blockchain-Based Trust Mechanism for Edge NetworksabstractMobile edge computing (MEC) raises the issue of resisting selfish edge attackers that use less computation resources than promised to process offloading tasks or provide faked computation results. In this paper, we present a blockchain based trust mechanism to help MEC address selfish edge attacks and faked service record attacks. This mechanism evaluates the computational performance of the edge devices and broadcasts such information to the neighboring edge devices and mobile devices. By building a reputation assignment method for the edge devices, the edge reputation system chooses the miner of the blockchain, which applies the joint Proof-of-Work and Proof-of-Stake consensus protocol to append a block recording the new service reputations onto the MEC blockchain. We propose a reinforcement learning (RL) based edge central processing unit (CPU) allocation algorithm without knowing the mobile service generation model and the network model in the dynamic edge computing process and a deep RL version to further improve the computational performance. The security performance is analyzed and the performance bound of the edge utility is provided. Experimental results show that this framework suppresses the selfish edge attacks, decreases the response latency and saves the energy compared with a benchmark MEC scheme. Liang Xiao 0003, Yuzhen Ding, Donghua Jiang 0002, Jinhao Huang, Dongming Wang 0002, Jie Li 0002, H. Vincent Poor |
IEEE Trans. Commun. | 1 |
| 2020 | Reinforcement Learning-Based Mobile Offloading for Edge Computing Against Jamming and InterferenceabstractMobile edge computing systems help improve the performance of computational-intensive applications on mobile devices and have to resist jamming attacks and heavy interference. In this paper, we present a reinforcement learning based mobile offloading scheme for edge computing against jamming attacks and interference, which uses safe reinforcement learning to avoid choosing the risky offloading policy that fails to meet the computational latency requirements of the tasks. This scheme enables the mobile device to choose the edge device, the transmit power and the offloading rate to improve its utility including the sharing gain, the computational latency, the energy consumption and the signal-to-interference-plus-noise ratio of the offloading signals without knowing the task generation model, the edge computing model, and the jamming/interference model. We also design a deep reinforcement learning based mobile offloading for edge computing that uses an actor network to choose the offloading policy and a critic network to update the actor network weights to improve the computational performance. We discuss the computational complexity and provide the performance bound that consists of the computational latency and the energy consumption based on the Nash equilibrium of the mobile offloading game. Simulation results show that this scheme can reduce the computational latency and save energy consumption. Liang Xiao 0003, Xiaozhen Lu, Tangwei Xu, Xiaoyue Wan, Wen Ji 0003, Yanyong Zhang |
IEEE Trans. Commun. | 1 |
| 2020 | Heterogeneous Edge Offloading With Incomplete Information: A Minority Game ApproachabstractTask offloading is one of key operations in edge computing, which is essential for reducing the latency of task processing and boosting the capacity of end devices. However, the heterogeneity among tasks generated by various users makes it challenging to design efficient task offloading algorithms. In addition, the assumption of complete information for offloading decision-making does not always hold in a distributed edge computing environment. In this article, we formulate the problem of heterogeneous task offloading in a distributed environment as a minority game (MG), in which each player must make decisions independently in each turn and the players who end up on the minority side win. The multi-player MG incentivizes players to cooperate with each other in the scenarios with incomplete information, where players don't have full information about other players (e.g., the number of tasks, the required resources). To address the challenges incurred by task heterogeneity and the divergence of naive MG approaches, we propose an MG based scheme, in which tasks are divided into subtasks and instructed to form into a set of groups as possible, and the left ones are scheduled to perform decision adjustment in a probabilistic manner. We prove that our proposed algorithm can converge to a near-optimal point, and also investigate its stability and price of anarchy in terms of task processing time. Finally, we conduct a series of simulations to evaluate the effectiveness of our proposed scheme and the results indicate that our scheme can achieve around 30% reduction of task processing time compared with other approaches. Moreover, our proposed scheme can converge to a near-optimal point, which cannot be guaranteed by naive MG approaches. Miao Hu 0001, Di Wu 0001, Yipeng Zhou, Xu Chen 0004, Liang Xiao 0003 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2020 | Reinforcement Learning-Based Downlink Interference Control for Ultra-Dense Small CellsabstractThe dense deployment of small cells in 5G cellular networks raises the issue of controlling downlink inter-cell interference under time-varying channel states. In this paper, we propose a reinforcement learning based power control scheme to suppress downlink inter-cell interference and save energy for ultra-dense small cells. This scheme enables base stations to schedule the downlink transmit power without knowing the interference distribution and the channel states of the neighboring small cells. A deep reinforcement learning based interference control algorithm is designed to further accelerate learning for ultra-dense small cells with a large number of active users. Analytical convergence performance bounds including throughput, energy consumption, inter-cell interference, and the utility of base stations are provided and the computational complexity of our proposed scheme is discussed. Simulation results show that this scheme optimizes the downlink interference control performance after sufficient power control instances and significantly increases the network throughput with less energy consumption compared with a benchmark scheme. Liang Xiao 0003, Hailu Zhang, Yilin Xiao 0001, Xiaoyue Wan, Sicong Liu 0002, Li-Chun Wang 0001, H. Vincent Poor |
IEEE Trans. Wirel. Commun. | 1 |
| 2019 | QoE-Aware Power Control for UAV-Aided Media Transmission with Reinforcement LearningabstractUnmanned aerial vehicles (UAVs) are widely utilized to capture and compress videos of the target area and then transmit the processed videos to the control station (CS) on the ground. The media transmissions in the UAV-aided network face many challenges due to the highly dynamic network topology and limited resources such as bandwidth and energy. This paper introduces a media transmission scheme in the UAV-aided network utilizing reinforcement learning algorithms to efficiently process and transmit the captured video, which is able to improve the quality-of-experience (QoE) and reduce the energy consumption. Exploiting the proposed reinforcement learning algorithm, the UAV dynamically selects the quantization parameter in the source coding process and determines the transmit power without knowing the video transmission model. Simulation results demonstrate that the proposed scheme is capable of achieving a higher video quality and utility with lower energy consumption compared with the state-of-the-art schemes. Yuzhen Ding, Donghua Jiang 0002, Jinhao Huang, Liang Xiao 0003, Sicong Liu 0002, Yuliang Tang, Huaiyu Dai |
GLOBECOM | 4 |
| 2019 | Privacy Aware Recommendation: Reinforcement Learning Based User Profile PerturbationabstractUser profile release in recommendation systems can apply the user profile perturbation technique to protect user privacy, in which each user sends a perturbed user profile such as the a list of clicked items to receive a recommendation service from a server. The perturbation policy such as the privacy budget determines the recommendation quality and the privacy level, while its optimization usually depends on the known attack model, which is rarely known by the users. In this paper, we propose a reinforcement learning based user profile perturbation scheme that applies differential privacy to protect user privacy for recommendation systems. According to reinforcement learning, the privacy budget to perturb the released user profile depends on the features of the actual user profiles and the released user profiles, and the estimated user privacy level. This scheme enables a user to optimize his or her perturbation policy in terms of both the user privacy level and the received recommendation quality without being aware of the attack model. We evaluate the computational complexity of this scheme and analyze a case study, a privacy aware movie recommendation system. Simulation results show that this scheme improves user privacy protection for a given level of recommendation quality compared with a benchmark profile perturbation scheme. Yilin Xiao 0001, Liang Xiao 0003, Hailu Zhang, Shui Yu 0001, H. Vincent Poor |
GLOBECOM | 2 |
| 2019 | Reinforcement Learning with Safe Exploration for Network SecurityabstractSafe reinforcement learning is important for the safety critical applications especially network security, as the exploration of some dangerous actions can result in huge short-term losses such as network failure or large scale privacy leakage. In this paper, we propose a reinforcement learning algorithm with safe exploration and uses transfer learning to reduce the initial random exploration. A blacklist is maintained to record the most dangerous state-action pairs as a safety constraint. A safe deep reinforcement learning version uses a convolutional neural network to estimate the risk levels and thus further improves the safety of the exploration and accelerates the learning speed for the learning agent. As a case study, the proposed reinforcement learning with safe exploration is applied in the anti-jamming robot communications. Experimental results show that the proposed algorithms can improve the jamming resistance of the robot and reduce the outage rate to enter the most dangerous states compared with the benchmark algorithms. Canhuang Dai, Liang Xiao 0003, Xiaoyue Wan, Ye Chen 0011 |
ICASSP | 2 |
| 2019 | Protecting Semantic Trajectory Privacy for VANET with Reinforcement LearningabstractLocation-based services in vehicular ad hoc networks (VANETs) have to protect user privacy and address the challenge due to the disclosure of the vehicle movement trajectory. In this paper, we propose an reinforcement learning (RL) based differential privacy mechanism that randomizes the released vehicle locations to protect the semantic trajectory of the vehicle and uses RL to select the obfuscation policy. Based on the semantic location of the vehicle and the attack history, this scheme enables a vehicle to optimize the obfuscation policy in terms of the privacy gain and the quality of service loss without being aware of the current attack model in a dynamic privacy protection game. Simulation results show that this scheme can increase the privacy gain, decrease the quality of service loss, and thus improve the utility of the vehicle in comparison with a benchmark scheme. Weihang Wang 0001, Minghui Min, Liang Xiao 0003, Ye Chen 0011, Huaiyu Dai |
ICC | 3 |
| 2019 | Voltage Based Authentication for Controller Area Networks with Reinforcement LearningabstractController area networks (CANs) are vulnerable to spoofing attacks such as frame falsifying attacks, as electronic control units (ECUs) send and receive messages without any authentication and encryption. In this paper, we propose a physical authentication scheme that exploits the voltage features of the ECU signals on the CAN bus and applies reinforcement learning to choose the authentication mode such as the protection level and test threshold. This scheme enables a monitor node to optimize the authentication mode via trial-and-error without knowing the CAN bus signal model and spoofing model. Experimental results show that the proposed authentication scheme can significantly improve the authentication accuracy and response compared with a benchmark scheme. Tangwei Xu, Xiaozhen Lu, Liang Xiao 0003, Yuliang Tang, Huaiyu Dai |
ICC | 3 |
| 2019 | SMDP Based Cross-Area Resource Management for Vehicular Cloud NetworksabstractRecent years have witnessed the emerging concept of Vehicular Cloud Networks (VCN) with the development of Internet of Vehicles and cloud computing. There is a common phenomenon that the loads and resources of local clouds (LCs) in different regions are seriously unbalanced in dynamic VCN as LCs are confronted with the shortage of resources and vehicles are featured by high mobility. As a consequence, we propose a cross-area resource management scheme (CRMS) based on semi-Markov decision process (SMDP) to alleviate this problem. In this scheme, a service migration mechanism plays an imperative role, which the local cloud needs to take the service requests from both local and neighboring cloud into account when allocating computing resources that we focus on in this paper. Considering the impact of different types of service requests on system revenue, we obtain the optimal policy of resource allocation adaptively through SMDP to maximize the long-term expected reward. Numerical results show the performance of the proposed CRMS has been significantly improved. Zhuyue Yu, Jiayou Xie, Yuliang Tang, Liang Xiao 0003 |
VTC Spring | 4 |
| 2019 | Eliminating NB-IoT Interference to LTE System: A Sparse Machine Learning-Based ApproachabstractNarrowband Internet-of-Things (NB-IoT) is a competitive 5G technology for massive machine-type communication scenarios, but meanwhile introduces narrowband interference (NBI) to existing broadband transmission such as the Long Term Evolution (LTE) systems in enhanced mobile broadband (eMBB) scenarios. In order to facilitate the harmonic and fair coexistence in wireless heterogeneous networks, it is important to eliminate NB-IoT interference to LTE systems. In this paper, a novel sparse machine learning-based framework and a sparse combinatorial optimization problem is formulated for accurate NBI recovery, which can be efficiently solved using the proposed iterative sparse learning algorithm called sparse cross-entropy minimization (SCEM). To further improve the recovery accuracy and convergence rate, regularization is introduced to the loss function in the enhanced algorithm called regularized SCEM. Moreover, exploiting the spatial correlation of NBI, the framework is extended to multiple-input multiple-output systems. Simulation results demonstrate that the proposed methods are effective in eliminating NB-IoT interference to LTE systems, and significantly outperform the state-of-the-art methods. Sicong Liu 0002, Liang Xiao 0003, Zhu Han 0001, Yuliang Tang |
IEEE Internet Things J. | 2 |
| 2019 | Reinforcement Learning-Based Microgrid Energy Trading With a Reduced Power Plant ScheduleabstractWith dynamic renewable energy generation and power demand, microgrids (MGs) exchange energy with each other to reduce their dependence on power plants. In this article, we present a reinforcement learning (RL)-based MG energy trading scheme to choose the electric energy trading policy according to the predicted future renewable energy generation, the estimated future power demand, and the MG battery level. This scheme designs a deep RL-based energy trading algorithm to address the supply-demand mismatch problem for a smart grid with a large number of MGs without relying on the renewable energy generation and power demand models of other MGs. A performance bound on the MG utility and dependence on the power plant is provided. Simulation results based on a smart grid with three MGs using wind speed data from Hong Kong Observation and electricity prices from ISO New England show that this scheme significantly reduces the average power plant schedule and thus increases the MG utility in comparison with a benchmark methodology. Xiaozhen Lu, Xingyu Xiao, Liang Xiao 0003, Canhuang Dai, Mugen Peng, H. Vincent Poor |
IEEE Internet Things J. | 3 |
| 2019 | Learning-Based Privacy-Aware Offloading for Healthcare IoT With Energy HarvestingabstractMobile edge computing helps healthcare Internet of Things (IoT) devices with energy harvesting provide satisfactory quality of experiences for computation intensive applications. We propose a reinforcement learning (RL)-based privacy-aware offloading scheme to help healthcare IoT devices protect both the user location privacy and the usage pattern privacy. More specifically, this scheme enables a healthcare IoT device to choose the offloading rate that improves the computation performance, protects user privacy, and saves the energy of the IoT device without being aware of the privacy leakage, IoT energy consumption, and edge computation model. This scheme uses transfer learning to reduce the random exploration at the initial learning process and applies a Dyna architecture that provides simulated offloading experiences to accelerate the learning process. A post-decision state learning method uses the known channel state model to further improve the offloading performance. We provide the performance bound of this scheme regarding the privacy level, the energy consumption, and the computation latency for three typical healthcare IoT offloading scenarios. Simulation results show that this scheme can reduce the computation latency, save the energy consumption, and improve the privacy level of the healthcare IoT device compared with the benchmark scheme. Minghui Min, Xiaoyue Wan, Liang Xiao 0003, Ye Chen 0011, Minghua Xia, Di Wu 0001, Huaiyu Dai |
IEEE Internet Things J. | 3 |
| 2019 | Deep Reinforcement Learning-Enabled Secure Visible Light Communication Against EavesdroppingabstractThe inherent broadcast characteristics of the visible light communication (VLC) channel makes VLC downlinks susceptible to unauthorized terminals in many actual VLC scenarios, such as offices and shopping centers. This paper considers a multiple-input-single-output (MISO) VLC scenario with multiple light fixtures acting as the transmitter, a VLC receiver as the legitimate user, and an eavesdropper attempting to intercept the undisclosed information. To improve the confidentiality of VLC links, a physical-layer anti-eavesdropping framework is proposed to obscure the unauthorized eavesdroppers and diminishes their capability of inferring the information through smart beamforming over the MISO VLC wiretap channel. To cope with the intractable problem of finding the theoretically optimal solution of the secrecy rate and utility for the MISO VLC wiretapping channel, a reinforcement learning (RL)-based VLC beamforming control scheme is proposed to achieve the optimal beamforming policy against the eavesdropper. Furthermore, a deep RL-based VLC beamforming control scheme is proposed to handle the curse of dimensionality for both observation space and action space and avoid the quantization error of the RL-based algorithm. Simulation results show that the proposed learning-based VLC beamforming control schemes can significantly decrease the bit error rate of the legitimate receiver and increase the secrecy rate and utility of the anti-eavesdropping MISO VLC system, compared with the benchmark strategy. Liang Xiao 0003, Geyi Sheng, Sicong Liu 0002, Huaiyu Dai, Mugen Peng, Jian Song 0004 |
IEEE Trans. Commun. | 1 |
| 2019 | Learning Driven Computation Offloading for Asymmetrically Informed Edge ComputingabstractEdge computing emerges as a promising paradigm to decentralize computation power to the edge of the network and thus improve user experience by task offloading. A user can perfectly schedule his tasks to be executed on edge servers if the execution time of all tasks can be known beforehand. However, it is difficult to know the task execution time (TET) before performing actual offloading, which normally varies on edge servers with different software and hardware configurations. Moreover, such configuration information is not always available to end users due to security concerns. In this paper, we first propose a learning-driven algorithm to accurately predict TETs of all tasks in such an asymmetrically informed edge computing environment. The basic idea is to predict unknown TETs using only a small sampled set of TETs by exploiting the underlying correlation between TETs and edge server configurations. Next, we formulate the problem of task offloading into a constrained optimization problem, which is unfortunately proved to be NP-hard. To address the above challenge, we design a task offloading algorithm, called Maximum Efficiency First Ordered (MEFO), to achieve near-optimal efficiency. Field measurements and experiments have been conducted to demonstrate that our proposed learning-driven algorithm can predict TETs more accurately than other algorithms as long as the fraction of sampled TETs is larger than a small predefined threshold, and our proposed MEFO algorithm achieves a much higher success rate of task offloading and a shorter processing delay with very limited information of edge servers. Miao Hu 0001, Di Wu 0001, Yipeng Zhou, Xu Chen 0004, Liang Xiao 0003 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2018 | Reinforcement Learning Based Power Control for VANET Broadcast against JammingabstractBroadcast of critical information such as emergency traffic messages in vehicular ad hoc networks (VANETs) has to address jamming with dynamic network topology. In this paper, we propose a deep reinforcement learning based cooperative power control scheme for VANET broadcast against reactive jammers who can observe the ongoing broadcast states. The neural episodic control based cooperative power control scheme uses the convolutional neural network and differentiate neural dictionary to accelerate the learning speed for the VANETs with dynamic topology. Simulation results have shown that the proposed scheme can effectively improve the packet delivery rate and reduce the energy consumption of the broadcast compared with other power control schemes. Canhuang Dai, Xingyu Xiao, Liang Xiao 0003, Peng Cheng 0001 |
GLOBECOM | 3 |
| 2018 | Learning Based Power Control for mmWave Massive MIMO against JammingabstractMillimeter-wave (mmWave) massive multiple-input multiple-output (MIMO) systems have to address smart jammers that use smart radio devices to choose the jamming policy with the goal of interrupting the ongoing transmissions. In this paper, we propose a reinforcement learning based power control strategy for the downlink mmWave massive MIMO systems. More specifically, we present a fast policy hill-climbing based power control algorithm for a base station to choose the transmit power over multiple antennas. Based on the signal-to-interference-plus-noise ratio (SINR) of the signals and the jamming strength, we evaluate the impact of the number of transmit antennas on the communication performance. Simulation results verify that the proposed schemes can increase the average SINR, sum data rate and the utility of the mmWave massive MIMO against smart jamming compared with the benchmark strategy. Zhongcheng Xiao, Sicong Liu 0002, Liang Xiao 0003 |
GLOBECOM | 4 |
| 2018 | Reinforcement Learning-Based Interference Control for Ultra-Dense Small CellsabstractThe densification deployment of small cells emerging into 5G cellular networks can achieve high capacity, but is faced with the challenge of how to manage energy consumption and inter-cell interference well in time-varying channels. In this paper, we propose a reinforcement learning based downlink power control algorithm to manage interference for the ultra-dense small cell networks. More specifically, base stations of the small cells use Q-learning to select the downlink transmit powers. A transfer learning method called hotbooting is applied to further accelerate the learning speed and save the energy consumption based on the estimated user density without being aware of the network and channel model of the other small cells. Simulation results demonstrate this scheme significantly improves the network throughput and saves the energy consumption compared with the benchmark, a data-driven based transmission power adaptation scheme. Hailu Zhang, Minghui Min, Liang Xiao 0003, Sicong Liu 0002, Peng Cheng 0001, Mugen Peng |
GLOBECOM | 3 |
| 2018 | Learning-Based Rogue Edge Detection in VANETs with Ambient Radio SignalsabstractEdge computing for mobile devices in vehicular ad hoc networks (VANETs) has to address rogue edge attacks, in which a rogue edge node claims to be the serving edge in the vehicle to steal user secrets and help launch other attacks such as man-in-the-middle attacks. Rogue edge detection in VANETs is more challenging than the spoofing detection in indoor wireless networks due to the high mobility of onboard units (OBUs) and the large-scale network infrastructure with roadside units (RSUs). In this paper, we propose a physical (PHY)- layer rogue edge detection scheme for VANETs according to the shared ambient radio signals observed during the same moving trace of the mobile device and the serving edge in the same vehicle. In this scheme, the edge node under test has to send the physical properties of the ambient radio signals, including the received signal strength indicator (RSSI) of the ambient signals with the corresponding source media access control (MAC) address during a given time slot. The mobile device can choose to compare the received ambient signal properties and its own record or apply the RSSI of the received signals to detect rogue edge attacks, and determines test threshold in the detection. We adopt a reinforcement learning technique to enable the mobile device to achieve the optimal detection policy in the dynamic VANET without being aware of the VANET model and the attack model. Simulation results show that the Q-learning based detection scheme can significantly reduce the detection error rate and increase the utility compared with existing schemes. Xiaozhen Lu, Xiaoyue Wan, Liang Xiao 0003, Yuliang Tang, Weihua Zhuang |
ICC | 3 |
| 2018 | DQN-Based Power Control for IoT Transmission against JammingabstractInternet of Things (IoTs) have to address jammers, with goal to interrupt the communication of the energy- constrained IoT devices and sometimes even cause denial-of-service attacks. In this paper, we propose a deep reinforcement learning based power control scheme for IoT devices to improve the transmission efficiency and save energy. This scheme depends on the current IoT transmission status and the jamming strength and applies deep Q-network (DQN) to determine the transmit power without being aware of the IoT topology and the jamming model. This scheme is implemented on the universal software radio peripherals for the anti- jamming communication performance evaluation. Experimental results show that this scheme improves the signal-to-interference-plus-noise of the IoT signals compared with the benchmark Q-learning based power control scheme against jamming. Ye Chen 0011, Yanda Li, Dongjin Xu, Liang Xiao 0003 |
VTC Spring | 4 |
| 2018 | Learning-Based Defense against Malicious Unmanned Aerial VehiclesabstractAdversary unmanned aerial vehicles (UAVs) seriously threaten public security and user privacy. In this paper, we propose a reinforcement learning (RL) based defense framework to address malicious UAVs close to a target estate such as a company or an institute. This framework uses Q-learning to choose the defense policy such as jamming the global positioning system signals (GPS) and hacking, and laser shooting. According to the defense history and the current security status of the target estate, this scheme can improve the UAV defense performance in the dynamic game without being aware of the UAV attack policy and environment model in the area of interests. Simulation results show that this scheme can reduce the risk rate of the estate and improve the utility compared with the benchmark scheme against malicious UAVs. Minghui Min, Liang Xiao 0003, Dongjin Xu, Lianfen Huang, Mugen Peng |
VTC Spring | 2 |
| 2018 | Utility Maximization of Cloud-Based In-Car Video Recording Over Vehicular Access NetworksabstractWith the advance of cloud computing and 4G/5G technology, video contents recorded by in-car cameras (i.e., vehicular digital video recorders) can be uploaded to the cloud to facilitate accident analysis, online surveillance, video sharing, etc. However, the cost of uploading such huge volume of video contents via unstable vehicular access networks (including cellular base stations and road-side units) can be considerable by considering the increasing video quality requirement, time constraint, and limited local buffer space. In this paper, we propose an adaptive video recording and uploading scheme to maximize the overall utility of cloud-based in-car video uploading over vehicular access networks. Specifically, the utility function is defined as the weighted sum of bandwidth cost and video quality and we formulate the problem into a constrained Markov decision process (MDP). Based on the theoretic foundation of MDP, we design and implement an algorithm to obtain an adaptive chunk uploading policy for video contents over vehicular access networks. Extensive simulations have been conducted to demonstrate that our policy can achieve the best performance compared with other alternative strategies. Zhaobin Deng, Yipeng Zhou, Di Wu 0001, Guoqiao Ye, Min Chen 0003, Liang Xiao 0003 |
IEEE Internet Things J. | 6 |
| 2018 | Defense Against Advanced Persistent Threats in Dynamic Cloud Storage: A Colonel Blotto Game ApproachabstractAdvanced persistent threat (APT) attackers apply multiple sophisticated methods to continuously and stealthily steal information from the targeted cloud storage systems and can even induce the storage system to apply a specific defense strategy and attack it accordingly. In this paper, the interactions between an APT attacker and a defender allocating their central processing units (CPUs) over multiple storage devices in a cloud storage system are formulated as a Colonel Blotto game. The Nash equilibria of the CPU allocation game are derived for both symmetric and asymmetric CPUs between the APT attacker and the defender to evaluate how the limited CPU resources, the data storage size and the number of storage devices impact the expected data protection level and the utility of the cloud storage system. A CPU allocation scheme based on “hotbooting” policy hill-climbing that exploits the experiences in similar scenarios to initialize the quality values to accelerate the learning speed is proposed for the defender to achieve the optimal APT defense performance in the dynamic game without being aware of the APT attack model and the data storage model. A hotbooting deep${Q}$-network-based CPU allocation scheme further improves the APT detection performance for the case with a large number of CPUs and storage devices. Simulation results show that our proposed reinforcement learning-based CPU allocation can improve both the data protection level and the utility of the cloud storage system compared with the${Q}$-learning-based CPU allocation against APTs. Minghui Min, Liang Xiao 0003, Caixia Xie, Mohammad Hajimirsadeghi, Narayan B. Mandayam |
IEEE Internet Things J. | 2 |
| 2018 | A Secure Mobile Crowdsensing Game With Deep Reinforcement LearningabstractMobile crowdsensing (MCS) is vulnerable to faked sensing attacks, as selfish smartphone users sometimes provide faked sensing results to the MCS server to save their sensing costs and avoid privacy leakage. In this paper, the interactions between an MCS server and a number of smartphone users are formulated as a Stackelberg game, in which the server as the leader first determines and broadcasts its payment policy for each sensing accuracy. Each user as a follower chooses the sensing effort and thus the sensing accuracy afterward to receive the payment based on the payment policy and the sensing accuracy estimated by the server. The Stackelberg equilibria of the secure MCS game are presented, disclosing conditions to motivate accurate sensing. Without knowing the smartphone sensing models in a dynamic version of the MCS game, an MCS system can apply deep Q-network (DQN), which is a deep reinforcement learning technique combining reinforcement learning and deep learning techniques, to derive the optimal MCS policy against faked sensing attacks. The DQN-based MCS system uses a deep convolutional neural network to accelerate the learning process with a high-dimensional state space and action set, and thus improve the MCS performance against selfish users. Simulation results show that the proposed MCS system stimulates high-quality sensing services and suppresses faked sensing attacks, compared with a Q-learning-based MCS system. Liang Xiao 0003, Yanda Li, Guoan Han, Huaiyu Dai, H. Vincent Poor |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2018 | Attacker-Centric View of a Detection Game against Advanced Persistent ThreatsabstractAdvanced persistent threats (APTs) are a major threat to cyber-security, causing significant financial and privacy losses each year. In this paper, cumulative prospect theory (CPT) is applied to study the interactions between a cyber system and an APT attacker when each of them makes subjective decisions to choose their scan interval and attack interval, respectively. Both the probability distortion effect and the framing effect are applied to model the deviation of subjective decisions of end-users from the objective decisions governed by expected utility theory, under uncertain attack durations in a pure-strategy game and scan interval in a mixed-strategy game. The CPT-based APT detection game incorporates both the probability weighting distortion and the framing effect of the subjective attacker and security agent of the cyber system, rather than discrete decision weights, as in earlier prospect theoretic study of APT detection. The Nash equilibria of the APT detection game are derived, showing that a subjective attacker becomes risk-seeking if the frame of reference for evaluating the utility is large, and becomes risk-averse if the frame of reference for evaluating the utility is small. A policy hill-climbing (PHC) based detection scheme is proposed to increase the policy uncertainty to fool the attacker in the dynamic game, and a “hotbooting” technique that exploits experiences in similar scenarios to initialize the quality values is developed to accelerate the learning speed of PHC-based detection. A practical example of a mobile network is presented to evaluate the performance of the proposed detection strategy. Simulation results show that the proposed strategy can improve detection performance with a higher data protection level and utilities of the cloud in the presence of an attacker compared with a standard Q-learning strategy. Liang Xiao 0003, Dongjin Xu, Narayan B. Mandayam, H. Vincent Poor |
IEEE Trans. Mob. Comput. | 1 |
| 2018 | PHY-Layer Authentication With Multiple Landmarks With Reduced OverheadabstractPhysical (PHY)-layer authentication systems can exploit channel state information of radio transmitters to detect spoofing attacks in wireless networks. The use of multiple landmarks each with multiple antennas enhances the spatial resolution of radio transmitters, and thus improves the spoofing detection accuracy of PHY-layer authentication. Unlike most existing PHY-layer authentication schemes that apply hypothesis tests and rely on the known radio channel model, we propose a logistic regression-based authentication to remove the assumption on the known channel model, and thus be applicable to more generic wireless networks. The Frank-Wolfe algorithm is used to estimate the parameters of the logistic regression model, in which the convex problem under a ℓ1-norm constraint is solved for weight sparsity to avoid over-fitting in the learning process. We design a distributed Frank-Wolfe-based PHY-layer authentication to further reduce the communication overhead between the landmarks and the security agent. Then, we construct an incremental aggregated gradient-based scheme to provide online authentication with a higher accuracy and lower computation overhead. Simulation and experimental results validate the accuracy of the proposed authentication schemes, and show the reduced communication and computation overheads. Liang Xiao 0003, Xiaoyue Wan, Zhu Han 0001 |
IEEE Trans. Wirel. Commun. | 1 |
| 2017 | Anti-Jamming Communication Game for UAV-Aided VANETsabstractVehicular ad-hoc networks (VANETs) are vulnerable to jamming attacks, and frequency hopping-based anti- jamming techniques are not always applicable in VANETs due to the high mobility of the onboard units (OBUs) especially under a large scale network topology. In this paper, we use unmanned aerial vehicles (UAVs) to deal with VANET jamming, especially smart jamming that changes the jamming policy based on the ongoing communication status of the VANET. More specifically, the UAV relays the data of OBUs to another roadside unit (RSU) with a better transmission condition if the serving RSU is located in a heavily jammed area. The interactions between the UAV and the jammer are formulated as an anti-jamming UAV relay game, in which the UAV decides whether or not to relay the data of the OBU to another RSU that is far away from the jammer, and the latter chooses the jamming power. The Nash equilibria (NE) of the game are derived to reveal how the best UAV relay strategy depends on the transmission cost and the radio channel model. A hotbooting policy hill climbing (PHC)-based UAV relay strategy is proposed to address jamming in the dynamic UAV-aided VANET game without the knowledge of network model and jamming model. Simulation results show that the proposed relay strategy can efficiently reduce the bit error rate (BER) of OBU data and thus increase the utility of VANET in comparison with a Q-learning based scheme. Xiaozhen Lu, Dongjin Xu, Liang Xiao 0003, Lei Wang 0009, Weihua Zhuang |
GLOBECOM | 3 |
| 2017 | Anti-Jamming Power Control Game in Unmanned Aerial Vehicle NetworksabstractIn this paper, the anti-jamming issue in unmanned aerial vehicle (UAV) networks is analyzed in a static game and a dynamic game. We investigate the effect of wireless channel fading characteristics from a UAV to a ground station and flying cost on the performance of a closed-form Nash equilibrium (NE) in the static game. Besides, in a Stackelberg dynamic game, wherein the system model is hard to determine, we propose a Q- learning based anti-jamming scheme and evaluate its performance via exhaustive simulations, which can achieve relatively higher average utility and Signal to Interference plus Noise Ratio (SINR) than a benchmark method. Shichao Lv, Liang Xiao 0003, Xiaoshan Wang, Changzhen Hu, Limin Sun 0001 |
GLOBECOM | 2 |
| 2017 | Reinforcement Learning Based Mobile Offloading for Cloud-Based Malware DetectionabstractCloud-based malware detection improves the detection performance for mobile devices that offload their malware detection tasks to security servers with much larger malware database and powerful computational resources. In this paper, we investigate the competition of the radio transmission bandwidths and the data sharing of the security server in the dynamic malware detection game, in which each mobile device chooses its offloading rate of the application traces to the security server. As the Q-learning technique has a slow learning rate in the game with high dimension, we have designed a mobile malware detection based on hotbooting-Q techniques, which initiates the quality values based on the malware detection experience. We propose an offloading strategy based on deep Q-network technique with a deep convolutional neural network to further improve the detection speed, the detection accuracy, and the utility. Preliminary simulation results verify the detection gain of the scheme compared with the Q- learning based strategy. Xiaoyue Wan, Geyi Sheng, Yanda Li, Liang Xiao 0003, Xiaojiang Du |
GLOBECOM | 4 |
| 2017 | Two-dimensional anti-jamming communication based on deep reinforcement learningabstractIn this paper, a two-dimensional anti-jamming communication scheme for cognitive radio networks is developed, in which a secondary user (SU) exploits both spread spectrum and user mobility to address jamming attacks, while not interfering with primary users. By applying a deep Q-network algorithm, this scheme determines whether to recommend that the SU leave an area of heavy jamming and chooses a frequency hopping pattern to defeat smart jammers. Without knowing the jamming model and the radio channel model, the SU derives an optimal anti-jamming communication policy using Q-learning in a proposed dynamic game, and applies a deep convolution neural network to accelerate the learning speed with a large number of frequency channels. The proposed scheme can increase the signal-to-interference-plus-noise ratio and improve the utility of the SU against cooperative jamming, compared with a Q-learning-only based benchmark system. Guoan Han, Liang Xiao 0003, H. Vincent Poor |
ICASSP | 2 |
| 2017 | Game theoretic study of protecting MIMO transmissions against smart attacksabstractMultiple-input multiple-output (MIMO) systems are threatened by smart attackers, who apply programmable radio devices such as software defined radios to perform multiple types of attacks such as eavesdropping, jamming and spoofing. In this paper, MIMO transmission in the presence of smart attacks is formulated as a noncooperative game, in which a MIMO transmitter chooses its transmit power level and a smart attacker determines its attack type accordingly. A Nash equilibrium of this secure MIMO transmission game is derived and conditions assuring its existence are provided to reveal the impact of the number of antennas and the costs of the attacker to launch each type of attack. A power control strategy based on Q-learning is proposed for the MIMO transmitter to suppress the attack motivation of smart attackers in a dynamic version of MIMO transmission game without being aware of the attack and the radio channel model. Simulation results show that our proposed scheme can reduce the attack rate of smart attackers and improve the secrecy capacity compared with the benchmark strategy. Yanda Li, Liang Xiao 0003, Huaiyu Dai, H. Vincent Poor |
ICC | 2 |
| 2017 | Defense against advanced persistent threats: A Colonel Blotto game approachabstractAn Advanced Persistent Threat (APT) attacker applies multiple sophisticated methods to continuously and stealthily attack targeted cyber systems. In this paper, the interactions between an APT attacker and a cloud system defender in their allocation of the Central Processing Units (CPUs) over multiple devices are formulated as a Colonel Blotto game (CBG), which models the competition of two players under given resource constraints over multiple battlefields. The Nash equilibria (NEs) of the CBG-based APT defense game are derived for the case with symmetric players and the case with asymmetric players each with different total number of CPUs. The expected data protection level and the utility of the defender are provided for each game at the NE. An APT defense strategy based on the policy hill-climbing (PHC) algorithm is proposed for the defender to achieve the optimal CPU allocation distribution over the devices in the dynamic defense game without being aware of the APT attack model. Simulation results have verified the efficacy of our proposed algorithm, showing that both the data protection level and the utility of the defender are improved compared with the benchmark greedy allocation algorithm. Minghui Min, Liang Xiao 0003, Caixia Xie, Mohammad Hajimirsadeghi, Narayan B. Mandayam |
ICC | 2 |
| 2017 | FHY-layer authentication with multiple landmarks with reduced communication overheadabstractIn this paper, we propose a physical (PHY)-layer authentication system that exploits the channel state information of radio transmitters to detect spoofing attacks in wireless networks. By using multiple landmarks and multiple antennas in the channel estimation, this authentication system enhances the spatial resolution of the channel information and thus improves the spoofing detection accuracy. Unlike most existing hypothesis test based PHY-layer authentication schemes that rely on the known radio channel model, our proposed authentication system uses the logistic regression to remove the assumption on the known channel model and is applicable to more generic wireless networks. The Frank-Wolfe algorithm is then used to estimate the parameters of the logistic regression model, which solves the convex problem under a ℓ1-norm constraint for weight sparsity to avoid over-fitting in the learning process. The distributed Frank-Wolfe algorithm can further reduce the communication overhead between the landmarks and the security agent while keeping the spoofing detection accuracy. Simulation results can validate the accuracy of the proposed PHY-layer authentication with multiple landmarks and show the performance gain regarding the overall communication overhead. Xiaoyue Wan, Liang Xiao 0003, Qiangda Li, Zhu Han 0001 |
ICC | 2 |
| 2017 | Defense Against Advanced Persistent Threats with Expert System for Internet of Things
Shichao Lv, Zhiqiang Shi, Limin Sun 0001, Liang Xiao 0003 |
WASA | 5 |
| 2017 | Cloud Storage Defense Against Advanced Persistent Threats: A Prospect Theoretic StudyabstractCloud storage is vulnerable to advanced persistent threats (APTs), in which an attacker launches stealthy, continuous, and targeted attacks on storage devices. In this paper, prospect theory (PT) is applied to formulate the interaction between the defender of a cloud storage system and an APT attacker who makes subjective decisions that sometimes deviate from the results of expected utility theory, which is a basis of traditional game theory. In the PT-based cloud storage defense game with pure strategy, the defender chooses a scan interval for each storage device and the subjective APT attacker chooses his or her interval of attack against each device. A mixed-strategy subjective storage defense game is also investigated, in which each subjective defender and APT attacker acts under uncertainty about the action of its opponent. The Nash equilibria (NEs) of both games are derived, showing that the subjective view of an APT attacker can improve the utility of the defender. A Q-learning-based APT defense scheme that the storage defender can apply without being aware of the APT attack model or the subjectivity model of the attacker in the dynamic APT defense game is also proposed. Simulation results show that the proposed defense scheme suppresses the attack motivation of subjective APT attackers and improves the utility of the defender, compared with the benchmark greedy defense strategy. Liang Xiao 0003, Dongjin Xu, Caixia Xie, Narayan B. Mandayam, H. Vincent Poor |
IEEE J. Sel. Areas Commun. | 1 |
| 2017 | Active authentication with reinforcement learning based on ambient radio signals
Jinliang Liu 0002, Liang Xiao 0003, Guolong Liu |
Multim. Tools Appl. | 2 |
| 2017 | Cloud-Based Malware Detection Game for Mobile Devices with OffloadingabstractAs accurate malware detection on mobile devices requires fast process of a large number of application traces, cloud-based malware detection can utilize the data sharing and powerful computational resources of security servers to improve the detection performance. In this paper, we investigate the cloud-based malware detection game, in which mobile devices offload their application traces to security servers via base stations or access points in dynamic networks. We derive the Nash equilibrium (NE) of the static malware detection game and present the existence condition of the NE, showing how mobile devices share their application traces at the security server to improve the detection accuracy, and compete for the limited radio bandwidth, the computational and communication resources of the server. We design a malware detection scheme with Q-learning for a mobile device to derive the optimal offloading rate without knowing the trace generation and the radio bandwidth model of other mobile devices. The detection performance is further improved with the Dyna architecture, in which a mobile device learns from the hypothetical experience to increase its convergence rate. We also design a post-decision state learning-based scheme that utilizes the known radio channel model to accelerate the reinforcement learning process in the malware detection. Simulation results show that the proposed schemes improve the detection accuracy, reduce the detection delay, and increase the utility of a mobile device in the dynamic malware detection game, compared with the benchmark strategy. Liang Xiao 0003, Yanda Li, Xueli Huang, Xiaojiang Du |
IEEE Trans. Mob. Comput. | 1 |
| 2016 | Channel-Based Authentication Game in MIMO SystemsabstractIn this paper, we investigate the PHY-layer authentication that exploits radio channel information to detect spoofing attacks in multiple- input multiple-output (MIMO) systems. We formulate the interactions between a receiver and a spoofing node in the spoofing detection as a zero-sum game. In this game, the receiver chooses the test threshold of the hypothesis test in the PHY-layer authentication to maximize its utility based on the Bayesian risk in the spoofing detection, while the adversary chooses its attack frequency, i.e., how often a spoofing packet is sent over multiple antennas. The unique Nash equilibrium of the static MIMO authentication game is derived and the condition for its existence is discussed. We investigate the impact of the number of antennas on the performance of the dynamic authentication game. We propose a PHY-layer spoofing detection based on Q-learning for MIMO systems to achieve the optimal test threshold in the spoofing detection via trials, and implement it over universal software radio peripherals. The performance of the spoofing detection algorithm is evaluated via experiments in indoor environments. Liang Xiao 0003, Tianhua Chen, Guoan Han, Weihua Zhuang, Limin Sun 0001 |
GLOBECOM | 1 |
| 2016 | Prospect Theoretic Study of Cloud Storage Defense against Advanced Persistent ThreatsabstractCloud storage is vulnerable to Advanced Persistent Threats (APTs), which are stealthy, continuous, well funded and targeted. In this paper, prospect theory is applied to study the interactions between a subjective cloud storage defender and a subjective APT attacker. Two subjective APT games are formulated, in which the defender chooses its interval to scan the storage device and the attacker decides its duration between launching two attacks under uncertain APT attack durations and action of the opponent, respectively. The Nash equilibria of the static subjective APT games are derived. We also study the dynamic APT game and propose a Q-learning based APT defense strategy for cloud storage. Simulation results show that the APT defense benefits from the subjective view of the attacker and the proposed defense strategy can improve detection performance with a higher utility. Dongjin Xu, Yanda Li, Liang Xiao 0003, Narayan B. Mandayam, H. Vincent Poor |
GLOBECOM | 3 |
| 2016 | Secure routing and resource allocation based on game theory in cooperative cognitive radio networksabstractSummary The era of big data is here now, and spectrum resources are increasingly scarce in heterogeneous network environment. The spectrum efficiency and secure transmission of big data are important issues. Cognitive radio has been proposed to address the issue of spectrum efficiency, and is a hot topic in the literatures. In multi‐hop cooperative cognitive radio networks (CCRNs), secondary users need the primary users' authorization to be relays. Most existing centralized route selection schemes ignore the energy allocation, and thus are inefficient. Moreover, the incomplete of information in multi‐hop network leads to many difficulties in cooperation. Inspired by the game theory, a novel strategy is proposed in this paper to defend against insider attacks based on trust. This strategy is denoted as secure routing and resource allocation based on game theory in CCRNs (SRGC). With a reputation updating process and distributed learning algorithm, the proposed strategy can find a ‘best’ route, which is relatively safe for each primary transmitter, and at the same time fully utilizes the spectrum and energy. Using NS2, simulations indicate that SRGC can well fit into CCRNs, improve the network performance and defend against the routing disruption attacks. Compared with other schemes, the SRGC results in a performance with better adaptability to the distributed environment. Moreover, SRGC can maximize the average throughput and minimize the data drop ratios. Copyright © 2015 John Wiley & Sons, Ltd. He Fang, Li Xu 0002, Liang Xiao 0003 |
Concurr. Comput. Pract. Exp. | 3 |
| 2015 | Jamming Games in Underwater Sensor Networks with Reinforcement LearningabstractJamming attacks that can further lead to denial of service attacks have thrown serious threats to underwater sensor networks (UWSNs). However, due to the narrow bandwidth of underwater acoustic signals and time variant propagation environments, jamming in UWSNs cannot be fully addressed by spread spectrum techniques, one type of widely-used antijamming methods in wireless networks for decades. In this work, we investigate jamming attacks in underwater sensor networks. More specifically, the interactions between the underwater sensors and jammers in UWSNs are formulated as an underwater jamming game, in which the players choose their transmit power levels to maximize their individual utilities based on the signal to interference plus noise ratio of the legal signals and transmission costs. The Nash equilibrium (NE) of a static jamming game is presented in a closed-form expression for the jamming scenario with known acoustic channel gains. For the dynamic and unknown underwater environments, we propose a reinforcement learning-based anti-jamming method for UWSNs, in which each sensor chooses its transmit power without knowing the channel gain of the jammers. Simulations are performed to evaluate the NE in the static jamming game in underwater sensor networks and to validate the efficacy of the proposed anti-jamming power control scheme against jamming in dynamic environments. Liang Xiao 0003, Qiangda Li, Tianhua Chen, En Cheng, Huaiyu Dai |
GLOBECOM | 1 |
| 2015 | Spoofing Detection with Reinforcement Learning in Wireless NetworksabstractIn this paper, we investigate the PHY-layer authentication in wireless networks, which exploits PHY-layer channel information such as the received signal strength indicators to detect spoofing attacks. The interactions between a legitimate receiver node and a spoofer are formulated as a PHY- authentication game. More specifically, the receiver chooses the test threshold in the hypothesis test of the spoofing detection to maximize its expected utility based on Bayesian risk to detect the spoofer. On the other hand, the spoofing node decides its attack strength, i.e., the frequency to send a spoofing packet that claims to use another node's MAC address, based on its individual utility in the zero-sum game. As it is challenging for most radio nodes to obtain the exact channel models in advance in a dynamic radio environment, we propose a spoofing detection scheme based on reinforcement learning techniques, which achieves the optimal test threshold in the spoofing detection via Q-learning and implement it over universal software radio peripherals (USRP). Experimental results are presented to validate its efficiency in spoofing detection. Liang Xiao 0003, Yan Li 0076, Guolong Liu, Qiangda Li, Weihua Zhuang |
GLOBECOM | 1 |
| 2015 | Secure mobile crowdsensing gameabstractBy recruiting sensor-equipped smartphone users to report sensing data, mobile crowdsensing (MCS) provides location-based services such as environmental monitoring. However, due to the distributed and potentially selfish nature of smartphone users, mobile crowdsensing applications are vulnerable to faked sensing attacks by users who bid a low price in an MCS auction and provide faked sensing reports to save sensing costs and avoid privacy leakage. In this paper, the interactions among an MCS server and smartphone users are formulated as a mobile crowdsensing game, in which each smartphone user chooses its sensing strategy such as its sensing time and energy to maximize its expected utility while the MCS server classifies the received sensing reports and determines the payment strategy accordingly to stimulate users to provide accurate sensing reports. Nash equilibrium (NE) of a static MCS game is evaluated and a closed-form expression for the NE in a special case is presented. Moreover, a dynamic mobile crowdsensing game is investigated, in which the sensing parameters of a smartphone are unknown by the server and the other users. A Q-learning discriminated pricing strategy is developed for the server to determine the payment to each user. Simulation results show that the proposed pricing mechanism stimulates users to provide high-quality sensing services and suppress faked sensing attacks. Liang Xiao 0003, Jinliang Liu 0002, Qiangda Li, H. Vincent Poor |
ICC | 1 |
| 2015 | Collaborative Anti-Jamming Broadcast with Uncoordinated Frequency Hopping over USRPabstractCognitive radio networks (CRNs) are threatened by smart jammers that aim to block the ongoing transmissions of secondary users according to their transmission patterns obtained from public control channels or compromised secondary users. Without requiring any pre-shared PHY-layer keys at the receivers, the uncoordinated frequency hopping (UFH) technique that was proposed to address smart jammers suffers from a low communication efficiency. In this paper, we develop a UFHbased collaborative broadcast system over universal software radio peripherals (USRPs), which exploits the node cooperation to improve the broadcast efficiency and jamming resistance of CRNs. Experiments are performed over USRPs to evaluate the broadcast performance against smart jammers under various network topologies. We investigate the impact of the CRN bandwidth, the prediction accuracy of smart jammers regarding the CRN frequency hopping pattern, and the jamming power and signal pattern. Experimental results show that the proposed broadcast system is more robust than three benchmark broadcast systems in most jamming scenarios. Guiquan Chen, Yan Li 0076, Liang Xiao 0003, Lianfen Huang |
VTC Spring | 3 |
| 2015 | Jamming Detection of Smartphones for WiFi SignalsabstractIn this paper, we investigate the impact of jamming attacks on the performance of smartphones regarding their WiFi access and propose a real-time jamming detection method based on the received signal strength indicator and the packet loss rate of WiFi signals, which can be easily implemented on Android smartphones. Experiments are performed to evaluate the proposed jamming detection method, in which universal software radio peripherals are used as jammers to block the WiFi signals between smartphone phones and wireless routers. Experimental results show that the proposed application can detect jamming attacks with small false alarm rate and miss detection raaaaaate. Guolong Liu, Jinliang Liu 0002, Yan Li 0076, Liang Xiao 0003, Yuliang Tang |
VTC Spring | 4 |
| 2015 | User-Centric View of Jamming Games in Cognitive Radio NetworksabstractJamming games between a cognitive radio enabled secondary user (SU) and a cognitive radio enabled jammer are considered, in which end-user decision making is modeled using prospect theory (PT). More specifically, the interactions between a user and a smart jammer regarding their respective choices of transmit power are formulated as a game under the assumption that the end-user decision making under uncertainty does not follow the traditional objective assumptions stipulated by expected utility theory, but rather follows the subjective deviations specified by PT. Two PT-based static jamming games are formulated to describe how subjective SU and jammer choose their transmit power to maximize their individual signal-to-interference-plus-noise ratio (SINR)-based utilities under uncertainties regarding the opponent's actions and channel states, respectively. The Nash equilibria of the games are presented under various channel models and transmission costs. Moreover, a PT-based dynamic jamming game is presented to investigate the long-term interactions between a subjective and a smart jammer according to a Markov decision process with uncertainty on the SU's future actions and the channel variations. Simulation results show that the subjective view of an SU tends to exaggerate the jamming probabilities and decreases its transmission probability, thus reducing the average SINR. On the other hand, the subjectivity of a jammer tends to reduce its jamming probability, and thus increases the SU throughput. Liang Xiao 0003, Jinliang Liu 0002, Qiangda Li, Narayan B. Mandayam, H. Vincent Poor |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2015 | Power control with reinforcement learning in cooperative cognitive radio networks against jamming
Liang Xiao 0003, Yan Li 0076, Jinliang Liu 0002 |
J. Supercomput. | 1 |
| 2014 | Prospect theoretic analysis of anti-jamming communications in cognitive radio networksabstractAn anti-jamming communication game between a cognitive radio enabled secondary user (SU) and a cognitive radio enabled jammer is considered, in which end-user decision making is modeled using prospect theory (PT). More specifically, the interactions between a user and a smart jammer (i.e., their respective choices of transmission probability) are formulated as a game under the assumption that end-user decision making under uncertainty does not follow the traditional objective assumptions stipulated by expected utility theory (EUT), but rather follows the subjective deviations specified by PT. Under the assumption that the capacity of the system is governed by the primary user activity, the Nash equilibria of the game are characterized under various conditions and the impact of the players' subjectivity (deviation from EUT behavior) on the SU's throughput is measured. Simulation results show that the subjective view of an SU tends to exaggerate the jamming probabilities and decreases its transmission probability, thus reducing the average throughput. On the other hand, the subjectivity of a jammer tends to reduce its jamming probability and thus increases the SU throughput. Liang Xiao 0003, Jinliang Liu 0002, Yan Li 0076, Narayan B. Mandayam, H. Vincent Poor |
GLOBECOM | 1 |
| 2014 | Anti-cheating prosumer energy exchange based on indirect reciprocityabstractEquipped with renewable energy generators, pro-sumers produce, store and consume energy and become an important entity in smart grids. With time-variant power demands and generation capacities, prosumers can exchange energy via local power lines with less transmission loss than the long-distance power lines to the traditional power plants. In this paper, we develop a local power exchange game and address the cheating problem that seller prosumers actually supply insufficient energy instead of the promised amount in the trade. We apply the indirect reciprocity principle to build an anti-cheating local energy exchange system that maintains reputations for the prosumers according to their selling histories. By supplying less energy to the buyers who cheated in previous trades, this system motivates the autonomous prosumers to cooperate instead of cheating. Simulation results have shown that this strategy significantly reduces the number of cheating prosumers and improves the overall system utility. Compared with the direct reciprocity-based counterpart, this system reaches the desirable equilibrium much faster and requires less energy from the traditional energy plants. Liang Xiao 0003, Yan Chen 0007, K. J. Ray Liu |
ICC | 1 |
| 2013 | Proximity-based security using ambient radio signalsabstractIn this paper, we propose a privacy-preserving proximity-based security strategy for location-based services in wireless networks, without requiring any pre-shared secret, trusted authority or public key infrastructure. More specifically, radio clients build their location tags according to the unique physical features of their ambient radio signals, which cannot be forged by attackers outside the proximity range. The proximity-based authentication and session key generation is based on the public location tag, which incorporates the received signal strength indicator (RSSI), sequence number and MAC address of the ambient radio packets. Meanwhile, as the basis for the session key generation, the secret location tag consisting of the arrival time interval of the ambient packets, is never broadcast, making it robust against eavesdroppers and spoofers. The proximity test utilizes the nonparametric Bayesian method called infinite Gaussian mixture model, and provides range control by selecting different features of various ambient radio sources. The authentication accuracy and key generation rate are evaluated via experiments using laptops in typical indoor environments. Liang Xiao 0003, Qiben Yan 0001, Wenjing Lou, Y. Thomas Hou 0001 |
ICC | 1 |
| 2013 | Proximity-Based Security Techniques for Mobile Users in Wireless NetworksabstractIn this paper, we propose a privacy-preserving proximity-based security system for location-based services in wireless networks, without requiring any pre-shared secret, trusted authority, or public key infrastructure. In this system, the proximity-based authentication and session key establishment are implemented based on spatial temporal location tags. Incorporating the unique physical features of the signals sent from multiple ambient radio sources, the location tags cannot be easily forged by attackers. More specifically, each radio client builds a public location tag according to the received signal strength indicators, sequence numbers, and media access control (MAC) addresses of the ambient packets. Each client also keeps a secret location tag that consists of the packet arrival time information to generate the session keys. As clients never disclose their secret location tags, this system is robust against eavesdroppers and spoofers outside the proximity range. The system improves the authentication accuracy by introducing a nonparametric Bayesian method called infinite Gaussian mixture model in the proximity test and provides flexible proximity range control by taking into account multiple physical-layer features of various ambient radio sources. Moreover, the session key establishment strategy significantly increases the key generation rate by exploiting the packet arrival time of the ambient signals. The authentication accuracy and key generation rate are evaluated via experiments using laptops in typical indoor environments. Liang Xiao 0003, Qiben Yan 0001, Wenjing Lou, Guiquan Chen, Y. Thomas Hou 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2012 | Indirect reciprocity game modelling for secure wireless networksabstractWe formulate the wireless security problem as an indirect reciprocity game, and propose a security mechanism that applies the indirect reciprocity principle to suppress attacks in wireless networks. In this system, a large number of nodes cooperate to reject the network access requests from attackers during the punishment periods. If the punishment time is so long that the cost due to the loss of network services exceeds the illegal security gains of the attack, rational nodes do not have incentive to attack, and hence our system can reduce the attacking probability in the network. We develop a social norm and reputation updating process to build such an indirect reciprocity mechanism for the network. We evaluate the evolutionarily stable strategy (ESS) in the game, and provide the optimal action strategy and its corresponding stationary reputation distribution. Our system is robust against collusion attacks, and can significantly reduce the attacking rate for a wide range of attacks. Simulation results show that our system has much better security performance than the direct reciprocity mechanism, especially in the large-scale wireless network with terminal mobility. Our system can be applied to many wireless networks including cognitive radio networks to improve their security performance. Liang Xiao 0003, Wan-Yi Sabrina Lin, Yan Chen 0007, K. J. Ray Liu |
ICC | 1 |
| 2012 | Indirect Reciprocity Security Game for Large-Scale Wireless NetworksabstractRadio nodes can obtain illegal security gains by performing attacks, and they are motivated to do so if the illegal gains are larger than the resulting costs. Most existing direct reciprocity-based works assume constant interaction among players, which does not always hold in large-scale networks. In this paper, we propose a security system that applies the indirect reciprocity principle to combat attacks in wireless networks. Because network access is highly desirable for most nodes, including potential attackers, our system punishes attackers by stopping their network services. With a properly designed social norm and reputation updating process, the aim is to incur a cost due to the loss of network access to exceed the illegal security gain. Thus rational nodes are motivated to abandon adversary behavior for their own interests. We derive the optimal strategy and the corresponding stationary reputation distribution, and evaluate the stability condition of the optimal strategy using the evolutionarily stable strategy concept. This security system is robust against collusion attacks and can significantly reduce the attacker population for a wide range of attacks when the stability condition is satisfied. Simulation results show that the proposed system significantly outperforms the existing direct reciprocity-based systems, especially in the large-scale networks with terminal mobility. This technique can be extended to many wireless networks, including cognitive radio networks, to improve their security performance. Liang Xiao 0003, Yan Chen 0007, Wan-Yi Sabrina Lin, K. J. Ray Liu |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2012 | Jamming-Resistant Collaborative Broadcast Using Uncoordinated Frequency HoppingabstractWe propose a jamming-resistant collaborative broadcast scheme for wireless networks, which utilizes the Un coordinated Frequency Hopping (UFH) technique to counteract jamming without preshared keys, and exploits node cooperation to achieve higher communication efficiency and stronger jamming resistance. In this scheme, nodes that already obtain the broadcast message serve as relays to help forward it to other nodes. Relying on the sheer number of relay nodes, our scheme provides a new angle for jamming countermeasure, which not only significantly enhances the performance of jamming-resistant broadcast, but can readily be combined with other existing or emerging antijamming approaches in various applications. We present the collaborative broadcast protocol, and analyze its successful packet reception rate and the corresponding cooperation gain for both synchronous and asynchronous relays for a snapshot scenario. We also investigate the full broadcast process based on a Markov chain model and derive a closed-form expression of the average broadcast delay. Simulation results in both single-hop and multihop networks indicate that our scheme is a promising antijamming technique in wireless networks. Liang Xiao 0003, Huaiyu Dai, Peng Ning |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2011 | Jamming-Resistant Collaborative Broadcast in Wireless Networks, Part II: Multihop NetworksabstractWe propose in [1] a collaborative broadcast scheme for wireless networks, which applies the Uncoordinated Frequency Hopping (UFH) technique to counteract jamming and exploits node cooperation to enhance broadcast efficiency. In this scheme, some nodes that already obtain the broadcast message are selected to relay the message to other nodes. In this paper, we extend the study to the generalized multihop network scenarios, and provide solutions for important related issues, such as the relay node selection, multiple access control, relay channel selection and packet scheduling.We also study the spatial and frequency (channel) diversity provided by the collaborative broadcast. Simulation results show that the collaborative broadcast achieves low broadcast delay, with low energy consumption and small computational overhead in multihop networks. Liang Xiao 0003, Huaiyu Dai, Peng Ning |
GLOBECOM | 1 |
| 2011 | Jamming-Resistant Collaborative Broadcast in Wireless Networks, Part I: Single-Hop NetworksabstractWe propose a collaborative broadcast scheme for wireless networks, which is based on the Uncoordinated Frequency Hopping (UFH) technique and exploits the node cooperation to achieve higher communication efficiency and stronger jamming resistance. In the collaborative broadcast, nodes that already obtain the broadcast message help forward the message to other nodes. Relying on the sheer number of relay nodes, which grows with time surely, our scheme is fundamentally more powerful than most recent attempts for anti-jamming broadcast. Potential applications include emergency alert broadcast and distribution of key system information in the presence of jamming. We provide three relay channel selection strategies for collaborative broadcast, analyze the corresponding successful packet reception rates for both synchronous and asynchronous scenarios, and present the corresponding cooperation gain. Simulation results in a practical setting show that our scheme significantly reduces broadcast delay and energy consumption against the most powerful jamming, - responsive-sweep jamming. Liang Xiao 0003, Huaiyu Dai, Peng Ning |
GLOBECOM | 1 |
| 2010 | PHY-Authentication Protocol for Spoofing Detection in Wireless NetworksabstractWe propose a PHY-authentication protocol to detect spoofing attacks in wireless networks, exploiting the rapid-decorrelation property of radio channels with distance. In this protocol, a PHY-authentication scheme that exploits channel estimations that already exist in most wireless systems, cooperates with any existing-either simple or advanced-higher-layer process, such as IEEE 802.11i. With little additional system overhead, our scheme reduces the workload of the higher-layer process, or provides some degree of spoofing protection for ``naked" wireless systems, such as some sensor networks. We describe the performance of our approach as a function of the spoofing pattern and the snapshot performance that can be easily measured through field tests. We discuss the implementation issues of the authentication protocol on 802.11 testbeds and verify its performance via field tests in a typical office building. Liang Xiao 0003, Alex Reznik, Wade Trappe, Chunxuan Ye, Yogendra Shah, Larry J. Greenstein, Narayan B. Mandayam |
GLOBECOM | 1 |
| 2009 | Channel-based detection of Sybil attacks in wireless networksabstractDue to the broadcast nature of the wireless medium, wireless networks are especially vulnerable to Sybil attacks, where a malicious node illegitimately claims a large number of identities and thus depletes system resources. We propose an enhanced physical-layer authentication scheme to detect Sybil attacks, exploiting the spatial variability of radio channels in environments with rich scattering, as is typical in indoor and urban environments. We build a hypothesis test to detect Sybil clients for both wideband and narrowband wireless systems, such as WiFi and WiMax systems. Based on the existing channel estimation mechanisms, our method can be easily implemented with low overhead, either independently or combined with other physical-layer security methods, e.g.,spoofingattack detection. The performance of our Sybil detector is verified, via both a propagation modeling software and field measurements using a vector network analyzer, for typical indoor environments. Our evaluation examines numerous combinations of system parameters, including bandwidth, signal power, number of channel estimates, number of total clients, number of Sybil clients, and number of access points. For instance, both the false alarm rate and the miss rate of Sybil attacks are usually below 0.01, with three tones, pilot power of 10 mW, and a system bandwidth of 20 MHz. Liang Xiao 0003, Larry J. Greenstein, Narayan B. Mandayam, Wade Trappe |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2009 | Channel-based spoofing detection in frequency-selective rayleigh channelsabstractThe radio channel response decorrelates rapidly as the transmitter changes location in an environment with rich scatterers and reflectors. Based on this fact, a channel-based authentication scheme was previously proposed to discriminate between transmitters at different locations, and thus to detect spoofing attacks in wireless networks. In this paper, we study its application in frequency-selective Rayleigh channels, considering channel time variations due to environmental changes and terminal mobility, as well as the channel estimation errors due to the interference from other radios. We propose a generalized likelihood ratio test (GLRT) that is optimal but computationally cumbersome, and a simplified version that requires no a priori knowledge of channel parameters and is therefore more practical. We verify the efficacy of the channel-based spoofing detectors via numerical analysis, showing how performance is improved by using multiple antennas, higher transmit power, and wider system bandwidth. We show that, under a wide variety of practical conditions, spoofing can be detected with better than 90% probability while keeping the probability of falsely rejecting valid transmissions below 10%. Liang Xiao 0003, Larry J. Greenstein, Narayan B. Mandayam, Wade Trappe |
IEEE Trans. Wirel. Commun. | 1 |
| 2008 | A Physical-Layer Technique to Enhance Authentication for Mobile TerminalsabstractWe propose an enhanced physical-layer authentication scheme for multi-carrier wireless systems, where transmission bursts consist of multiple frames. More specifically, it is based on the spatial variability characteristic of wireless channels, and able to work with moderate terminal mobility. For the authentication of the first frame in each data burst, the legal transmitter uses the saved channel response from the previous burst as the key for authentication of the first frame in the next burst. The key is obtained either via feedback from the receiver, or using the symmetric channel property of a TDD system. Then the authentication of the following frames in the burst is performed either by a Neyman-Pearson hypothesis test, or a least-squares adaptive channel estimator. Simulations in a typical indoor building show that the scheme based on the Neyman-Pearson test is more robust against terminal mobility, and is able to detect spoofing attacks efficiently with small system overhead when the terminal moves with a typical pedestrian speed. Liang Xiao 0003, Larry J. Greenstein, Narayan B. Mandayam, Wade Trappe |
ICC | 1 |
| 2008 | Distributed measurements for estimating and updating cellular system performanceabstractWe investigate the use of distributed measurements for estimating and updating the performance of a cellular system. Specifically, we discuss the number and placement of sensors in a given cell for estimating its signal coverage. Here, an "outage" is said to occur at a location if a mobile receiver there has inadequate signal-to-noise ratio (SNR -based outage) or, using another criterion, inadequate signal-to-interference ratio (SIR- based outage); and the "outage probability" is the fraction of the cell area over which outage occurs. A design goal is to improve measurement efficiency (i.e., minimizing the required number of measurement sensors) while accurately estimating the outage probability and mapping the coverage holes. The investigation uses a generic path loss model incorporating distance effects and spatially correlated shadow fading. Our emphasis is on the performance prediction accuracy of the sensor network, rather than on cellular system analysis per se. Through analysis and simulation, we assess several approaches to estimating the outage probability. Applying the principle of importance sampling to the sensor placement, we show that a cell outage probability of Pocan be accurately estimated using ~ 10/Popower-measuring sensors distributed in a random uniform way over the area with base-sensor distances from 50% to 100% of the cell radius. This result applies to both SNR-based and SIR-based outage estimation for both indoor and outdoor environments. Liang Xiao 0003, Larry J. Greenstein, Narayan B. Mandayam, Shalini S. Periyalwar |
IEEE Trans. Commun. | 1 |
| 2008 | Using the physical layer for wireless authentication in time-variant channelsabstractThe wireless medium contains domain-specific information that can be used to complement and enhance traditional security mechanisms. In this paper we propose ways to exploit the spatial variability of the radio channel response in a rich scattering environment, as is typical of indoor environments. Specifically, we describe a physical-layer authentication algorithm that utilizes channel probing and hypothesis testing to determine whether current and prior communication attempts are made by the same transmit terminal. In this way, legitimate users can be reliably authenticated and false users can be reliably detected. We analyze the ability of a receiver to discriminate between transmitters (users) according to their channel frequency responses. This work is based on a generalized channel response with both spatial and temporal variability, and considers correlations among the time, frequency and spatial domains. Simulation results, using the ray-tracing tool WiSE to generate the time-averaged response, verify the efficacy of the approach under realistic channel conditions, as well as its capability to work under unknown channel variations. Liang Xiao 0003, Larry J. Greenstein, Narayan B. Mandayam, Wade Trappe |
IEEE Trans. Wirel. Commun. | 1 |
| 2007 | Fingerprints in the Ether: Using the Physical Layer for Wireless AuthenticationabstractThe wireless medium contains domain-specific information that can be used to complement and enhance traditional security mechanisms. In this paper we propose ways to exploit the fact that, in a typically rich scattering environment, the radio channel response decorrelates quite rapidly in space. Specifically, we describe a physical-layer algorithm that combines channel probing (M complex frequency response samples over a bandwidth W) with hypothesis testing to determine whether current and prior communication attempts are made by the same user (same channel response). In this way, legitimate users can be reliably authenticated and false users can be reliably detected. To evaluate the feasibility of our algorithm, we simulate spatially variable channel responses in real environments using the WiSE ray-tracing tool; and we analyze the ability of a receiver to discriminate between transmitters (users) based on their channel frequency responses in a given office environment. For several rooms in the extremities of the building we considered, we have confirmed the efficacy of our approach under static channel conditions. For example, measuring five frequency response samples over a bandwidth of 100 MHz and using a transmit power of 100 mW, valid users can be verified with 99% confidence while rejecting false users with greater than 95% confidence. Liang Xiao 0003, Larry J. Greenstein, Narayan B. Mandayam, Wade Trappe |
ICC | 1 |
| 2007 | Sensor-assisted localization in cellular systemsabstractWe investigate the use of an auxiliary network of sensors to locate mobiles in a cellular system, based on the received signal strength at the sensor receivers from a mobile's transmission. The investigation uses a generic path loss model incorporating distance effects and spatially correlated shadow fading. We describe four simple localization schemes and show that they all meet E-911 requirements in most environments. Performance can be further improved by implementing the MMSE algorithm, which ideally reaches the Cramer-Rao bound. We compare the MMSE algorithm and the four simple schemes when the model parameters are estimated via inter-sensor measurements. Liang Xiao 0003, Larry J. Greenstein, Narayan B. Mandayam |
IEEE Trans. Wirel. Commun. | 1 |
| 2006 | Sensor Networks for Estimating and Updating the Performance of Cellular SystemsabstractWe investigate the use of an auxiliary network of sensors to assist radio resource management in a cellular system. Specifically, we discuss the number and placement of sensors in a given cell for estimating its signal coverage. Here, an "outage" is said to occur at a location if the mobile receiver there has inadequate signal-to-noise ratio (SNR-based outage) or, using another criterion, inadequate signal-to-interference ratio (SIR-based outage); and the "outage probability" is the fraction of the cell area over which outage occurs. A design goal is to confine the number of sensors per cell to an acceptable level while accurately estimating the outage probability. The investigation uses a generic path loss model incorporating distance effects and spatially correlated shadow fading. Our emphasis is the performance prediction accuracy of the sensor network, rather than cellular system analysis per se. Through analysis and simulation, we assess several approaches to estimating the outage probability. Applying the principle of importance sampling to the sensor placement, we show that a cell outage probability of ~ Po can be accurately estimated using ~ 10/Po power-measuring sensors distributed in a random uniform way over base-mobile distances from 50% to 100% of the cell radius. This result applies to both SNR-based and SIR-based cases, in both indoor and outdoor environments. Liang Xiao 0003, Larry J. Greenstein, Narayan B. Mandayam, Shalini S. Periyalwar |
ICC | 1 |
| 2003 | QoS-oriented scheduling algorithm for mobile multimedia in OFDMabstractIn this paper, we propose and evaluate a new algorithm to schedule multimedia data in the multi-user OFDM wireless system. This algorithm allocates the resource dynamically in both time domain and frequency domain, based on the feedback channel state, packet QoS requirement and the amount of the data to be transmitted. It tries to guarantees the maximum delay, BER requirement and priority of the packet and improve system throughput, even with worse channel condition. Several neighboring subcarriers in OFDIM are allocated together and by utilizing adaptive coding and modulation, they are transmitted with multiple data rates based on the channel state information. We study its performance on the channel model with Doppler fading, multi-path fading and path-loss. Simulation results show it is efficient to guarantee QoS for high rate multimedia data in OFDM system. Liang Xiao 0003, Yan Yao 0002 |
PIMRC | 1 |