VLDB 2026 Research / reviewers in the wild / expert
Zefang Lv
dblp:236/4524
· DBLP profile ↗
28ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0003-3880-5069ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 24 · 5 first-author · 23 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RLCVP: Collaborative Vehicular Perception Against Data Fabrication Attacks
Zhiping Lin 0002, Liang Xiao 0003, Zefang Lv, Chen Chen 0037, Liqing Ye |
WCNC | 4 |
| 2026 | Reinforcement Learning Based Model Adaptation and Partition for Dynamic DNN Empowered Mobile VisionabstractMost mobile vision systems adopt background subtraction or task offloading, but inference performance remains constrained by background dynamics. Dynamic deep neural networks (DNNs) scale computation without relying on historical information. However, a dynamic DNN tuned to a specific spatial granularity cannot excel across all images. For example, resolution level models perform well on images with sparse and large objects but poorly on images with dense and small objects. Changing system overheads make the choice of an optimal edge-assisted scheme for dynamic DNNs difficult. In this paper, we propose an efficient reinforcement-learning-based model adaptation and partition scheme for mobile vision empowered by dynamic DNNs, which jointly adapts the backbone network and the partition point of the inference model to accelerate inference without degrading inference accuracy. The scheme selects hyperparameters for dynamic DNNs, such as spatial granularity and complexity level, based on image characteristics and measured performance without relying on previous frames. A model partition algorithm is proposed to split the changing inference model in polynomial complexity using the Ford–Fulkerson algorithm. Experimental results based on object detection and image classification show that the proposed scheme reduces inference latency and energy consumption by over 71.2% and 31.8%, without degrading inference accuracy compared to the benchmarks. Liang Xiao 0003, Tuhao Li, Zefang Lv, Ziyue Qiao, Hui Xiong 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2026 | Reinforcement Learning-Based Edge-Assisted Inference With Multimodal DataabstractDeep neural networks (DNNs) extract embeddings from multimodal data such as audio, images, and LiDAR, each with heterogeneous computational demands and representation abilities to support multimodal services such as audio-visual speech recognition. Reinforcement learning (RL)-based model selection and splitting schemes determine the model variants from DNNs and the components to offload in unimodal models to reduce inference latency, aiming for a three-fold trade-off among inference accuracy, computation cost, and communication overhead but ignore the heterogeneity of different modalities within multimodal DNNs. In this paper, we propose an RL-based edge-assisted multimodal inference scheme that optimizes model selection at modality level and edge-assisted policies, including collaborative servers and partition points for each feature extractor to perform multimodal DNNs on mobile devices. Based on the information complexity and historical influence of each modality, as well as real-time observations such as channel gain, and previous inference performance, the policy distributions are designed to maximize the utility, as a weighted sum of inference latency and energy consumption, and inference accuracy. Safe policy exploration further mitigates risks such as low inference accuracy, intolerable inference latency, and improper allocation of computational resources to specific modalities. We analyze the computational complexity affected by the number of model variants, edge servers, and partition points and derive performance bounds for inference latency, energy consumption, and utility under specific sample sizes and data rates. Experimental results show that the proposed schemes improve inference performance compared to benchmark schemes. Liang Xiao 0003, Chuxuan Wang, Zefang Lv, Yiwen Zhan 0002, Yilin Xiao 0001, Helin Yang |
IEEE Trans. Mob. Comput. | 4 |
| 2026 | RL-Based Anti-Jamming Maritime Communications for LLM InferenceabstractReinforcement learning (RL)-based anti-jamming communication schemes select the transmit power and channel to send images or point clouds, but have low transmission reliability due to lack of maritime environmental features and the delayed feedback under harsh channel conditions caused by sea surface reflections and wave fluctuations. In this paper, we propose an RL-based anti-jamming maritime communication scheme for LLM inference that enables user equipment (UE) to optimize uplink power, channel and data compression ratio for transmitting multi-modal data based on data size of each modality, number of received packets, environmental features such as weather conditions and jamming features such as the received jamming power with quality-of-service guarantee. The historical anti-jamming experiences are used to recover the delayed feedback or loss and reduce communication outage under harsh maritime channel. The upper bound in terms of latency, UE energy consumption and inference accuracy is provided based on a maritime anti-jamming game to show the impact of the number of modalities, bandwidth and channel gains. The proposed scheme is implemented based on UEs equipped with Jetson Orin, camera and BME280 sensors for transmitting 200-KB temperature, humidity and pressure as well as images to the control center for LLaVA-based environmental feature extraction and marine object detection. Experimental results based on 4 UEs at Harbor show the performance gain with 40.5% less latency, 46.4% lower UE energy consumption and 10.9% higher inference accuracy against the smart jammer. Liang Xiao 0003, Pengli Zhang, Haoyu Chen 0005, Zefang Lv, Manhao Jiang |
IEEE Trans. Wirel. Commun. | 6 |
| 2026 | Learning-Based Anti-Jamming Energy-Efficient Wide-Area CommunicationsabstractReinforcement learning (RL)-based low-power wide-area networks (LPWANs) such as Long Range (LoRa) enable the access point (AP) to optimize the uplink power and channel of user equipment (UE), but the energy efficiency and quality of service (QoS) are low due to frequent downlink feedback against smart jamming. In this paper, we propose an RL-based anti-jamming energy-efficient wide-area communication scheme to optimize the transmit power, channel and downlink feedback policy with QoS guarantee. The similarities of channel states and UE communication tasks are evaluated to update learning parameters and accelerate policy selection under low duty cycles in LPWANs. We further propose a deep RL-based scheme that exploits neural networks to address the quantization error of the received power at AP and UE battery level and update weights according to prioritized experience replay to improve learning efficiency. The computational complexity and the performance bound in terms of signal-to-noise ratio and UE energy consumption are provided based on the Stackelberg equilibrium of the anti-jamming game. Our proposed schemes are implemented based on Raspberry Pi and BME280 sensors to support the LoRa-based transmission of temperature, humidity and pressure data. Experimental results on campus show the performance gain with less packet loss rate and UE energy consumption over the benchmark against smart jamming. Liang Xiao 0003, Shuohua Wang, Zefang Lv, Yiwen Zhan 0002, Haoyu Chen 0005 |
IEEE Trans. Wirel. Commun. | 5 |
| 2025 | Reinforcement Learning Based Edge-Assisted Dynamic Inference for Mobile VisionabstractMobile device can deploy high-complexity vision tasks by reducing spatial redundancy in the inference of deep neural networks (DNNs) and offloading computations to remote servers via wireless channels. However, a trained DNN model with fixed network architecture faces challenges in dynamic scenarios with less overlap between previous and current frames. In this paper, we propose a RL-based edge-assisted mobile vision scheme based on dynamic DNN models to balance inference latency and capacity without relying on historical information for dynamic visual scenarios. This scheme optimizes the hyperparameter settings and complexity level of dynamic neural networks as well as the collaborative server and the partition points of the DAG structured inference model. Based on the image embeddings, channel gain, and previous inference performance, the inference policy is selected to optimize the utility function, which is a weighted sum of inference latency, local energy consumption, and inference accuracy. Experimental results demonstrate the performance improvements of our proposed scheme over the benchmarks. Liang Xiao 0003, Tuhao Li, Zefang Lv, Ziyue Qiao, Hui Xiong 0001 |
GLOBECOM | 4 |
| 2025 | Reinforcement Learning Based UAV Swarm Enabled 3-D Multimodal Jamming DetectionabstractReinforcement learning based jamming detection that chooses the test threshold to evaluate the received signal strength indicator (RSSI) and the packet loss rate is inaccurate for unmanned aerial vehicles (UAVs) lacking target indication against smart jammers. In this paper, we propose a UAV swarm-enabled multimodal jamming detection scheme to optimize the test thresholds based on vision information such as object classes and relative distances, along with the RSSI, the channel gain and the communication performance, including packet delivery ratio, bit error rate and packet delivery delay. A machine learning classifier is used to assess the joint variation among RSSI, channel gain and communication performance, with the output outlier score compared with the test threshold to detect the jammer. The detection results shared from neighboring UAVs are exploited in the update of policy distribution to refine jamming signal resolution and thus reduce the miss detection rate. The utility bound is derived based on the Nash equilibrium of the jamming detection game between the UAV swarm and the jammer. Experimental results based on 5 UAVs to detect a smart jammer show that our proposed scheme enhances the detection accuracy compared with the benchmarks. Qiaoxin Chen, Liang Xiao 0003, Jieling Li, Yuxiao Ren, Zefang Lv, Hongbin Jin |
ICC | 6 |
| 2025 | Reinforcement Learning Based Anti-Jamming FANET Routing with QoS GuaranteeabstractReinforcement learning (RL) based flying ad-hoc network (FANET) routing enables unmanned aerial vehicles (UAVs) to choose the next-hop, but the quality of service (QoS) and energy efficiency have to be enhanced against jamming due to the inaccurate path quality estimation. In this paper, we propose an RL based anti-jamming FANET routing with QoS guarantee to optimize the transmit power and the originator message broadcast interval for the route discovery to estimate the path quality to select the next-hop from the routing table. Based on the path availability history, the channel conditions, the received jamming power and the transmission quality regarding the number of the received originator messages and the bit error rate during route discovery, the routing policy is selected to enhance the throughput and the energy efficiency. The backward estimation of the throughput in the utility evaluation addresses the delayed feedback from the destination. The performance bound is derived in terms of network topology and channel gain based on the Nash equilibrium of the cooperative game among the UAVs. Simulation results provide the performance gain of the throughput and energy consumption over the benchmarks. Jieling Li, Chuxuan Wang, Liang Xiao 0003, Zefang Lv, Pengli Zhang, Helin Yang |
ICC | 4 |
| 2025 | RL-based Anti-Jamming Maritime Communications for Multi-Modal PerceptionabstractReinforcement learning (RL)-based communication scheme selects the transmit power and channel to improve communication reliability, but energy consumption and transmission delay degrades due to multi-modal data transmission such as images and point clouds for object detection under harsh maritime channel against jamming. In this paper, we propose a RL-based anti-jamming maritime communication scheme for multi-modal perception that optimizes the uplink transmit power, channel and compression ratios of each modality to improve perception accuracy and reduce communication energy consumption and transmission delay. The transmission policy of multi-modal data is determined by the probability distribution based on the data size and maritime environmental features such as rainfall and wind speed. The upper bound in terms of transmission delay and communication energy consumption is provided based on the Nash equilibrium of anti-jamming perception game to show the impact of the number of modalities and channel gain. Simulation results for ship detection based on the images and radar point clouds show the performance gain over benchmark against the smart jammer. Liang Xiao 0003, Pengli Zhang, Haoyu Chen 0005, Zefang Lv |
VTC2025-Fall | 6 |
| 2025 | Reinforcement-Learning-Based APT Defense for Large-Scale Smart GridsabstractReinforcement-learning-(RL)-based advanced persistent threats (APTs) defense schemes choose the scan interval to enhance the detection accuracy, and the data protection level can be further improved by optimizing the repair rate corresponding to methods, such as malicious files removal, security software version update and passwords reset, each requiring different computational resources, such as CPUs. In this article, we propose an RL-based APT defense scheme to optimize both the continuous scan interval of metering data and repair rate to mitigate the potential loss for meter data management systems (MDMSs) in smart grids with large-scale meters. Based on the size of the metering data stored at MDMS and the compromised data, the data tag granularity and the number of CPUs for repairing, neural networks extract the state feature, address the quantization error of defense policy, and update the weights based on the APT defense experiences and the shared weights of neighboring MDMSs to improve the defense performance as a weighted sum of the defense duration, detection accuracy and data protection level. The computational complexity of the proposed scheme and the performance bounds according to the Nash Equilibrium of the game between the MDMS and the APT attacker are provided. Simulation results based on three MDMSs that receive the metering data from 50–150 smart meters and an APT attacker with the selected attack interval up to 5 s show the performance gain over the benchmark based on the optimal control theory and Q-learning in large-scale smart grids. Liang Xiao 0003, Zefang Lv, Zhiping Lin 0002, Yousong Du |
IEEE Internet Things J. | 3 |
| 2025 | Reinforcement Learning-Based Accurate Worm Detection for Smart GridsabstractReinforcement learning (RL) based worm detection chooses the test threshold to evaluate the network traffic features such as the spectral flatness measure (SFM), but the detection of the evasive worm that modifies the scan rate and the worm propagation speed to manipulate the network traffic features is inaccurate due to the estimation and quantization error in the test threshold. In this paper, we propose an RL based accurate worm detection for smart grids that enables the control center to optimize the test threshold based on the number of meters, the infection time series and the number of connections to new destination IP addresses, besides the traffic log size received from each data concentrator and the number of the previously infected meters. A constraint on the maximum missed detection rate required by the smart grids is exploited in the detection policy distribution to support the reliable data transmission. A deep RL version addresses the quantization error in terms of the test threshold in the SFM evaluation and the infection time series in the state formulation, and compress the state space of the traffic log size received from a large number of data concentrators. Based on a worm detection game, the performance bound is provided under the specified detection window size and the propagation speed of evasive worm. Simulation results for 3600 meters show that the performance gain of the detection accuracy and latency against evasive worm over the benchmarks. Liang Xiao 0003, Jieling Li, Yilin Xiao 0001, Zefang Lv, Chuxuan Wang, Pengmin Li |
IEEE Internet Things J. | 4 |
| 2025 | Learning-Based Energy-Efficient Anti-Jamming FANET Routing With QoS GuaranteeabstractReinforcement learning (RL) based flying ad-hoc network (FANET) routing enables unmanned aerial vehicles (UAVs) to choose the next-hop to forward the packets, but the quality of service (QoS) and energy efficiency have to be enhanced due to the inaccurate path quality estimation under jamming attacks. In this paper, we propose an RL based energy-efficient anti-jamming FANET routing scheme with QoS guarantee to optimize both the transmit power and the originator message broadcast interval for the route discovery to estimate the path quality to select the next-hop from the routing table. Based on the path availability history, the channel conditions and the received jamming power, as well as the transmission quality regarding the number of the received originator messages and the bit error rate during route discovery, the routing policy is selected to enhance the throughput and the energy efficiency under jamming attacks with changing power. The backward estimation of the throughput in the utility evaluation addresses the delayed feedback from the destination under large-scale networks. The deep neural networks are further designed to address the quantization error of the transmission quality and the channel gain for UAVs with high mobility to enhance the path exploration efficiency. In addition, the upper bound in terms of network topology and channel gain is derived based on the Nash equilibrium of the anti-jamming routing game. The proposed routing scheme is implemented to improve the image transmission quality against jamming in outdoor environments. Experimental results based on UAVs equipped with Raspberry Pi show the performance gain of the throughput and the energy consumption. Jieling Li, Liang Xiao 0003, Chuxuan Wang, Zefang Lv, Pengli Zhang, Helin Yang |
IEEE Trans. Commun. | 4 |
| 2025 | Reinforcement Learning-Based False Data Injection Attacks in Smart GridsabstractFalse data injection (FDI) attacks construct attack vectors to inject false data into tampered meters with the goal of falsifying state estimation, but resulting in low successful attack rate with high attack costs in terms of the number of tampered meters in large-scale smart grids, because the bad data detection at the control center chooses the dynamic detection thresholds to identify the modified meter measurements. In this article, we propose a reinforcement learning-based FDI attack scheme that optimizes both the tampered meters and the false data to enhance the success attack rate and injected errors while reducing attack costs. Based on meter measurements and previous performance, the attack vector is constructed to induce more errors in state estimation and bypass bad data detection. The performance bounds regarding the successful attack rate and the injected error are derived in terms of the number of bus phase angles, the susceptance of the transmission line, and the maximum false data based on the Nash equilibrium of the FDI game. Simulations performed on both the IEEE 14-bus and IEEE 118-bus systems demonstrate the performance gain over the benchmarks. Liang Xiao 0003, Haoyu Chen 0005, Zefang Lv, Chuxuan Wang, Yilin Xiao 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | Collaborative Perception Against Data Fabrication Attacks in Vehicular NetworksabstractCollaborative perception in vehicular networks enables the connected autonomous vehicle (CAV) to gather sensing data, such as feature maps of light detection and ranging (LiDAR) point clouds, from neighboring CAVs to achieve higher perception accuracy, which has performance degradation against data fabrication attacks that share falsified sensing data with random probability. In this paper, we exploit the spatial consistency check to detect the potentially manipulated regions in LiDAR point clouds and measure the inconsistency degree of the received sensing data based on the number of conflict regions, which is the basis for determining the falsified sensing data if the inconsistency degree exceeds the threshold of the hypothesis test. The reinforcement learning (RL)-based collaborative vehicular perception scheme against data fabrication attacks is further proposed to choose CAVs based on the inconsistency degrees, the data quality measured by the confidence scores, the channel gains and the CAV reputations, which enhances the utility as the weighted sum of perception accuracy, speed and minimum latency requirement for data transmission. In addition, the multi-layer perceptron-based neural networks extract the perception features of sensing data from historical experiences, such as the data quality of received feature maps, as well as compress the RL state that linearly increases with the network scales and the spatial granularity of LiDAR point clouds for faster learning. Experimental results based on 10 CAVs equipped with LiDAR sensors and NVIDIA computational units to detect 20 vehicles against data fabrication attacks show that our proposed scheme outperforms the benchmarks in terms of perception accuracy and speed. Zhiping Lin 0002, Liang Xiao 0003, Zefang Lv |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Edge-Assisted Collaborative Perception Against Jamming and Interference in Vehicular NetworksabstractCollaborative perception of connected autonomous vehicles (CAVs) that offload the sensing data, such as the feature map extracted from light detection and ranging (LiDAR) point clouds, to an edge device such as the roadside unit (RSU) to detect traffic objects has severe performance degradation due to the offloading latency and packet loss rate (PLR) under jamming and interference. In this paper, we propose an edge-assisted reinforcement learning (RL)-based collaborative perception scheme for CAVs to enhance the accuracy and speed against jamming and interference in LiDAR-based object detection. Based on the spatial confidence score of the feature map, the data size, the channel gains, the received jamming power and interference level, this scheme chooses the critical regions of the feature map, radio channel and transmit power with the hierarchical structure to enhance the learning efficiency. The risk level of the selected policy evaluates the time asynchronization and information loss of the shared feature map using the multi-level risk function based on multiple thresholds of the offloading latency and PLR, with assigning different penalties to mitigate the selection of high-risk policies that degrade perception performance. The upper performance bound in terms of the perception accuracy, latency and utility is provided based on the Stackelberg equilibrium of the game between the jammer and CAVs. Experimental results based on the Robosense RS-LiDAR-16 sensors and the Raspberry Pi to detect 10 vehicles in an$8.5\times 4\times 3.5$m3area show the performance gain with 22.4% higher perception accuracy and 41.3% less latency compared with the benchmark against a smart jammer. Zhiping Lin 0002, Liang Xiao 0003, Zefang Lv, Yunjun Zhu, Yanyong Zhang, Yong-Jin Liu 0001 |
IEEE Trans. Wirel. Commun. | 4 |
| 2024 | Reinforcement Learning Based Energy-Efficient Anti-Jamming NB-IoT CommunicationsabstractNarrowband Internet of Things (NB-IoTs) with limited power supply and long range requirements can apply reinforcement learning (RL) based anti-jamming techniques to choose the transmit power and channel to enhance the communication reliability, but have low energy efficiency for the energy-constrained user equipment (UE). In this paper, we propose a RL-based energy-efficient anti-jamming NB-IoT communication scheme to choose the transmission policy including the UE transmit channel and power as well as whether to send them to UE over the downlink channel. Based on the received power on each channel, UE battery level and message size, the transmission policy distribution is formulated according to the risk value that indicates communication reliability exceeding the narrowband bit error rate (BER) requirement and the long-term utility as the weighted sum of UE energy consumption, BER level and frequency hopping overhead for higher energy efficiency and reliability. The similarity of the state-action pairs is evaluated using the risk value to update the long-term risk levels of feasible policies via the knowledge reuse technique for faster learning. The computational complexity and the upper bounds of our proposed scheme in terms of BER and energy consumption are provided. Simulation results show the performance gain over the benchmark against jamming. Shuohua Wang, Zefang Lv, Jieling Li, Liang Xiao 0003 |
GLOBECOM | 3 |
| 2024 | Reinforcement Learning Based QoS-Aware Anti-Jamming Underwater Video TransmissionabstractUnderwater video transmission has to ensure quality-of-service (QoS) against jamming with severe multipath effect and narrow bandwidth limitation that degrade the communication performance under variable channel state. In this paper, we propose a reinforcement learning (RL)-based QoS-aware underwater video transmission scheme to optimize the video compression ratio, modulation format and transmit power based on the state consisting of the channel gain and previous transmission performance. This scheme evaluates the risk level that indicates the probability of failing the QoS and the long-term expected utility of each transmission policy under the current state to improve the anti-jamming communication performance. We derive the performance bound of the utility and analyze its relationship with transmission policy. Simulation results illustrate that our scheme improves the QoS by reducing the frame loss rate (FLR), transmission delay and increasing spatial-spectral entropy-based quality (SSEQ) compared with the benchmark. Shaoxuan Li, Zefang Lv, Liang Xiao 0003, Wei Su 0002 |
WCNC | 4 |
| 2024 | Reinforcement Learning Based Energy-Efficient Fast Routing for FANETsabstractReinforcement learning (RL) based flying ad-hoc network (FANET) routing enables unmanned aerial vehicles (UAVs) to choose the next-hop to increase the packet delivery ratio, but the routing latency and energy consumption have to be further reduced over inaccurate feedback for large-scale networks. In this paper, we propose an RL based energy-efficient fast routing for each UAV to choose the forwarding decision and the power. Based on the state consisting of the battery level, channel conditions and forwarding decisions of the one-hop neighbors, the routing policy is chosen to enhance the utility as the weighted sum of the delivery success indicator, the latency and the energy consumption. The number of the latency violations and the learning parameters shared among the one-hop neighbors are exploited in the update of the routing policy distribution following the latency constraint with the reduced energy consumption. The deep neural networks address the state quantization error of the latency and the channel gain for UAVs with high mobility under large-scale networks. The performance bound regarding the end-to-end latency and the energy consumption is derived in terms of network topology and channel gain based on the packet forwarding game. The performance gain over the benchmark is provided via both simulation and experimental results. Jieling Li, Liang Xiao 0003, Xuchen Qi, Zefang Lv, Qiaoxin Chen, Yong-Jin Liu 0001 |
IEEE Trans. Commun. | 4 |
| 2024 | Safe Multi-Agent Reinforcement Learning for Wireless Applications Against Adversarial CommunicationsabstractBased on the network observations and learning parameters shared by the neighboring learning agents, multi-agent reinforcement learning (RL) has to enhance the performance over adversarial communications, in which spoofing attackers send fake learning messages to fool the learning agent and thus degrade the performance of wireless applications. In this paper, we propose a safe multi-agent RL algorithm for wireless applications against adversarial communications, in which each learning agent chooses the cooperative agents to share the learning information and authenticates the received learning messages before integrating them into the RL state formulation and the learning parameter update. The communication policy distribution for the cooperative agent selection is formulated based on the long-term discounted reward and the sharing reputation for each neighboring agent, which is updated based on the authentication results to indicate the probability as a spoofing attacker. Neural networks are designed to estimate the long-term discounted reward and the sharing reputation for the learning agent with sufficient computational resources in large-scale wireless networks to enhance the agent selection security. As a case study, our proposed algorithm is implemented in the unmanned aerial vehicle swarm anti-jamming video transmission against spoofing attackers that send fake received jamming power as well as Q-values and neural network weights in the anti-jamming transmission policy learning. Both simulation and experimental results are provided to verify the performance gain over the benchmark. Zefang Lv, Liang Xiao 0003, Haoyu Chen 0005, Xiangyang Ji |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Efficient Communications in Multi-Agent Reinforcement Learning for Mobile ApplicationsabstractThe environment observations and learning experiences shared by the cooperative learning agents accelerate multi-agent reinforcement learning (MARL) with partial observations for mobile applications but the performance degrades due to the redundant and outdated observations under severe channel fading in wireless networks. In this paper, we propose an efficient communication scheme in MARL for mobile applications that enables each learning agent to optimize the cooperative agents and the learning parameters to integrate the shared information. The cooperative agents are chosen according to the learning environment observations, the channel states, and the task similarity with neighboring agents. The learning parameters are chosen based on the attention mechanism that exploits the correlation with the local observation to enhance the agent receptive field for efficient policy exploration. Neural networks with weights updated based on the learning factors determined by the task similarity are designed to further improve the learning efficiency. The performance bounds including the information gain from the learning agent cooperation, the communication cost and the utility are provided based on the Nash equilibrium of the cooperative MARL communication game. The proposed scheme is implemented in the anti-jamming video transmission of the unmanned aerial vehicle swarms to optimize the transmit channel and power and experimental results verify the performance gain over the benchmark. Zefang Lv, Liang Xiao 0003, Yousong Du, Yunjun Zhu, Shuai Han 0002, Yong-Jin Liu 0001 |
IEEE Trans. Wirel. Commun. | 1 |
| 2023 | Multi-Agent Reinforcement Learning for Wireless Networks Against Adversarial CommunicationsabstractBased on the efficient and reliable exchange of learning messages containing both the policy selection experiences such as the learning parameters and observations among the learning agents, multi -agent reinforcement learning (RL) has to address adversarial communications that send fake learning messages to learning agents with the goal of decreasing the RL rewards or even failing the learning tasks. In this paper, we propose a multi-agent RL (MARL) communication framework for wireless networks against adversarial communications, in which each learning agent chooses the cooperative agents to share learning messages based on the agent reputation that indicates the probability to send fake learning messages. By comparing with the learning history, each learning agent authenticates the received learning messages before integrating them in the RL task state formulation and the learning parameter update for robust task learning. The learning factor that increases with the correlation between the local and the shared observation is calculated to update the task Q-values and neural network weights based on the shared learning parameters. As a case study, the multi-agent deep Q-network within our proposed MARL communication framework is implemented in the UAV swarm video transmission system and the performance gain over the benchmark is provided in the simulation results based on 5-UAV swarm against an attacker that sends fake observations and neural network weights. Zefang Lv, Liang Xiao 0003, Helin Yang, Xiangyang Ji |
GLOBECOM | 1 |
| 2023 | Efficient Communications for Multi-Agent Reinforcement Learning in Wireless NetworksabstractMulti-agent reinforcement learning (RL) utilizes the observations and learning experiences shared among the agents to accelerate learning speed under partial observations and the resulting learning efficiency depends on the cooperative agent selection and the RL task state formulation. In this paper, we propose an efficient communication scheme for multi-agent RL that enables each learning agent to optimize the cooperative agent selection and the task state formulation to improve the learning performance and the quality of service for RL-based applications in wireless networks. Based on the local observation, the radio channel states, the similarity of RL task with neighboring agents and previous communication cost, this scheme formulates a communication state, which is input to a neural network to estimate the communication policy distribution. The RL task state of the learning agent, which consists of the local observation such as channel states and previous task performance, as well as the correlation between the shared and the local observation extracted based on the attention mechanism, is formulated to enhance the agent receptive field. In addition, the shared learning information is also exploited to update the local learning parameters such as the task Q-values and neural network weights and further improve the RL task policy exploration. As a case study, the proposed communication scheme is implemented in the multi-agent deep Q-network based anti-jamming unmanned aerial vehicle swarm communications and the performance gain over the benchmark is verified via simulation results. Zefang Lv, Yousong Du, Liang Xiao 0003, Shuai Han 0002, Xiangyang Ji |
GLOBECOM | 1 |
| 2023 | Reinforcement Learning Based Energy-Efficient Routing with Latency Constraints for FANETsabstractReinforcement learning (RL) enables flying ad-hoc networks (FANETs) to choose the next hop unmanned aerial vehicles (UAV s) with shorter routing path, but may raise the retransmission rate and fails to guarantee the quality of service (QoS) under the high mobility and fast fading channels. In this paper, we propose an RL based routing scheme that optimizes both the routing and the power allocation to protect the latency QoS and save routing energy consumption of the FANET. Based on the routing history, the channel conditions, the battery level and the shared knowledge from the neighbors, this scheme formulates the routing policy distribution with safe exploration to select the stable path and thus reduce the retransmission rate. Specifically, the risk value with respect to end-to-end latency constraint is designed to evaluate the routing policy and reduce the exploration probability of the high-latency routing. Based on the distributed value function approach, the learning parameter such as the state value functions shared among neighbors is exploited to accelerate the routing process and enhance the routing stability under the dynamic network topology. Simulation results verify the routing performance gain of our proposed scheme over the benchmark. Xuchen Qi, Jieling Li, Zefang Lv, Liang Xiao 0003 |
GLOBECOM | 3 |
| 2023 | Reinforcement Learning Based UAV Swarm Communications Against JammingabstractReinforcement learning based unmanned aerial vehicle (UAV) swarm communications have to address the challenges raised by the large-scale dynamic network and strong jamming and interference. In this paper, we propose a multiagent reinforcement learning based UAV swarm anti-jamming communication scheme to optimize the UAV relay selection and power allocation based on the network topology, channel states, previous performance and the network states shared by neighboring UAVs. This scheme formulates the policy distribution to improve the policy space exploration and designs a soft learning mechanism to guide the policy update and stabilize the learning process. According to transfer learning, the shared swarm experiences are exploited to accelerate the initial policy learning. We investigate the computational complexity of the proposed scheme and derive the performance bound regarding the message bit error rate, the swarm energy consumption and the utility. Simulation results show that the proposed scheme improves the swarm communication performance and saves energy consumption compared with the benchmark scheme. Zefang Lv, Guohang Niu, Liang Xiao 0003, Chengwen Xing, Wenyuan Xu 0001 |
ICC | 1 |
| 2023 | Multi-Agent Reinforcement Learning Based UAV Swarm Communications Against JammingabstractThe swarm relay and power allocation policy determines the bit error rate and the energy consumption of unmanned aerial vehicles (UAVs) and can be optimized based on the network and jamming model, which is rarely known by UAVs. In this paper, we propose a multi-agent reinforcement learning (RL)-based UAV swarm communication scheme to optimize the relay selection and power allocation against jamming. Based on the network topology, channel states, previous performance and observations shared by the neighboring UAVs, this scheme formulates the policy distribution to improve the policy exploration and applies a policy learning mechanism to stabilize the learning process. Based on transfer learning, the shared swarm experiences are exploited to accelerate the initial learning and improve policy optimization. A deep RL-based scheme is proposed to mitigate the state quantization error for the rapidly changing channel states under high swarm moving speed and thus further improve the anti-jamming performance. This scheme designs a policy network with four fully connected layers to approximate the policy distribution and uses another two neural networks to estimate the average policy distribution and the expected long-term utility, respectively, to update the policy network for stabilized deep learning. We investigate the computational complexity and derive the performance bound regarding the bit error rate, the energy consumption and the utility. Simulation and experimental results verify the performance gain of our proposed schemes over related works. Zefang Lv, Liang Xiao 0003, Yousong Du, Guohang Niu, Chengwen Xing, Wenyuan Xu 0001 |
IEEE Trans. Wirel. Commun. | 1 |
| 2022 | Reinforcement Learning Based Vulnerability Analysis for Smart Grids Against False Data Injection Attacks
Liang Xiao 0003, Zefang Lv |
WASA (1) | 4 |
| 2020 | Autonomous and Privacy-preserving Energy Trading Based on Redactable Blockchain in Smart GridabstractWith the development of information and communication technologies in smart grid, peer-to-peer (P2P) energy trading for distributed energy resources (DER) has achieved an efficient two-way flow of information and power. The adoption of blockchain technology makes the P2P energy trading more secure and transparent. Considering that users with extra energy may be reluctant to participate in the energy trading due to privacy concerns, many researchers have focused on potential privacy issues. However, most of the existing works are built on top of a semi-decentralized energy blockchain in which only a few certified third-party nodes are authorized to manage and verify transactions - once these authorized nodes are attacked, the system will be threatened. In this paper, we use the Ciphertext Policy Attribute-Based Encryption (CP-ABE) scheme to establish a blockchain based P2P energy trading approach with privacy preservation, which allows peer nodes, including sellers and purchasers, to manage and verify transactions autonomously without needing any additional third-party nodes. In addition, we introduce the redactable blockchain technology into our scheme to ensure users can modify their sensitive information uploaded to the blockchain. Furthermore, we improve the CP-ABE scheme to provide low latency in the system. The experimental evaluations show that our scheme is efficient and practical. Wenti Yang, Zhitao Guan, Longfei Wu, Xiaojiang Du, Zefang Lv, Mohsen Guizani |
GLOBECOM | 5 |
| 2019 | Achieving data utility-privacy tradeoff in Internet of Medical Things: A machine learning approach
Zhitao Guan, Zefang Lv, Xiaojiang Du, Longfei Wu, Mohsen Guizani |
Future Gener. Comput. Syst. | 2 |