VLDB 2026 Research / reviewers in the wild / expert
Yousong Du
dblp:363/9941
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0002-1480-4443ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reinforcement-Learning-Based APT Defense for Large-Scale Smart GridsabstractReinforcement-learning-(RL)-based advanced persistent threats (APTs) defense schemes choose the scan interval to enhance the detection accuracy, and the data protection level can be further improved by optimizing the repair rate corresponding to methods, such as malicious files removal, security software version update and passwords reset, each requiring different computational resources, such as CPUs. In this article, we propose an RL-based APT defense scheme to optimize both the continuous scan interval of metering data and repair rate to mitigate the potential loss for meter data management systems (MDMSs) in smart grids with large-scale meters. Based on the size of the metering data stored at MDMS and the compromised data, the data tag granularity and the number of CPUs for repairing, neural networks extract the state feature, address the quantization error of defense policy, and update the weights based on the APT defense experiences and the shared weights of neighboring MDMSs to improve the defense performance as a weighted sum of the defense duration, detection accuracy and data protection level. The computational complexity of the proposed scheme and the performance bounds according to the Nash Equilibrium of the game between the MDMS and the APT attacker are provided. Simulation results based on three MDMSs that receive the metering data from 50–150 smart meters and an APT attacker with the selected attack interval up to 5 s show the performance gain over the benchmark based on the optimal control theory and Q-learning in large-scale smart grids. Liang Xiao 0003, Zefang Lv, Zhiping Lin 0002, Yousong Du |
IEEE Internet Things J. | 6 |
| 2024 | Efficient Communications in Multi-Agent Reinforcement Learning for Mobile ApplicationsabstractThe environment observations and learning experiences shared by the cooperative learning agents accelerate multi-agent reinforcement learning (MARL) with partial observations for mobile applications but the performance degrades due to the redundant and outdated observations under severe channel fading in wireless networks. In this paper, we propose an efficient communication scheme in MARL for mobile applications that enables each learning agent to optimize the cooperative agents and the learning parameters to integrate the shared information. The cooperative agents are chosen according to the learning environment observations, the channel states, and the task similarity with neighboring agents. The learning parameters are chosen based on the attention mechanism that exploits the correlation with the local observation to enhance the agent receptive field for efficient policy exploration. Neural networks with weights updated based on the learning factors determined by the task similarity are designed to further improve the learning efficiency. The performance bounds including the information gain from the learning agent cooperation, the communication cost and the utility are provided based on the Nash equilibrium of the cooperative MARL communication game. The proposed scheme is implemented in the anti-jamming video transmission of the unmanned aerial vehicle swarms to optimize the transmit channel and power and experimental results verify the performance gain over the benchmark. Zefang Lv, Liang Xiao 0003, Yousong Du, Yunjun Zhu, Shuai Han 0002, Yong-Jin Liu 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2023 | Efficient Communications for Multi-Agent Reinforcement Learning in Wireless NetworksabstractMulti-agent reinforcement learning (RL) utilizes the observations and learning experiences shared among the agents to accelerate learning speed under partial observations and the resulting learning efficiency depends on the cooperative agent selection and the RL task state formulation. In this paper, we propose an efficient communication scheme for multi-agent RL that enables each learning agent to optimize the cooperative agent selection and the task state formulation to improve the learning performance and the quality of service for RL-based applications in wireless networks. Based on the local observation, the radio channel states, the similarity of RL task with neighboring agents and previous communication cost, this scheme formulates a communication state, which is input to a neural network to estimate the communication policy distribution. The RL task state of the learning agent, which consists of the local observation such as channel states and previous task performance, as well as the correlation between the shared and the local observation extracted based on the attention mechanism, is formulated to enhance the agent receptive field. In addition, the shared learning information is also exploited to update the local learning parameters such as the task Q-values and neural network weights and further improve the RL task policy exploration. As a case study, the proposed communication scheme is implemented in the multi-agent deep Q-network based anti-jamming unmanned aerial vehicle swarm communications and the performance gain over the benchmark is verified via simulation results. Zefang Lv, Yousong Du, Liang Xiao 0003, Shuai Han 0002, Xiangyang Ji |
GLOBECOM | 2 |
| 2023 | Reliable Communications for Hypersonic Vehicles: A Reinforcement Learning ApproachabstractThe ultra-high speed (e.g., typically moving with 10–20 Mach) of the hypersonic vehicle (HSV) causes a plasma sheath, which severely degrades the communication performance and results in communication blackouts. In this paper, we propose a deep reinforcement learning (RL)-based HSV reliable communications scheme against jamming, which enables the HSV to select the carrier frequency and transmit power according to the signal quality, the estimated voltage standing wave ratio, flight altitude, flight speed, and angle of attack. Specifically, we design a deep two-level hierarchical structure to compress the high-dimensional state and action space, with the added advantage of leveraging transfer learning to reduce initial exploration and expedite the optimization process. To optimize the learning speed, the dueling architecture is implemented in the deep network to measure the state value and the advantage function of the policies. In contrast to the benchmark, the simulation results indicate that the proposed scheme yields a significant reduction in both bit error rate and transmit power. Jingchen Xu, Zhiping Lin 0002, Yousong Du, Helin Yang, Liang Xiao 0003 |
GLOBECOM | 4 |
| 2023 | Multi-Agent Reinforcement Learning Based UAV Swarm Communications Against JammingabstractThe swarm relay and power allocation policy determines the bit error rate and the energy consumption of unmanned aerial vehicles (UAVs) and can be optimized based on the network and jamming model, which is rarely known by UAVs. In this paper, we propose a multi-agent reinforcement learning (RL)-based UAV swarm communication scheme to optimize the relay selection and power allocation against jamming. Based on the network topology, channel states, previous performance and observations shared by the neighboring UAVs, this scheme formulates the policy distribution to improve the policy exploration and applies a policy learning mechanism to stabilize the learning process. Based on transfer learning, the shared swarm experiences are exploited to accelerate the initial learning and improve policy optimization. A deep RL-based scheme is proposed to mitigate the state quantization error for the rapidly changing channel states under high swarm moving speed and thus further improve the anti-jamming performance. This scheme designs a policy network with four fully connected layers to approximate the policy distribution and uses another two neural networks to estimate the average policy distribution and the expected long-term utility, respectively, to update the policy network for stabilized deep learning. We investigate the computational complexity and derive the performance bound regarding the bit error rate, the energy consumption and the utility. Simulation and experimental results verify the performance gain of our proposed schemes over related works. Zefang Lv, Liang Xiao 0003, Yousong Du, Guohang Niu, Chengwen Xing, Wenyuan Xu 0001 |
IEEE Trans. Wirel. Commun. | 3 |