Zhiyu Mou

dblp:43/8491 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2024
0000-0002-3260-1420ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Complete Coverage Path Planning for Data Collection with Multiple UAVs
abstract
The utilization of unmanned aerial vehicles (UAVs) for communication data collection across all areas can be modeled as a complete coverage path planning (CCPP) problem. To address the challenge of lengthy coverage time in traditional CCPP algorithms, we propose a weighted balanced graph partitioning based complete coverage path planning scheme (WBGPP), which consists of two sub-algorithm: weighted balanced graph partitioning (Weighted B-GRAP) and single agent path planning (SAPP). The Weighted B-GRAP algorithm can decompose the multi-UAV CCPP problem into multiple single UAV CCPP problems by assigning each UAV a responsibility area according to its capability. Then, we optimize the backtracking strategy through breadth-first search and design a SAPP algorithm to reduce the number of repeated visits and shorten the coverage time. The simulation results show that the proposed WBGPP scheme effectively reduce the coverage time of multiple UAVs in CCPP problems and can be applied to various maps.
Zhiyu Mou, Bo Lin 0010, Feifei Gao 0001
WCNC2
2022 Sustainable Online Reinforcement Learning for Auto-bidding
abstract
Recently, auto-bidding technique has become an essential tool to increase the revenue of advertisers. Facing the complex and ever-changing bidding environments in the real-world advertising system (RAS), state-of-the-art auto-bidding policies usually leverage reinforcement learning (RL) algorithms to generate real-time bids on behalf of the advertisers. Due to safety concerns, it was believed that the RL training process can only be carried out in an offline virtual advertising system (VAS) that is built based on the historical data generated in the RAS. In this paper, we argue that there exists significant gaps between the VAS and RAS, making the RL training process suffer from the problem of inconsistency between online and offline (IBOO). Firstly, we formally define the IBOO and systematically analyze its causes and influences. Then, to avoid the IBOO, we propose a sustainable online RL (SORL) framework that trains the auto-bidding policy by directly interacting with the RAS, instead of learning in the VAS. Specifically, based on our proof of the Lipschitz smooth property of the Q function, we design a safe and efficient online exploration (SER) policy for continuously collecting data from the RAS. Meanwhile, we derive the theoretical lower bound on the safety degree of the SER policy. We also develop a variance-suppressed conservative Q-learning (V-CQL) method to effectively and stably learn the auto-bidding policy with the collected data. Finally, extensive simulated and real-world experiments validate the superiority of our approach over the state-of-the-art auto-bidding algorithm.
Zhiyu Mou, Yusen Huo, Rongquan Bai, Mingzhou Xie, Chuan Yu 0002, Jian Xu 0015, Bo Zheng 0007
NeurIPS1
2022 Resilient UAV Swarm Communications With Graph Convolutional Neural Network
abstract
In this paper, we study the self-healing problem of unmanned aerial vehicle (UAV) swarm network (USNET) that is required to quickly rebuild the communication connectivity under unpredictable external destructions (UEDs). Firstly, to cope with theone-off UEDs, we propose a graph convolutional neural network (GCN) that can find the recovery topology of the USNET in an on-line manner. Secondly, to cope withgeneral UEDs, we develop a GCN based trajectory planning algorithm that can make UAVs rebuild the communication connectivity during the self-healing process. We also design a meta learning scheme to facilitate the on-line executions of the GCN. Numerical results show that the proposed algorithms can rebuild the communication connectivity of the USNET more quickly than the existing algorithms under both one-off UEDs and general UEDs. The simulation results also show that the meta learning scheme can not only enhance the performance of the GCN but also reduce the time complexity of the on-line executions.
Zhiyu Mou, Feifei Gao 0001, Jun Liu 0063, Qihui Wu 0001
IEEE J. Sel. Areas Commun.1
2021 Three-Dimensional Area Coverage with UAV Swarm based on Deep Reinforcement Learning
abstract
In this paper, we study the fast coverage problem of 3D irregular terrain surfaces with a hierarchical UAV swarm. We first build a 3D model of a random irregular terrain and project the 3D terrain surface into many weighted 2D patches. Then we develop a two-level hierarchical UAV swarm architecture, including the low-level follower UAVs (FUAVs) and the high-level leader UAVs (LUAVs). For FUAVs, we adopt the traditional coverage trajectory algorithm to carry out specific coverage tasks within patches based on the star communication topology. For LUAVs, we propose a swarm deep Q-learning (SDQN) reinforcement learning algorithm to select patches. The numerical results show that the total coverage time of the SDQN is less than that of existing methods, which demonstrates the effectiveness of the proposed algorithm.
Zhiyu Mou, Yu Zhang 0047, Feifei Gao 0001, Tao Zhang 0006, Zhu Han 0001
ICC1
2021 Hierarchical Deep Reinforcement Learning for Backscattering Data Collection With Multiple UAVs
abstract
The emerging backscatter communication technology is recognized as a promising solution to the battery problem of Internet of Things (IoT) devices. For example, the wireless sensor network with backscatter communication technology can monitor the environment in remote areas without battery maintenance or replacement. Unfortunately, the transmission range of backscatter communication is limited. To tackle this challenge, we propose a multi-UAV-aided data collection scenario where the unmanned aerial vehicle (UAV) can fly close to the backscatter sensor node (BSN) to activate it and then collects the data. We aim to minimize the total flight time of the rechargeable UAVs when the collection mission is finished. During the data collection process, the UAVs can return to the charging station to recharge itself when the energy of UAV is not sufficient to complete the mission. To reduce the complexity of the task, we first use the Gaussian mixture model clustering method to divide the BSNs into multiple clusters. Then we consider the deterministic boundary and ambiguous boundary for the UAV flying regions, respectively. For the deterministic boundary scenario, we propose a single-agent deep option learning (SADOL) algorithm, where each UAV cannot fly beyond the deterministic boundary. For the ambiguous boundary scenario, we propose a multiagent deep option learning (MADOL) algorithm to enable the UAVs to cooperatively learn the ambiguous BSNs assignment. In the simulation, we compare the proposed algorithms with multiagent deep deterministic policy gradient (MADDPG), deep deterministic policy gradient (DDPG), and deep Q-network (DQN) algorithms, which proves the proposed algorithms can achieve better performance.
Yu Zhang 0047, Zhiyu Mou, Feifei Gao 0001, Ling Xing 0001, Jing Jiang 0026, Zhu Han 0001
IEEE Internet Things J.2
2021 Deep Reinforcement Learning Based Three-Dimensional Area Coverage With UAV Swarm
abstract
Unmanned aerial vehicle (UAV) technology is recognized as a promising solution to area coverage problems (ACPs) and has been extensively studied recently. In this paper, we study the 3D irregular terrain surface coverage problem with a hierarchical UAV swarm. We first build the 3D model of a random irregular terrain and propose a geometric way to project the 3D terrain surface into many weighted 2D patches. Then we develop a two-level hierarchical UAV swarm architecture, including the low-level follower UAVs (FUAVs) and the high-level leader UAVs (LUAVs). For FUAVs, we design a coverage trajectory algorithm to carry out specific coverage tasks within patches based on the star communication topology. For LUAVs, we propose a swarm deep Q-learning (SDQN) reinforcement learning algorithm to select patches. Moreover, an observation history model based on convolutional neural networks (CNNs) and the mean embedding method is integrated into SDQN to address the communication limitation problems of LUAVs. The numerical results show that FUAVs can cover the entire area of each patch with little redundancies, and the total coverage time of the SDQN is less than that of existing methods, which demonstrates the effectiveness of the proposed algorithms.
Zhiyu Mou, Yu Zhang 0047, Feifei Gao 0001, Tao Zhang 0006, Zhu Han 0001
IEEE J. Sel. Areas Commun.1