EDBT 2026 Demo / reviewers in the wild / expert
Mengjie Yi
dblp:237/5302
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0007-7851-8335ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 4 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fast and Generalizable Task Scheduling in Double-Layered Satellite Network: A Graph-Based Deep Reinforcement Learning Approach
Zhonghe Liu, Runzi Liu, Yan Zhang 0006, Mengjie Yi |
WCNC | 4 |
| 2025 | Meta-Reinforcement Learning for Timely and Energy-Efficient Data Collection in Solar-Powered AAV-Assisted IoT NetworksabstractAutonomous aerial vehicles (AAVs) have the potential to greatly aid Internet of Things (IoT) networks in mission-critical data collection, thanks to their flexibility and cost-effectiveness. However, challenges arise due to the AAV’s limited onboard energy and the unpredictable status updates from sensor nodes (SNs), which impact the freshness of collected data. In this paper, we investigate the energy-efficient and timely data collection in IoT networks through the use of a solar-powered AAV. Each SN generates status updates at stochastic intervals, while the AAV collects and subsequently transmits these status updates to a central data center. Furthermore, the AAV harnesses solar energy from the environment to maintain its energy level above a predetermined threshold. To minimize both the average age of information (AoI) for SNs and the energy consumption of the AAV, we jointly optimize the AAV trajectory, SN scheduling, and offloading strategy. Then, we formulate this problem as a Markov decision process (MDP) and propose a meta-reinforcement learning algorithm to enhance the generalization capability. Specifically, the compound-action deep reinforcement learning (CADRL) algorithm is proposed to handle the discrete decisions related to SN scheduling and the AAV’s offloading policy, as well as the continuous control of AAV flight. Moreover, we incorporate meta-learning into CADRL to improve the adaptability of the learned policy to new tasks. To validate the effectiveness of our proposed algorithms, we conduct extensive simulations and demonstrate their superiority over other baseline algorithms. Mengjie Yi, Xijun Wang 0001, Juan Liu 0002, Yan Zhang 0006, Ronghui Hou |
IEEE Trans. Commun. | 1 |
| 2024 | Satellite-Assisted UAV Data Collection for Information Freshness in IoRT NetworksabstractUtilizing UAVs and satellites can offer an effective means to collect data for the Internet of remote things (IoRT) networks. However, due to the limited energy of UAVs and the high cost of satellite communication, ensuring the reduction of UAV energy consumption and communication costs while collecting fresh data poses a significant challenge. In this paper, we explore the issue of data gathering in IoRT networks with the assistance of UAVs and satellites. The UAV gathers data from sensor nodes (SNs) and decides whether to relay the collected data via satellite or send it directly to the data processing center. We handle this problem by formulating it as a Markov decision process to minimize the combined weighted sum of the average age of information, the energy consumption of the UAV, and communication costs through the implementation of a compound-action proximal policy optimization (CPPO) method. It can handle the compound actions of the UAV. This approach simultaneously optimizes the UAV's path, SN scheduling, and transmission decisions. Simulation results demonstrate that our algorithm can achieve better performance compared to baseline methods. Mengjie Yi, Yan Zhang 0006, Xijun Wang 0001, Juan Liu 0002 |
WCNC | 2 |
| 2024 | Meta-Learning Deep Reinforcement Learning for Fresh Data Collection in UAV-Assisted Wireless Sensor Networks
Mengjie Yi, Xijun Wang 0001, Juan Liu 0002, Yan Zhang 0006, Ronghui Hou |
WiOpt | 2 |
| 2023 | Deep Reinforcement Learning for Energy-Efficient Fresh Data Collection in Rechargeable UAV-assisted IoT NetworksabstractThe unmanned aerial vehicle (UAV) can act as the edge server in delay-sensitive monitoring for data collection and processing in the Internet of things (IoT) networks due to its flexibility and low operational cost. One of its major disadvantages is the limited battery level. This paper focuses on a problem with the rechargeable UAV-assisted energy-efficient and fresh data collection in the IoT networks. In particular, the UAV takes off from the initial position to collect data packets from sensor nodes (SNs) in the IoT networks and needs to reach the final position at a given time. Some charging stations (CSs) are in the IoT networks, which can recharge the UAV by the wireless power transfer technique to keep the UAV’s energy level from falling below the threshold energy. To minimize the weighted sum of the average age of information (AoI) and the average recharging price, we design a Markov Decision Process (MDP) to determine the UAV’s flight trajectory, the scheduling of SNs, and energy recharging. The MDP is then solved using a rechargeable UAV-assisted data collection algorithm based on dueling double deep Q-networks (D3QN). Numerous simulations show that the proposed D3QN algorithm can reduce the weighted sum of the average AoI and the average recharging price more effectively than the baseline algorithms. Mengjie Yi, Xijun Wang 0001, Juan Liu 0002, Yan Zhang 0006, Ronghui Hou |
WCNC | 1 |
| 2023 | Multitask Transfer Deep Reinforcement Learning for Timely Data Collection in Rechargeable-UAV-Aided IoT NetworksabstractThanks to their high-flexibility and low-operational cost, unmanned aerial vehicles (UAVs) can be used to support mission-critical applications in the Internet of Things (IoT). However, due to the limited onboard energy, it is difficult for UAVs to provide continuous data collection. In this article, we study the problem of rechargeable-UAV-aided timely data collection in IoT networks, where the UAV collects status updates from multiple sensors and gets recharged from the charging stations (CSs) to keep its energy level above a threshold. To tradeoff the information freshness and energy consumption, we formulate a Markov decision process (MDP) with the objective of minimizing the weighted sum of the average total Age of Information and average recharging price. Under the dynamics and uncertainty of the environment, we propose a multitask transfer deep reinforcement learning method to jointly optimize the UAV ’ s flight trajectory, transmission scheduling, and battery recharging. To enable the application of the learned policy to new environments with similar settings and avoid starting from scratch, we develop a multitask network made up of common knowledge layers and task-specific knowledge layers. It specifically makes it possible for the transfer of common knowledge between environments with different network scales (e.g., different numbers of sensors/CSs) and/or topologies (e.g., different locations of sensors/CSs). Simulation results demonstrate that the proposed algorithm can adapt to new environments and achieve superior performance compared to the baseline algorithms. Mengjie Yi, Xijun Wang 0001, Juan Liu 0002, Yan Zhang 0006, Ronghui Hou |
IEEE Internet Things J. | 1 |
| 2023 | Cooperative Data Collection With Multiple UAVs for Information Freshness in the Internet of ThingsabstractMaintaining the freshness of information in the Internet of Things (IoT) is a critical yet challenging problem. In this paper, we study cooperative data collection using multiple Unmanned Aerial Vehicles (UAVs) with the objective of minimizing the total average Age of Information (AoI). We consider various constraints of the UAVs, including kinematic, energy, trajectory, and collision avoidance, in order to optimize the data collection process. Specifically, each UAV, which has limited on-board energy, takes off from its initial location and flies over sensor nodes to collect update packets in cooperation with the other UAVs. The UAVs must land at their final destinations with non-negative residual energy after the specified time duration to ensure they have enough energy to complete their missions. It is crucial to design the trajectories of the UAVs and the transmission scheduling of the sensor nodes to enhance information freshness. We model the multi-UAV data collection problem as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP), as each UAV is unaware of the dynamics of the environment and can only observe a part of the sensors. To address the challenges of this problem, we propose a multi-agent Deep Reinforcement Learning (DRL)-based algorithm with centralized learning and decentralized execution. In addition to the reward shaping, we use action masks to filter out invalid actions and ensure that the constraints are met. Simulation results demonstrate that the proposed algorithms can significantly reduce the total average AoI compared to the baseline algorithms, and the use of the action mask method can improve the convergence speed of the proposed algorithm. Xijun Wang 0001, Mengjie Yi, Juan Liu 0002, Yan Zhang 0006, Meng Wang 0019, Bo Bai 0001 |
IEEE Trans. Commun. | 2 |
| 2021 | Deep Reinforcement Learning for User Association in Heterogeneous Networks with Dual ConnectivityabstractThe dual connectivity is emerging as a promising solution to boost capacity in heterogeneous networks. However, it is challenging to obtain an optimal user association in heterogeneous networks with dual connectivity, due to its non-convex and combinatorial nature. In this paper, we propose a user association scheme based on deep reinforcement learning to maximize the overall network utility, which takes both throughput and user fairness into account, in the downlink of a heterogeneous network. Particularly, each user associates with the macro base station (BS) and a micro BS. We apply a deep Q-network (DQN) to obtain the nearly optimal policy to associate the users and micro BSs. Simulation results demonstrate that DQN-based user association performs better compared to the conventional user association schemes in heterogeneous networks with dual connectivity, and it behaves good scalability when the environment changes. Mengjie Yi, Yan Zhang 0006, Xijun Wang 0001, Chao Xu 0007, Xiao Ma 0007 |
WCNC | 1 |