VLDB 2026 Research / reviewers in the wild / expert
Yijun Hao
dblp:340/7210
· DBLP profile ↗
5ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0001-7636-2786ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | T³Planner: Multi-Phase Planning Across Structure-Constrained Optical, IP, and Routing TopologiesabstractNetwork topology planning is an essential multi-phase process to build and jointly optimize the multi-layer network topologies in wide-area networks (WANs). Most existing practices target single-phase/layer planning, and are incapable of satisfying all rigorous topological structure constraints (e.g., dual-homing rings) defined by network standards and operators, especially in large-scale networks. These significantly limit their usability and performance in production networks. We consider a general topology planning problem with typical structure constraints over three essential phases (greenfield, reconfiguration, and site expansion) and topological layers (optical, IP, and routing topologies). We present, T3Planner, a novel practical solver to this problem in production. Specifically, we develop a structure-driven encoder based on graph neural network (GNN) for concise structure encoding, and design a new learning framework with optical-centric layer compression/reconstruction and rule-aided reinforcement learning (RL) for fast convergence and high performance. Extensive experiments on nine real topologies demonstrate that T3Planner scales to large optical networks with hundreds of sites, saves 46.6% cost, and supports$3.12\times $more demand when compared to related existing approaches. Yijun Hao, Shusen Yang, Cong Zhao 0001, Xuebin Ren, Peng Zhao 0001, Chenren Xu, Shibo Wang 0002 |
IEEE J. Sel. Areas Commun. | 1 |
| 2025 | Learning Adaptive Multi-Timescale Scheduling for Mobile Edge ComputingabstractIn mobile edge computing (MEC), resource scheduling is crucial to task requests’ performance and service providers’ cost, involving multi-layer heterogeneous scheduling decisions. Existing MEC schedulers typically adopt static-timescale scheduling, where scheduling decisions are updated regularly at fixed intervals for all layers. The inflexible updating timescales lead to poor performance in the production networks. In this paper, we propose EdgeTimer, an unprecedented approach that automatically and adaptively determines respective updating timescales of multiple scheduling layers to achieve a better trade-off between the operation cost and service performance. Specifically, we design (i) a three-layer hierarchical deep reinforcement learning (DRL) framework for efficient learning of tightly coupled policies, (ii) a tailored multi-agent DRL algorithm for decentralized scheduling, with the convergence strictly proved, and (iii) a lightweight system defender for deterministic reliability assurance. Furthermore, we apply EdgeTimer to a wide range of Kubernetes scheduling rules, and evaluate it using production traces with different workload patterns. Through extensive trace-driven experiments, we demonstrate that EdgeTimer can significantly decrease the operation cost for service providers without sacrificing the delay performance, thereby improving overall profits, compared with the state-of-the-art approaches. Yijun Hao, Shusen Yang, Shibo Wang 0002, Xuebin Ren |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | OACR$^{2}$2: Online Admission Control and Resource Reservation for 5G Slice Networks With Deep Reinforcement LearningabstractNetwork slicing architecture is expected to fulfill network applications with heterogeneous requirements through efficient slice admission control (SAC) policies. Existing SAC approaches entirely rely on current limited observations to make admission decisions, ignoring the potential impact of future demands. The short-sighted behaviors lead to poor service performance and infrastructure providers’ (InPs’) revenue in practice. In this paper, we propose OACR$^{2}$, an online SAC approach based on deep reinforcement learning (DRL) that can exploit predictable future requests to make more precise admission control decisions for the long-term revenue, and reserve proper resources accordingly. Specifically, we design three novel schemes: (i) a requirement predictor based on long short-term memory (LSTM) and a novel input-output way to predict future unforeseen requests, (ii) a DRL admission controller based on the partially observable Markov decision process model to make precise admission decisions without accurate future request information, with the convergence strictly proved, and (iii) a decision defender to guarantee decision reliability. Extensive experiments on real-world traces demonstrate that compared to the No-wait, Wait-queue, and Wait-earliest time approaches, OACR$^{2}$improves InPs’ revenue and acceptance ratio by up to 40.9% and 16.7%, respectively, without sacrificing online inference time (within 0.9 milliseconds). Yijun Hao, Shusen Yang, Peng Zhao 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | EdgeTimer: Adaptive Multi-Timescale Scheduling in Mobile Edge Computing with Deep Reinforcement LearningabstractIn mobile edge computing (MEC), resource scheduling is crucial to task requests’ performance and service providers’ cost, involving multi-layer heterogeneous scheduling decisions. Existing schedulers typically adopt static timescales to regularly update scheduling decisions of each layer, without adaptive adjustment of timescales for different layers, resulting in potentially poor performance in practice.We notice that the adaptive timescales would significantly improve the trade-off between the operation cost and delay performance. Based on this insight, we propose EdgeTimer, the first work to automatically generate adaptive timescales to update multi-layer scheduling decisions using deep reinforcement learning (DRL). First, EdgeTimer uses a three-layer hierarchical DRL framework to decouple the multi-layer decision-making task into a hierarchy of independent sub-tasks for improving learning efficiency. Second, to cope with each sub-task, EdgeTimer adopts a safe multi-agent DRL algorithm for decentralized scheduling while ensuring system reliability. We apply EdgeTimer to a wide range of Kubernetes scheduling rules, and evaluate it using production traces with different workload patterns. Extensive trace-driven experiments demonstrate that EdgeTimer can learn adaptive timescales, irrespective of workload patterns and built-in scheduling rules. It obtains up to 9:1 more profit than existing approaches without sacrificing the delay performance. Yijun Hao, Shusen Yang, Shibo Wang 0002, Xuebin Ren |
INFOCOM | 1 |
| 2023 | Delay-Oriented Scheduling in 5G Downlink Wireless Networks Based on Reinforcement Learning With Partial Observationsabstract5G wireless networks are expected to satisfy different delay requirements of various traffics by network resource scheduling. Existing scheduling methods perform poorly in practice due to their unrealistic assumption on the access to the full channel state information (CSI) or the explicit mathematical expression of network delay. In this paper, we consider the delay-oriented packet scheduling problem in multi-cell 5G downlink networks with multiple users and traffic types (e.g., FTP, VoIP and video streaming), and formulate it as a partially observable Markov decision process (POMDP). We design a delay-oriented downlink scheduling framework based on deep reinforcement learning (DRL) to autonomously schedule the active traffic flows without the full channel information. Furthermore, a recurrent proximal policy optimization (RPPO) algorithm is proposed to perceive the underlying state and accelerate learning under different time granularities, with the policy gradient theorem under POMDP strictly proved. By incorporating the future traffic information provided by a proposed spatial-temporal prediction algorithm, RPPO can balance the load and achieve lower delay in real-time multi-cell multi-user scenarios. Results of extensive experiments on a realistic 5G simulator demonstrate that our framework significantly outperforms existing approaches in terms of both tail delay and average delay for up to 48% and 41.7%, respectively. Yijun Hao, Cong Zhao 0001, Shusen Yang |
IEEE/ACM Trans. Netw. | 1 |