VLDB 2026 Research / reviewers in the wild / expert
Yuhao Chai
dblp:337/7493
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0003-3546-0536ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Distributed Inference Optimization for Large Language Model in Edge-Cloud Collaborative NetworksabstractWith the progressive evolution of large language models (LLMs) and the increasing need of computing for$\mathbf{6 G}$, it becomes crucial for multiple network nodes with limited computing resources to share the need for large model inference. Model partition methods have been proposed to enable computation-intensive artificial intelligence (AI) services by splitting an AI model across multi-edge and cloud nodes. In this paper, a distributed inference optimization framework for transformer decoder-only based LLMs (DIO-LLMs) is proposed in edge-cloud collaborative networks. The partitioning and offloading strategy is determined based on the computing workload and network status. DIO-LLMs specifically accounts for the parallel execution capabilities of the transformer architecture. It employs a two-phase model partitioning strategy, comprising inter-layer and intra-layer partitions, to effectively distribute LLMs across edge and cloud nodes. Additionally, to mitigate inference latency under resource limitations, a Greedy Proximal Policy Optimization (GPPO) based algorithm has been developed to devise optimal strategies. Simulation results indicate that under memory constraints, the proposed algorithm can reduce inference latency more effectively than other baseline algorithms. Zideng Feng, Lu Lu 0016, Yuhao Chai, Zhenyu Zhang 0032, Yong Zhang 0025, Yinglei Teng, Da Guo |
ICC | 4 |
| 2025 | Heterogeneous Request Scheduling and Resource Optimization in Serverless Edge NetworksabstractWith the development of virtualization technology, serverless computing has been gaining significant attention in recent years, primarily due to its advantages in scalability and a pay-as-you-go pricing model. In edge networks, the deployment of fine-grained function instances to handle massive request data makes it more difficult to optimize the quality of user experience. This paper addresses the scheduling of heterogeneous function processing requests generated by users in serverless edge computing scenarios, considering the constraints of resources on edge nodes, and making decisions regarding the warm and cold start during the scheduling process. The problem is modeled as a constrained multi-objective optimization issue aimed at minimizing latency and energy consumption. A deep reinforcement learning strategy, grounded in Multi-Agent Proximal Policy Optimization (MAPPO), is introduced to address this challenge, with each user being represented as an autonomous agent. Simulations were conducted to assess the impacts of the learning rate, the request volume, and the size of the input data of the function. The experimental results indicate that, compared with Particle Swarm Optimization (PSO) and Genetic Algorithm (GA), under different scenarios of request scales, the average system delay is reduced by at most 31%, and the average energy consumption is reduced by at most 24%. Jialu Tian, Yuhao Chai, Nanxiang Shi, Yue Lian, Zhenyu Zhang 0032, Yong Zhang 0025, Yinglei Teng |
VTC2025-Fall | 2 |
| 2025 | Research on joint game theory and multi-agent reinforcement learning-based resource allocation in micro operator networks
Yuhao Chai, Yong Zhang 0025, Zhenyu Zhang 0032, Da Guo, Yinglei Teng |
Comput. Networks | 1 |
| 2025 | Joint AI Service Placement, Task Scheduling, and Resource Allocation for IoT in 6G NetworksabstractAs Internet of Things (IoT)-based artificial intelligence (AI) applications grow, the surge in computational and communication demands has raised concerns about energy consumption, making it critical for 6G networks to address this challenge. This paper examines the joint optimization of AI service placement, task scheduling, and computing resource allocation in an edge-network-cloud system to minimize long-term energy consumption. These problems are interdependent: AI service placement determines service locations, influencing task scheduling, which in turn dictates computing resource allocation. The key challenge lies in the coupling of these variables and the two time-scale nature of the problem, involving long-term (AI service placement) and short-term (task scheduling and computing resource allocation) strategies. To address this, a Hierarchical Markov Decision Process (HMDP) framework is proposed for efficient and coordinated optimization across time scales. A Hierarchical Mean-Field Dueling Double Deep Q-Network (HMFD3QN) algorithm is developed within this framework, where the upper layer optimizes AI service placement, and the lower layer manages task scheduling and computing resource allocation. By integrating mean-field theory, the algorithm reduces the complexity of multi-agent interactions. The computing resource allocation problem is shown to be convex when other variables are fixed, and an optimal strategy is derived using Karush-Kuhn-Tucker (KKT) conditions to simplify the action space for reinforcement learning. Experimental results demonstrate that the proposed method can reduce energy consumption by up to 34% compared to baseline methods, significantly improve queue stability, and increase the proportion of tasks meeting QoS requirements. Zhenyu Zhang 0032, Lu Lu 0016, Yuhao Chai, Di Wu 0078, Yong Zhang 0025 |
IEEE Internet Things J. | 4 |
| 2024 | AI Service Deployment and Resource Allocation Optimization Based on Human-Like Networking ArchitectureabstractIn the forthcoming sixth-generation (6G) era, edge-network-cloud collaboration is needed to support artificial intelligence as a service (AIaaS) with a strong demand for computing power. However, how to guarantee the Quality of AI Service (QoAIS) and utilize the edge-network-cloud collaboration to enhance the performance of AI service is a big challenge. In this paper, we propose an AI service management and network resource scheduling architecture based on human-like networking. Considering the Quality of Service (QoS) requirements and AI tasks, we propose a joint AI agent placement with deep neural network (DNN) deployment and dynamic bandwidth resource allocation algorithm (JAAPD-D). JAAPD-D is proposed to solve the short-term and long-term joint resource allocation problem which includes communication, computation, and memory resources in the network. We adjust the agent placement, DNN deployment, and schedule routing path to ensure effective service transmission in the long time interval and dynamically allocate bandwidth resources in the short time interval. We use Lyapunov optimization to ensure the system stability of the whole network, meet the QoS requirements of various services, and minimize the average end-to-end delay of services. Simulation results show that JAAPD-D outperforms existing algorithms in terms of delay, traffic accepted rate, network system throughput, and cost. Yuhao Chai, Di Wu 0078, Lu Lu 0016, Nanxiang Shi, Yinglei Teng, Yong Zhang 0025 |
IEEE Internet Things J. | 3 |
| 2024 | Joint Task Offloading, Resource Allocation and Model Placement for AI as a Service in 6G NetworkabstractIn the future, 6G network is expected to achieve deep integration of communication and computation, where computation-centric services will be ubiquitous in the network. There are differences in data size, computing power types (CPU/GPU), model complexity, and Quality of Service (QoS) requirements among various CPU computing services and artificial intelligence (AI) services. By providing AI as a Service (AIaaS) in 6G network, the deployment of AI models and the scheduling of task and computing resources can be accelerated. The fundamental challenge lies in the effective amalgamation of the long-term strategy of the model placement problem and the short-term strategy of the task scheduling problem to attain dynamic scheduling and management of tasks and heterogeneous computing resources. A two-timescale optimization method for joint task offloading, computing resource allocation and model placement is proposed in this article. We present an edge-network-cloud framework that configures AIaaS functional units, taking into account the heterogeneous computing requirements and QoS demands of different services. A long-term problem to minimize latency and energy consumption is formulated. To work out the coupled optimization parameters, the problem is decomposed into short-term deterministic sub-problems using Lyapunov optimization. We propose low-complexity algorithms for joint task offloading strategy based on deferred acceptance algorithm, computing resource allocation strategy based on convex optimization, and model placement strategy based on multi-armed bandits. Experimental results demonstrate that our approach outperforms reinforcement learning and other popular optimization algorithms in terms of complexity and effectiveness. Yuhao Chai, Kaice Gao, Guohan Zhang, Lu Lu 0016, Yong Zhang 0025 |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | Research on multi-service slice resource allocation over licensed and unlicensed bands
Yuhao Chai, Yong Zhang 0025, Tengteng Ma, Da Guo, Yinglei Teng |
Wirel. Networks | 1 |