VLDB 2026 Research / reviewers in the wild / expert
Chao Wang 0153
dblp:188/7759-153
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
0009-0007-3911-4645ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 2 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A logarithmic-approximation approach for bandwidth efficient MoE deployment in edge networks
Xiangning Lu, Chao Wang 0153, Danyang Zheng 0001, Xiangyi Chen, Huanlai Xing |
Comput. Networks | 3 |
| 2026 | Heterogeneous Dual-Agent DRL with generalization for SFC shared protection
Yihan Zhong, Chao Wang 0153, Honghui Xu 0001, Danyang Zheng 0001, Xiaojun Cao |
Comput. Networks | 2 |
| 2026 | A Provably Cost-Efficient Approach to Deploying MoE Inference Models at the Network Edge
Chao Wang 0153, Danyang Zheng 0001, Huanlai Xing, Chen Yang 0043, Xiaojun Cao, Jie Xu 0007, Fei Teng 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2025 | Heuristic-guided Migration-Agent-Based DRL for Compressed Model Placement in Edge NetworksabstractThe model compression techniques enable deploying compressed large language models (CMs) at network edge, facilitating convenient provision of AI-generated content (AIGC) services. To ensure timely delivery of these services, efficient placement of CMs across resource-constrained edge networks is essential. In this work, we investigate how to obtain latency-efficient CM placement across resource-constrained edge networks. With the objective of service latency optimization, we formulate the CM placement in resource-constrained network (CPRN) problem and establish its NP-hardness. We propose the Migration-agent-based Deep Reinforcement Learning (M-DRL) approach, which incorporates a specially designed migration agent tailored for such placement problems. To enhance training efficiency, we incorporate efficient heuristic placement results into the environment of M-DRL, developing our Heuristic-guided M-DRL (HM-DRL) approach. Our extensive simulation results demonstrate that HM-DRL outperforms an extended benchmark in service latency, while maintaining a low training overhead. Chao Wang 0153, Danyang Zheng 0001, Yihan Zhong, Honghui Xu 0001, Xiaojun Cao |
GLOBECOM | 1 |
| 2025 | Cost-Efficient Knowledge Distillation-enabled Student Models Placement in Edge NetworksabstractTo support edge intelligence, knowledge distillation (KD) is widely employed to compress large language models (LLMs) into smaller, domain-specific student models. However, due to the limited generalization capabilities of student models, they may fail to provide accurate responses across diverse domains. In such cases, the teacher model serves as a complementary component, handling queries that exceed the scope of student models. In line with the KD paradigm, this work investigates a collaborative deployment framework in which multiple student models are distributed across the network edge to serve the majority of client requests. In contrast, a centralized teacher model addresses more complex or ambiguous queries. To begin, we formally define the Student Model Placement in Edge Networks (SMP-EN) problem, aiming to minimize total access costs. We prove that SMP-EN is NP-hard, and to address this challenge, we introduce an Access Cost Measure (ACM) that quantifies the expected access costs. Building upon this measure, we propose the ACM-based Student Model Placement (ACM-SMP) algorithm to determine student model placement efficiently. Extensive simulations show that ACM-SMP significantly reduces the average expected client access cost compared to benchmarks. Weiqing Zeng, Danyang Zheng 0001, Huanlai Xing, Wenting Wei, Chao Wang 0153, Xiaojun Cao |
GLOBECOM | 5 |
| 2025 | Towards Expert Models Deployment Cost Optimization in Edge Computing NetworksabstractWith the widespread adoption of large language models (LLMs) like GPT, user experiences in various interactive applications have significantly improved. However, reports from OpenAI highlight that GPT clients are now facing high response delays and frequent interruptions, particularly during peak usage hours, due to limited computation resources. This challenge is expected to escalate as machines are interacting with GPT models at higher frequencies, with greater data volumes, and over longer lifecycles. A promising solution is to deploy LLMs across edge networks to efficiently distribute the huge resource demands. This work presents the very first efforts at exploring how to cost-effectively deploy expert models from a mixture of experts (MoE) LLM within edge networks. We introduce the expert models deployment in edge networks (EMD-EN) problem, focusing on optimizing deployment costs. To address this, we propose a novel least cost gain (LCG) measure for selecting appropriate physical nodes to host expert models and present a corresponding LCG-based expert models deployment (LCGEMD) algorithm. Extensive simulations show that our approach outperforms the benchmarks by an average of 17.31% and 36.98% in terms of deployment cost reduction. Chao Wang 0153, Yihan Zhong, Shaohua Cao, Danyang Zheng 0001, Xiaojun Cao |
ICC | 2 |
| 2025 | Towards Profits Optimization in LLM Inference Model Deployment at the Network EdgeabstractRecent advances in large language models (LLMs) have empowered robots and drones with autonomous decision-making capabilities. Due to the stringent real-time requirements of these applications, LLM inference must be performed at the network edge. However, hosting high-precision LLMs on a single edge server is often infeasible, creating challenges in efficiently distributing LLM deployments across edge networks. This work addresses these challenges by formulating and solving the profit maximization problem for distributed LLM inference deployment. We first formally define the Profit-Centric Inference Chain Deployment (PC-InCD) problem. To solve PC-InCD, we introduce a novel Local Maximal Profit (LMP) factor that enables effective edge server selection for hosting LLM sub-modules, and we propose the LMP-based Inference Chain Deployment (LMP-InCD) algorithm. Extensive simulations demonstrate that LMP-InCD significantly outperforms benchmark methods in maximizing profit across diverse network conditions. Danyang Zheng 0001, Huanlai Xing, Honghui Xu 0001, Chengzong Peng, Chao Wang 0153, Xiaojun Cao |
IPCCC | 6 |
| 2025 | Towards cost optimization in security-aware service function chaining and embedding over multi-vendor edge networks
Chao Wang 0153, Danyang Zheng 0001, Wenyi Tang, Honghui Xu 0001, Xiaojun Cao |
Comput. Networks | 1 |
| 2025 | A provably efficient in-network computing services deployment approach for security burst
Danyang Zheng 0001, Chao Wang 0153, Honghui Xu 0001, Wenyi Tang, Yihan Zhong, Xiaojun Cao |
Comput. Networks | 2 |