VLDB 2026 Research / reviewers in the wild / expert
Hao Wang 0116
dblp:181/2812-116
· DBLP profile ↗
12ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0001-9883-2400ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 2 first-author · 7 since 2021Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MFS: An Efficient Model Family Serving System for LLMsabstractLLM serving providers typically offer a suite of structurally similar models, known as model families, such as the open-source Llama2 series featuring 7B, 13B, and 70B models. While numerous optimizations for LLM serving have been proposed, the potential for leveraging synergies between models within the same family has not been thoroughly explored. This paper introduces MFS, an innovative multi-tiered LLM model family serving system to exploit the structural similarities and parameter redundancies across different scales of models within a family. By utilizing a novel fine-tuning technique called Knowledge Precipitation, MFS restructures the largest model in a family to encapsulate smaller models within its architecture, enabling a unified multi-tiered serving pipeline. Based on the multi-tiered model, MFS realizes a highly parallelized tiered-level batching approach, significantly enhancing system efficiency. It also enables the sharing of intermediate features and KV-cache between models and facilitates multi-level sampling techniques during the inference phase. Experimental results demonstrate that MFS achieves substantial improvements over existing methods, including a 56.1% reduction in end-to-end token generation latency and a 47.8% decrease in GPU memory footprint without compromising the quality of generated content. Yunxuan Zhang, Hao Wang 0116, Han Tian, Liu Yang 0008, Xudong Liao, Wenxue Li 0004, Ping Yin, Bowen Liu 0002, Kai Chen 0005 |
EuroSys | 2 |
| 2026 | DSA: Efficient Data-Plane Memory Scheduler for In-Network Aggregation to Accelerate Distributed TrainingabstractTo reduce the traffic volume and accelerate communication in distributed training (DT) jobs, recent works introduce In-Network Aggregation (INA) to move the gradient summation into network programmable switches. However, switch memory is a scarce resource, unable to support massive DT jobs in data centers, and existing INA solutions have not utilized switch memory to the best extent. We propose DSA, an Efficient Data-Plane switch memory Scheduler for in-network Aggregation. DSA introduces preemption to the switch memory management for INA jobs. Furthermore, under packet preemption scenarios, DSA optimizes the selective retransmission mechanism to reduce redundant retransimtting packets to alleviate congestion. In the data plane, DSA allows gradient tensors with high priority to preempt the switch aggregators (basic computation unit in INA) from tensors with low priority, which avoids an aggregator wasting time in idle. In the control plane, DSA devises a priority policy which assigns high priority to gradient tensors that benefit overall job efficiency more, e.g., communication-intensive jobs. We implement the prototype of DSA. The experimental results show that DSA can improve the average job completion time (JCT) by up to 1.35x compared with baseline solutions. Jinbin Hu 0001, Xinming Xu, Hao Wang 0116, Jin Wang 0001, Kai Chen 0005 |
IEEE Trans. Netw. | 3 |
| 2026 | Towards Optimal Communication Scheduling With Automatic Configuration for Distributed DNN TrainingabstractByteScheduler partitions and rearranges tensor transmissions to improve the communication efficiency of distributed Deep Neural Network (DNN) training. The configuration of hyper-parameters (i.e., the partition size and the credit size) is critical to the effectiveness of partitioning and rearrangement. Currently ByteScheduler adopts Bayesian Optimization (BO) to find the optimal configuration for the hyper-parameters beforehand. In practice, however, various runtime factors (such as worker node status and network conditions) change over time, making the statically-determined one-shot configuration result suboptimal for real-world DNN training. To address this problem, in this paper we present a realtime configuration method (called AutoByte) that automatically and timely searches the optimal hyper-parameters as the training systems dynamically change. AutoByte extends the ByteScheduler framework with a meta network, which takes the systems’ runtime statistics as its input, dynamically adjusts the triggering threshold based on system environment characteristics, and outputs predictions for speedups under specific configurations. Evaluation results on various DNN models show that AutoByte can dynamically tune the hyper-parameters with low resource usage, and deliver up to 33.2% higher performance than the best static configuration method on the ByteScheduler framework. Jinbin Hu 0001, Xinming Xu, Hao Wang 0116, Yiqing Ma, Yiming Zhang 0003, Jin Wang 0001, Kai Chen 0005 |
IEEE Trans. Netw. | 3 |
| 2025 | Design and Operation of Shared Machine Learning Clusters on CampusabstractThe rapid advancement of large machine learning (ML) models has driven universities worldwide to invest heavily in GPU clusters. Effectively sharing these resources among multiple users is essential for maximizing both utilization and accessibility. However, managing shared GPU clusters presents significant challenges, ranging from system configuration to fair resource allocation among users. This paper introduces SING, a full-stack solution tailored to simplify shared GPU cluster management. Aimed at addressing the pressing need for efficient resource sharing with limited staffing, SING enhances operational efficiency by reducing maintenance costs and optimizing resource utilization. We provide a comprehensive overview of its four extensible architectural layers, explore the features of each layer, and share insights from real-world deployment, including usage patterns and incident management strategies. As part of our commitment to advancing shared ML cluster management, we open-source SING's resources to support the development and operation of similar systems. Kaiqiang Xu, Decang Sun, Hao Wang 0116, Zhenghang Ren, Xinchen Wan, Xudong Liao, Zilong Wang 0007, Junxue Zhang 0001, Kai Chen 0005 |
ASPLOS (1) | 3 |
| 2025 | Enabling In-Network Acceleration Over the Cloud
Hao Wang 0116, Decang Sun, Jinbin Hu 0001, Kai Chen 0005 |
INFOCOM | 1 |
| 2025 | Towards Optimal Rack-scale μs-level CPU Scheduling through In-Network Workload Shaping
Xudong Liao, Han Tian, Xinchen Wan, Chaoliang Zeng, Hao Wang 0116, Junxue Zhang 0001, Mengyu Ma, Guyue Liu, Kai Chen 0005 |
USENIX ATC | 5 |
| 2024 | AutoPipe: Automatic Configuration of Pipeline Parallelism in Shared GPU ClusterabstractAs training Deep Neural Network (DNN) is time-consuming, people resort to parallelization across multiple accelerators. A plethora of solutions adopt data/model parallelization, but they suffer from frequent weight synchronization overhead or resource under-utilization. Recent work introduces pipeline parallelism to improve the utilization of accelerators, however, most existing pipeline parallelism approaches take a one-shot configuration, while ignoring the fluctuation of available resources, e.g., bandwidth and GPUs. Moreover, the heuristic work partition methods oversimplify the computation and communication process, leading to sub-optimal results. To address this challenge, we present AutoPipe, a self-adaptive pipeline parallelism optimization solution. At its core, AutoPipe introduces a reinforcement learning (RL) based work partitioning model, which takes into account both exact communication procedure and dynamic state switching. To mitigate the stalls on state switching, AutoPipe adopts layer-by-layer computation under switching. We have implemented an AutoPipe prototype and evaluated it via testbed experiments. Our results show that the AutoPipe-enhanced PipeDream can find better work partitioning and benefit from dynamic configuration, outperforming the vanilla solutions by up to 89% for exclusive tasks and 143% in dynamic workloads. Furthermore, we show that AutoPipe can also work well with other pipeline parallelism schemes and achieve considerable performance gains. Jinbin Hu 0001, Ying Liu 0064, Hao Wang 0116, Jin Wang 0001 |
ICPP | 3 |
| 2024 | Towards Domain-Specific Network Transport for Distributed DNN Training
Hao Wang 0116, Han Tian, Jingrong Chen 0004, Xinchen Wan, Jiacheng Xia, Gaoxiong Zeng, Wei Bai 0001, Junchen Jiang, Yong Wang 0046, Kai Chen 0005 |
NSDI | 1 |
| 2024 | Accelerating Neural Recommendation Training with Embedding Scheduling
Chaoliang Zeng, Xudong Liao, Xiaodian Cheng, Han Tian, Xinchen Wan, Hao Wang 0116, Kai Chen 0005 |
NSDI | 6 |
| 2022 | Herald: An Embedding Scheduler for Distributed Embedding Model TrainingabstractGiven the ability to represent categorical features, embedding models have gained great success on many internet services. State-of-the-art training frameworks enable embedding cache in GPU workers to benefit from hardware acceleration while supporting massive category representations (embeddings) in the limited-capacity GPU device memory. However, based on our measurements, naively adopting a cache system in embedding model training leads to non-negligible communications overhead between caches and the global parameter server. We observe that many such communications are avoidable, given the predictability and sparsity natures of embedding cache accesses in distributed training. Chaoliang Zeng, Xiaodian Cheng, Han Tian, Hao Wang 0116, Kai Chen 0005 |
APNet | 4 |
| 2022 | AutoByte: Automatic Configuration for Optimal Communication Scheduling in DNN TrainingabstractByteScheduler partitions and rearranges tensor transmissions to improve the communication efficiency of distributed Deep Neural Network (DNN) training. The configuration of hyper-parameters (i.e., the partition size and the credit size) is critical to the effectiveness of partitioning and rearrangement. Currently, ByteScheduler adopts Bayesian Optimization (BO) to find the optimal configuration for the hyper-parameters beforehand. In practice, however, various runtime factors (e.g., worker node status and network conditions) change over time, making the statically-determined one-shot configuration result suboptimal for real-world DNN training.To address this problem, we present a real-time configuration method (called AutoByte) that automatically and timely searches the optimal hyper-parameters as the training systems dynamically change. AutoByte extends the ByteScheduler framework with a meta-network, which takes the system’s runtime statistics as its input and outputs predictions for speedups under specific configurations. Evaluation results on various DNN models show that AutoByte can dynamically tune the hyper-parameters with low resource usage, and deliver up to 33.2% higher performance than the best static configuration in ByteScheduler. Yiqing Ma, Hao Wang 0116, Yiming Zhang 0003, Kai Chen 0005 |
INFOCOM | 2 |
| 2020 | RAT - Resilient Allreduce Tree for Distributed Machine LearningabstractParameter/gradient exchange plays an important role in large-scale distributed machine learning (DML). However, prior solutions such as parameter server (PS) or ring-allreduce (Ring) fall short since they are not resilient to issues or uncertainties like oversubscription, congestion or failures that may occur in datacenter networks (DCN). Xinchen Wan, Hong Zhang 0025, Hao Wang 0116, Shuihai Hu, Junxue Zhang 0001, Kai Chen 0005 |
APNet | 3 |