VLDB 2026 Research / reviewers in the wild / expert
Yunzhuo Liu
dblp:276/3196
· DBLP profile ↗
12ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 3 first-author · 7 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Revisiting Flow Control in Node-Centric Datacenter NetworksabstractNode-centric Data Centers (NDCs) are highly flexible, cost-efficient, and failure-resilient, and have gained growing popularity in recent years. However, RDMA technology used in NDC still faces challenges, including high retransmission overhead, Head-of-Line Blocking (HoLB) and deadlock problems. Existing solutions for traditional data centers cannot simultaneously address these issues due to the unique topology and server transmission characteristics of NDC. In this paper, we propose a per-port flow control named PortFC for NDC. PortFC addresses the above problems through the designs of a Pause/Resume control signal, a per-port queue allocation method, an egress-detecting per-port flow control mechanism, and a server-aware queue scheduling method. Our evaluation shows that PortFC is free from retransmission, capable of eliminating HoLB and avoiding deadlocks. PortFC achieves 1.7-8.0 times higher throughput and reduces latency by 11.7%-87.7% compared to the state-of-the-art lossy RDMA based on IRN and the lossless RDMA method based on PFC. In particular, PortFC still demonstrates good performance in a Rail-only NDC with heterogeneous bandwidth domains. Peirui Cao, Rui Ning, Guangyu Zhao, Zhaochen Zhang, Chang Liu 0001, Yunzhuo Liu, Rui Li 0020, Chengyuan Huang, Tao Sun 0010, Guihai Chen, Baochun Li, Chen Tian 0001 |
IEEE Trans. Netw. | 6 |
| 2025 | FastIOV: Fast Startup of Passthrough Network I/O Virtualization for Secure ContainersabstractSingle Root I/O Virtualization (SR-IOV) technology has advanced in recent years and can simultaneously satisfy the network requirements of high data plane performance, high deployment density, and fast startup for applications in traditional containers. However, it falls short with secure containers, which have become the mainstream choice in multi-tenant clouds. SR-IOV requires secure containers to use passthrough I/O for higher data plane performance, which hinders the container startup performance and prevents its usage in time-sensitive tasks like serverless computing. In this paper, we advocate that the startup performance of SR-IOV enabled secure containers can be further boosted, making SR-IOV suitable for building a Container Network Interface (CNI) for secure containers. We first dissect the end-to-end concurrent startup process and identify three key bottlenecks that lead to the slow startup, including Virtual Function I/O device set management, Direct Memory Access memory mapping, and Virtual Function (VF) driver initialization. We then propose a CNI named FastIOV that addresses these bottlenecks through lock decomposition, unnecessary mapping skipping, decoupled zeroing, and asynchronous VF driver initialization. Our evaluation shows that FastIOV reduces the overhead of enabling SR-IOV for secure containers by 96.1%, achieving 65.7% and 75.4% reductions in the average and 99th percentile end-to-end startup time. Yunzhuo Liu, Junchen Guo, Bo Jiang 0003, Yang Song 0031, Rong Wen, Biao Lyu, Shunmin Zhu, Xinbing Wang |
EuroSys | 1 |
| 2025 | PortFC: Designing High-performance Deadlock-free BCube NetworksabstractBCube is a modular data center network.Compared with other topologies, BCube has natural advantages, such as lower deployment costs and stronger failure recovery capabilities.However, RDMA technology used in BCube still faces challenges, including high retransmission overhead, Head-of-Line Blocking (HoLB) and deadlock problems.Existing solutions for traditional data centers cannot simultaneously address these issues due to the unique topology and server transmission characteristics of BCube.In this paper, we propose a per-port flow control named PortFC for BCube.PortFC addresses the above problems through the designs of a Pause/Resume control signal, a per-port queue allocation method, an egress-detecting per-port flow control mechanism, and a serveraware queue scheduling method.Our evaluation shows that PortFC is free from retransmission, capable of eliminating HoLB and avoiding deadlocks.PortFC achieves 1.7-8.0times higher throughput and reduces latency by 11.7%-87.7%compared to the state-of-the-art Peirui Cao, Rui Ning, Zhaochen Zhang, Chang Liu 0001, Rui Li 0020, Yongqi Yang, Yunzhuo Liu, Chengyuan Huang, Tao Sun 0010, Xiaodong Duan, Guihai Chen, Chen Tian 0001 |
ICS | 8 |
| 2025 | vClos: Network contention aware scheduling for distributed machine learning tasks in multi-tenant GPU clusters
Xinchi Han, Shizhen Zhao, Yongxi Lv, Peirui Cao, Qinwei Yang, Yunzhuo Liu, Shengkai Lin, Bo Jiang 0003, Ximeng Liu, Yong Cui 0001, Chenghu Zhou, Xinbing Wang |
Comput. Networks | 7 |
| 2024 | Understanding Network Startup for Secure Containers in Multi-Tenant Clouds: Performance, Bottleneck and OptimizationabstractIn this paper, we use empirical measurements to show that container network startup is a key factor that contributes to the slow startup of secure containers in multi-tenant clouds, especially in the scenario of serverless computing, where the issue is pronounced by high-volume concurrent container invocations. We conduct extensive and detailed analysis on existing Container Network Interface (CNI) plugins and show that even the fastest one doubles the startup time from the no-network scenario. We show that the major cause of the blowup in total startup time is that enabling networking significantly increases the contention among different startup stages, particularly for global Linux kernel locks, including the Routing Table NetLink (RTNL) mutex lock and various spin locks. We reveal that contending for these locks hinders startup performance in three ways, including directly increasing stage time, causing poor pipeline overlap and wasting CPU resources. To mitigate such kernel lock contention, we propose a multi-stage concurrency control mechanism based on Bayesian optimization to limit the concurrency of each contended stage. Our results show that this lightweight mechanism can effectively reduce the end-to-end container startup time by 18.8% with negligible extra overhead. Yunzhuo Liu, Junchen Guo, Bo Jiang 0003, Xiaoqing Sun, Yang Song 0031, Zhiyuan Hou, Biao Lyu, Rong Wen, Shunmin Zhu, Xinbing Wang |
IMC | 1 |
| 2024 | CE-NAS: An End-to-End Carbon-Efficient Neural Architecture Search FrameworkabstractThis work presents a novel approach to neural architecture search (NAS) that aims to increase carbon efficiency for the model design process. The proposed framework CE-NAS addresses the key challenge of high carbon cost associated with NAS by exploring the carbon emission variations of energy and energy differences of different NAS algorithms. At the high level, CE-NAS leverages a reinforcement-learning agent to dynamically adjust GPU resources based on carbon intensity, predicted by a time-series transformer, to balance energy-efficient sampling and energy-intensive evaluation tasks. Furthermore, CE-NAS leverages a recently proposed multi-objective optimizer to effectively reduce the NAS search space. We demonstrate the efficacy of CE-NAS in lowering carbon emissions while achieving SOTA results for both NAS datasets and open-domain NAS tasks. For example, on the HW-NasBench dataset, CE-NAS reduces carbon emissions by up to 7.22X while maintaining a search efficiency comparable to vanilla NAS. For open-domain NAS tasks, CE-NAS achieves SOTA results with 97.35% top-1 accuracy on CIFAR-10 with only 1.68M parameters and a carbon consumption of 38.53 lbs of CO2. On ImageNet, our searched model achieves 80.6% top-1 accuracy with a 0.78 ms TensorRT latency using FP16 on NVIDIA V100, consuming only 909.86 lbs of CO2, making it comparable to other one-shot-based NAS baselines. Our code is available at https://github.com/cake-lab/CE-NAS. Yunzhuo Liu, Bo Jiang 0003, Tian Guo 0001 |
NeurIPS | 2 |
| 2023 | Measuring the Impact of Gradient Accumulation on Cloud-based Distributed TrainingabstractGradient accumulation (GA) is a commonly adopted technique for addressing the GPU memory shortage problem in model training. It reduces memory consumption at the cost of increased computation time. Although widely used, its benefits to model training have not been systematically studied. Our work evaluates and summarizes the benefits of GA, especially in cloud-based distributed training scenarios, where training cost is determined by both execution time and resource consumption. We focus on how GA can be utilized to balance execution time and resource consumption to achieve the lowest bills. Through empirical evaluations on AliCloud platforms, we observe that the total training cost can be reduced by 31.2% on average with a 17.3% increase in training time, when GA is introduced in the large-model and small-bandwidth scenarios with data-parallel training strategies. Besides, taking micro-batch size into optimization can further decrease training time and cost by 21.2% and 24.8% on average, respectively, for hybrid-parallel strategies in large-model and GPU training scenarios. Zimeng Huang, Bo Jiang 0003, Tian Guo 0001, Yunzhuo Liu |
CCGrid | 4 |
| 2023 | Libra: Contention-Aware GPU Thread Allocation for Data Parallel Training in High Speed NetworksabstractOverlapping gradient communication with backward computation is a popular technique to reduce communication cost in the widely adopted data parallel S-SGD training. However, the resource contention between computation and All-Reduce communication in GPU-based training reduces the benefits of overlap. With GPU cluster network evolving from low bandwidth TCP to high speed networks, more GPU resources are required to efficiently utilize the bandwidth, making the contention more noticeable. Existing communication libraries fail to account for such contention when allocating GPU threads and have suboptimal performance. In this paper, we propose to mitigate the contention by balancing the overlapped computation and communication time. We formulate an optimization problem that decides the communication thread allocation to reduce overall backward time. We develop a dynamic programming based near-optimal solution and extend it to co-optimize thread allocation with tensor fusion. We conduct simulated study and real-world experiment using an 8-node GPU cluster with 50Gb RDMA network training four representative DNN models. Results show that our method reduces backward time by 10%-20% compared with Horovod-NCCL, by 6%-13% compared with tensor-fusion-optimization-only methods. Simulation shows that our method achieves the best scalability with a training speedup of 1.2x over the best-performing baseline as we scale up cluster size. Yunzhuo Liu, Bo Jiang 0003, Shizhen Zhao, Tao Lin 0001, Xinbing Wang, Chenghu Zhou |
INFOCOM | 1 |
| 2023 | X-Plane: A High-Throughput Large-Capacity 5G UPFabstractCloud providers, such as AWS and Azure, have started providing 5G services on their cloud infrastructure. In this paper, we present the design and implementation of X-Plane, a system that uses commercial programmable ASICs and DRAM servers on today's cloud infrastructure to implement high-performance 5G User Plane Function (UPF). Building X-Plane is hard because we need to address the following challenges: consistency issues when concurrently accessing UPF state data, slow UE table lookup due to repetitive and numerous Packet Detection Rule (PDR) matching, and the need to handle out-of-order packets from disconnected UEs. X-Plane addresses these challenges by designing three novel technologies: concurrent state data access protocol, fast flow table and paging buffer for handling out-of-order packets. We demonstrate its feasibility and practicality with our implementation on a Tofino-based programmable ASIC. Our evaluation shows that X-Plane can support over ~490Gbps throughput per ASIC pipeline, over 10 million UEs, and finish packet processing within predictable ~4 us on average. Yunzhuo Liu, Hao Nie, Bo Jiang 0003, Yirui Liu 0001, Yidong Yao, Xionglie Wei, Biao Lyu, Chenren Xu, Shunmin Zhu, Xinbing Wang |
MobiCom | 1 |
| 2023 | Threshold-Based Routing-Topology Co-Design for Optical Data CenterabstractDespite the bandwidth scaling limit of electrical switching and the high cost of building Clos data center networks (DCNs), the adoption of optical DCNs is still limited. There are two reasons. First, existing optical DCN designs usually face high deployment complexity. Second, these designs are not full-optical and the performance benefit over the non-blocking Clos DCN is not clear. After exploring the design tradeoffs of the existing optical DCN designs, we propose TROD (ThresholdRouting basedOpticalDatacenter), a low-complexity optical DCN with superior performance than other optical DCNs. There are two novel designs in TROD that contribute to its success. First, TROD performs robust topology optimization based on the recurring traffic patterns and thus does not need to react to every traffic change, which lowers deployment and management complexity. Second, TROD introduces tVLB (threshold-based Valiant Load Balance), which can avoid network congestion as much as possible even under unexpected traffic bursts. We conduct simulation based on both Facebook’s real DCN traces and our synthesized highly bursty DCN traces. TROD reduces flow completion time (FCT) by about 1.15-2.16$\times$compared to Google’s Jupiter DCN, at least 2$\times$compared to other optical DCN designs, and about 2.4-3.2$\times$compared to expander graph DCN. Compared with the non-blocking Clos, TROD reduces the hop count of the majority packets by one, and could even outperform the non-blocking Clos with proper bandwidth over-provision at the optical layer. Note that TROD can be built with commercially available hardware and does not require host modifications. Peirui Cao, Shizhen Zhao, Zhuotao Liu, Mingwei Xu 0001, Min Yee Teh, Yunzhuo Liu, Xinbing Wang, Chenghu Zhou |
IEEE/ACM Trans. Netw. | 7 |
| 2021 | TROD: Evolving From Electrical Data Center to Optical Data CenterabstractDespite the bandwidth scaling limit of electrical switching and the high cost of building Clos data center networks (DCNs), the adoption of optical DCNs is still limited. There are two reasons. First, existing optical DCN designs usually face tremendous deployment complexity. Second, these designs are not full-optical and the performance benefit against the non-blocking Clos DCN is not clear.After exploring the design tradeoffs of the existing optical DCN designs, we propose TROD (Threshold Routing based Optical Datacenter), a low-complexity optical DCN with superior performance than other optical DCNs. There are two novel designs in TROD that contribute to its success. First, TROD performs robust topology optimization based on the recurring traffic patterns and thus does not need to react to every traffic change, which lowers deployment and management complexity. Second, TROD introduces tVLB (threshold-based VLB), which can avoid network congestion as much as possible even under unexpected traffic bursts. We conduct simulation based on both Facebook’s real DCN traces and our synthesized highly bursty DCN traces. TROD reduces flow completion time (FCT) by at least 2× compared with the existing optical DCN designs, and by approximately 2.4-3.2× compared with expander graph DCN. Compared with the non-blocking Clos, TROD reduces the hop count of the majority packets by one, and could even outperform the non-blocking Clos with proper bandwidth over-provision at the optical layer. Note that TROD can be built with commercially available hardware and does not require host modifications. Peirui Cao, Shizhen Zhao, Min Yee Teh, Yunzhuo Liu, Xinbing Wang |
ICNP | 4 |
| 2020 | Grad: Learning for Overhead-aware Adaptive Video Streaming with Scalable Video CodingabstractVideo streaming commonly uses Dynamic Adaptive Streaming over HTTP (DASH) to deliver good Quality of Experience (QoE) to users. Videos used in DASH are predominantly encoded by single-layered video coding such as H.264/AVC. In comparison, multi-layered video coding such as H.264/SVC provides more flexibility for upgrading the quality of buffered video segments and has the potential to further improve QoE. However, there are two challenges for using SVC in DASH: (i) the complexity in designing ABR algorithms; and (ii) the negative impact of SVC's coding overhead. In this work, we propose a deep reinforcement learning method called Grad for designing ABR algorithms that take advantage of the quality upgrade mechanism of SVC. Additionally, we quantify the impact of coding overhead on the achievable QoE of SVC in DASH, and propose jump-enabled hybrid coding (HYBJ) to mitigate the impact. Through emulation, we demonstrate that Grad-HYBJ, an ABR algorithm for HYBJ learned by Grad, outperforms the best performing state-of-the-art ABR algorithm by 17% in QoE. Yunzhuo Liu, Bo Jiang 0003, Tian Guo 0001, Ramesh K. Sitaraman, Don Towsley, Xinbing Wang |
ACM Multimedia | 1 |