EDBT 2026 Demo / reviewers in the wild / expert
Zhuolong Yu
dblp:186/9734
· DBLP profile ↗
23ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-8846-5229ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 20 · 4 first-author · 13 since 2021Systems, architecture and hardware · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking Cloud Optimization: Volatility-Driven for Better OutcomesabstractCloud providers commonly employ oversubscription strategies to maximize profitability, leveraging the significant gap between the resources purchased by tenants and those actually consumed by their workloads. However, the temporal volatility of workloads may lead to overload on oversubscribed nodes. To address this issue, existing works typically focus on designing reactive rescheduling mechanisms triggered by overload events or adopt conservative oversubscription strategies to mitigate overload risks. Nonetheless, these solutions compromise either tenant experience or provider profitability. In fact, reducing the temporal volatility of workloads is key to addressing the above challenges. We observe that many workloads exhibit temporal complementarity. Aggregating such workloads can effectively mitigate temporal volatility, thereby improving overall resource utilization. Motivated by this insight, we first design a new metric, called Maximum-based Coefficient of Variation (MCV), to quantify the temporal volatility of workloads. We then propose Hestia, a framework that achieves long-term stable oversubscription through workload aggregation. Specifically, we propose a smoothing-based method to classify workloads suitable for aggregation according to their periodicity. Subsequently, we design an aggregation algorithm to minimize the overall MCV, and treat the aggregated workloads as the units for oversubscription. Experimental results show that, using CPU as a representative example, Hestia reduces MCV by 43.3% and increases oversubscription profit by 66.74%. Baoqing Wang, Gongming Zhao, Hongli Xu 0001, Shibo Wu, Zhuolong Yu, Jiawei Liu 0007, Junhong Lu, Shaohui Xu, Fanjie Meng |
SIGCOMM | 5 |
| 2025 | Towards High-Performance and Compatible RDMA Networks with Receiver-Based and Fine-Grained Congestion Control
Jianchun Liu, Hongli Xu 0001, Yangming Zhao, Zhuolong Yu |
ICCCN | 5 |
| 2025 | Uno: A One-Stop Solution for Inter- and Intra-Data Center Congestion Control and Reliable ConnectivityabstractCloud computing and AI workloads are driving unprecedented demand for efficient communication within and across datacenters. However, the coexistence of intra- and inter-datacenter traffic within datacenters plus the disparity between the RTTs of intra- and inter-datacenter networks complicates congestion management and traffic routing. Particularly, faster congestion responses of intra-datacenter traffic causes rate unfairness when competing with slower inter-datacenter flows. Additionally, inter-datacenter messages suffer from slow loss recovery and, thus, require reliability. Existing solutions overlook these challenges and handle inter- and intra-datacenter congestion with separate control loops or at different granularities. We propose Uno, a unified system for both inter- and intra-DC environments that integrates a transport protocol for rapid congestion reaction and fair rate control with a load balancing scheme that combines erasure coding and adaptive routing. Our findings show that Uno significantly improves the completion times of both inter- and intra-DC flows compared to state-of-the-art methods such as Gemini. Tommaso Bonato, Sepehr Abdous, Abdul Kabbani, Ahmad Ghalayini, Nadeen Gebara, Terry Lam, Anup Agarwal, Tiancheng Chen, Zhuolong Yu, Konstantin Taranov, Mahmoud Elhaddad, Daniele De Sensi, Soudeh Ghorbani, Torsten Hoefler |
SC | 9 |
| 2025 | SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA CommunicationabstractRDMA is vital for efficient distributed training across datacenters, but millisecond-scale latencies complicate the design of its reliability layer. We show that depending on long-haul link characteristics, such as drop rate, distance and bandwidth, the widely used Selective Repeat algorithm can be inefficient, warranting alternatives like Erasure Coding. To enable such alternatives on existing hardware, we propose SDR-RDMA, a software-defined reliability stack for RDMA. Its core is a lightweight SDR SDK that extends standard point-to-point RDMA semantics — fundamental to AI networking stacks — with a receive buffer bitmap. SDR bitmap enables partial message completion to let applications implement custom reliability schemes tailored to specific deployments, while preserving zero-copy RDMA benefits. By offloading the SDR backend to NVIDIA’s Data Path Accelerator (DPA), we achieve line-rate performance, enabling efficient inter-datacenter communication and advancing reliability innovation for inter-datacenter training. Mikhail Khalilov, Marcin Chrapek, Tiancheng Chen, Kenji Nakano, Nicola Mazzoletti, Peter-Jan Gootzen, Salvatore Di Girolamo, Rami Nudelman, Gil Bloch, Abdul Kabbani, Sreevatsa Anantharamu, Konstantin Taranov, Zhuolong Yu, Scott Moe, Mahmoud Elhaddad, Torsten Hoefler |
SC | 16 |
| 2024 | Accelerating Distributed Training With Collaborative In-Network AggregationabstractThe surging scale of distributed training (DT) incurs significant communication overhead in datacenters, while a promising solution is in-network aggregation (INA). It leverages programmable switches (e.g., Intel Tofino switches) for gradient aggregation to accelerate DT tasks. Due to switches’ limited on-chip memory size, existing solutions try to design the memory sharing mechanism for INA. This mechanism requires gradients to arrive at switches synchronously, while network dynamics make it common for the asynchronous arrival of gradients, resulting in existing solutions being inefficient (e.g., massive communication overhead). To address this issue, we propose GOAT, the first-of-its-kind work on gradient scheduling with collaborative in-network aggregation, so that switches can efficiently aggregate asynchronously arriving gradients. Specifically, GOAT first partitions the model into a set of sub-models, then decides which sub-model gradients each switch is responsible for aggregating exclusively and to which switch each worker should send its sub-model gradients. To this end, we design an efficient knapsack-based randomized rounding algorithm and formally analyze the approximation performance. We implement GOAT and evaluate its performance on a testbed consisting of 3 Intel Tofino switches and 9 servers. Experimental results show that GOAT can speed up the DT by$1.5 \times $compared to the state-of-the-art solutions. Hongli Xu 0001, Gongming Zhao, Zhuolong Yu, Bingchen Shen, Liguang Xie |
IEEE/ACM Trans. Netw. | 4 |
| 2024 | Achieving Cost Optimization for Tenant Task Placement in Geo-Distributed CloudsabstractCloud infrastructure has gradually displayed a tendency of geographical distribution in order to provide anywhere, anytime connectivity to tenants all over the world. The tenant task placement in geo-distributed clouds comes with three critical and coupled factors:regional diversity in electricity prices,access delay for tenants, andtraffic demand among tasks. However, existing works disregard either the regional difference in electricity prices or the tenant requirements in geo-distributed clouds, resulting in increased operating costs or low user QoS. To bridge the gap, we design a cost optimization framework for tenant task placement in geo-distributed clouds, called TanGo. However, it is non-trivial to achieve an optimization framework while meeting all the tenant requirements. To this end, we first formulate the electricity cost minimization for task placement problem as a constrained mixed-integer non-linear programming problem. We then propose a near-optimal algorithm with a tight approximation ratio$(1-1/e)$using an effective submodular-based method. Results of in-depth simulations based on real-world datasets show the effectiveness of our algorithm as well as the overall 10%-30% reduction in electricity expenses compared to commonly-adopted alternatives. Luyao Luo, Gongming Zhao, Hongli Xu 0001, Zhuolong Yu, Liguang Xie |
IEEE/ACM Trans. Netw. | 4 |
| 2023 | TanGo: A Cost Optimization Framework for Tenant Task Placement in Geo-distributed CloudsabstractCloud infrastructure has gradually displayed a tendency of geographical distribution in order to provide anywhere, anytime connectivity to tenants all over the world. The tenant task placement in geo-distributed clouds comes with three critical and coupled factors: regional diversity in electricity prices, access delay for tenants, and traffic demand among tasks. However, existing works disregard either the regional difference in electricity prices or the tenant requirements in geo-distributed clouds, resulting in increased operating costs or low user QoS. To bridge the gap, we design a cost optimization framework for tenant task placement in geo-distributed clouds, called TanGo. However, it is non-trivial to achieve an optimization framework while meeting all the tenant requirements. To this end, we first formulate the electricity cost minimization for task placement problem as a constrained mixed-integer non-linear programming problem. We then propose a near-optimal algorithm with a tight approximation ratio (1 − 1/e) using an effective submodular-based method. Results of in-depth simulations based on real-world datasets show the effectiveness of our algorithm as well as the overall 10%-30% reduction in electricity expenses compared to commonly-adopted alternatives. Luyao Luo, Gongming Zhao, Hongli Xu 0001, Zhuolong Yu, Liguang Xie |
INFOCOM | 4 |
| 2023 | COIN: Cost-Efficient Traffic Engineering with Various Pricing Schemes in CloudsabstractThe rapid growth of cloud services has brought a significant increase in inter-datacenter traffic. To transfer data among geographically distributed datacenters, cloud providers need to purchase bandwidth from ISPs. The data transferring cost has become one of the major expenses for cloud providers. Therefore, it is essential for a cloud provider to carefully allocate inter-datacenter traffic among the ISPs' links to minimize the costs. Exiting solutions mainly focus on the situations where all links adopt the same pricing scheme. However, in practice, ISPs usually provide multiple pricing schemes for their links due to market competition, which makes the existing solutions nonoptimal. Thus, a new traffic engineering approach that considers various pricing schemes is needed. This paper presents COIN, a new framework for cost-efficient traffic engineering with various pricing schemes. We propose a partition rounding traffic engineering algorithm based on linear independence analysis. The approximation factors and time complexity are formally analyzed. We further conduct large-scale simulations with real- world topologies and datasets. Extensive simulation results show that COIN can save the data transferring cost by up to 54.54% compared with the state-of-the-art solutions. Gongming Zhao, Jingzhou Wang, Hongli Xu 0001, Zhuolong Yu, Chunming Qiao |
INFOCOM | 4 |
| 2023 | GOAT: Gradient Scheduling with Collaborative In-Network Aggregation for Distributed TrainingabstractThe surging scale of distributed training (DT) incurs significant communication overhead in datacenters, while a promising solution is in-network aggregation (INA). It leverages programmable switches (e.g., Intel Tofino switches) for gradient aggregation to accelerate the DT. Due to switches' limited on-chip memory size, existing solutions try to design the memory sharing mechanism for INA. This mechanism requires gradients to arrive at switches synchronously, while network dynamics make it common for the asynchronous arrival of gradients, resulting in existing solutions being inefficient (e.g., massive communication overhead). To address this issue, we propose GOAT, the first-of-its-kind work on gradient scheduling with collaborative in-network aggregation, so that switches can efficiently aggregate asynchronously arriving gradients. Specifically, GOAT first partitions the model into a set of sub-models, then decides which sub-model gradients each switch is responsible for aggregating exclusively and to which switch each worker should send its sub-model gradients. To this end, we design an efficient knapsack-based randomized rounding algorithm and formally analyze the approximation performance. We implement GOAT and evaluate its performance on a testbed consisting of 3 Intel Tofino switches and 9 servers. Experimental results show that GOAT can speed up the DT by 1.5× compared to the state-of-the-art solutions. Gongming Zhao, Hongli Xu 0001, Zhuolong Yu, Bingchen Shen, Liguang Xie |
IWQoS | 4 |
| 2023 | Understanding the Micro-Behaviors of Hardware Offloaded Network Stacks with LuminaabstractHardware offloaded network stacks are widely adopted in modern datacenters to meet the demand for high throughput, ultra-low latency and low CPU overhead. To fully leverage their exceptional performance, users need to have a deep understanding of their behaviors. Despite many efforts on testing software network stacks, hardware network stacks impose unique challenges to testing tools due to their kernel bypass nature and high performance. Zhuolong Yu, Wei Bai 0001, Shachar Raindel, Vladimir Braverman, Xin Jin 0008 |
SIGCOMM | 1 |
| 2023 | GRID: Gradient Routing With In-Network Aggregation for Distributed TrainingabstractAs the scale of distributed training increases, it brings huge communication overhead in clusters. Some works try to reduce the communication cost through gradient compression or communication scheduling. However, these methods either downgrade the training accuracy or do not reduce the total transmission amount. One promising approach, called in-network aggregation, is proposed to mitigate the bandwidth bottleneck in clusters by aggregating gradients in programmable hardware (e.g., Intel Tofino switches). However, existing solutions mainly implement in-network aggregation through fixed (or default) routing paths, resulting in load imbalancing and long communication time. To deal with this issue, we propose GRID, the first-of-its-kind work on Gradient Routing with In-network Aggregation for Distributed Training. In the control plane, we present an efficient gradient routing algorithm based on randomized rounding and formally analyze the approximation performance. In the data plane, we realize in-network aggregation by carefully designing the logic of workers and programmable switches. We implement GRID and evaluate its performance on a small-scale testbed consisting of 3 Intel Tofino switches and 9 commodity servers. With a combination of testbed experiments and large-scale simulations, we show that GRID can reduce the communication time by 38.4%–60.1% and speed up distributed training by 17.4%–52.7% compared with state-of-the-art solutions. Gongming Zhao, Hongli Xu 0001, Changbo Wu, Zhuolong Yu |
IEEE/ACM Trans. Netw. | 5 |
| 2023 | Scalable and Robust East-West Forwarding Framework for Hyperscale CloudsabstractWith the broad deployment of distributed applications on clouds, east-west traffic is now dominating the majority of cloud networks. The existing communication solutions are tightly coupled with either the control plane (e.g., preprogrammed model) or the location of compute nodes (e.g., conventional gateway model). As a result, it is difficult to flexibly respond to the rapidly expanding networks and frequent abnormal events (e.g., burst traffic and device failures). Accordingly, they may not provide high-performance east-west forwarding while ensuring scalability and robustness. To address this issue, we design Zeta, a scalable and robust east-west forwarding framework with gateway clusters for hyperscale clouds. Zeta abstracts the traffic forwarding capability as a Gateway Cluster Layer, decoupled from the logic of control plane and the location of compute nodes. Specifically, Zeta adopts gateway clusters to support large-scale networks and cope with burst traffic. Moreover, a transparent Multi IPs Migration is proposed for fast recovery from unpredictable failures. We implement Zeta based on eXpress Data Path (XDP) and evaluate its scalability and robustness through comprehensive experiments with up to 100k container instances. Our evaluation shows that Zeta reduces the 99% RTT by$5.1 {\times }$in burst video traffic, and reduces the gateway pure recovery delay by$10.8 {\times }$compared with the state-of-the-art solutions. Qianyu Zhang 0001, Gongming Zhao, Liguang Xie, Hongli Xu 0001, Zhuolong Yu, Yangming Zhao, Chunming Qiao, Liusheng Huang |
IEEE/ACM Trans. Netw. | 5 |
| 2022 | Zeta: A Scalable and Robust East-West Communication Framework in Large-Scale Clouds
Qianyu Zhang 0001, Gongming Zhao, Hongli Xu 0001, Zhuolong Yu, Liguang Xie, Yangming Zhao, Chunming Qiao, Liusheng Huang |
NSDI | 4 |
| 2021 | Twenty Years After: Hierarchical Core-Stateless Fair Queueing
Zhuolong Yu, Jingfeng Wu, Vladimir Braverman, Ion Stoica, Xin Jin 0008 |
NSDI | 1 |
| 2021 | Programmable packet scheduling with a single queueabstractProgrammable packet scheduling enables scheduling algorithms to be programmed into the data plane without changing the hardware. Existing proposals either have no hardware implementations for switch ASICs or require multiple strict-priority queues. Zhuolong Yu, Chuheng Hu, Jingfeng Wu, Xiao Sun 0004, Vladimir Braverman, Mosharaf Chowdhury, Zhenhua Liu 0002, Xin Jin 0008 |
SIGCOMM | 1 |
| 2020 | NetLock: Fast, Centralized Lock Management Using Programmable SwitchesabstractLock managers are widely used by distributed systems. Traditional centralized lock managers can easily support policies between multiple users using global knowledge, but they suffer from low performance. In contrast, emerging decentralized approaches are faster but cannot provide flexible policy support. Furthermore, performance in both cases is limited by the server capability. Zhuolong Yu, Yiwen Zhang 0008, Vladimir Braverman, Mosharaf Chowdhury, Xin Jin 0008 |
SIGCOMM | 1 |
| 2019 | QPipe: quantiles sketch fully in the data planeabstractEfficient network management requires collecting a variety of statistics over the packet flows. Monitoring the flows directly in the data plane allows the system to detect anomalies faster. However, monitoring algorithms have to handle a throughput of 109 packets per second and to maintain a very low memory footprint. Widely adopted sampling-based approaches suffer from low accuracy in estimations. Thus, it is natural to ask: "Is it possible to maintain important statistics in the data plane using small memory footprint?". In this paper, we answer this question in affirmative for an important case of quantiles. We introduce QPipe, the first quantiles sketching algorithm that can be implemented entirely in the data plane. Our main technical contribution is an on-the-plane implementation of a variant of SweepKLL [27] algorithm. Specifically, we give novel implementations of argmin(), the major building block of SweepKLL which are usually not supported in the data plane of the commodity switch. We prototype QPipe in P4 and compare its performance with a sampling-based baseline. Our evaluations demonstrate 10× memory reduction for a fixed approximation error and 90× error improvement for a fixed amount of memory. We conclude that QPipe can be an attractive alternative to sampling-based methods. Nikita Ivkin, Zhuolong Yu, Vladimir Braverman, Xin Jin 0008 |
CoNEXT | 2 |
| 2017 | On the effect of flow table size and controller capacity on SDN network throughputabstractSoftware Defined Network (SDN) is an architectural trend in networking towards the use of the centralized controller to get better performance. However, due to limited resources (especially limited flow table size and controller processing capacity), it may result in low-throughput and long-delay for a set of bursty flows. In this paper, we first combine the flow table size constraint and the controller processing capacity constraint to define the Throughput Maximization with Limited Resources (TMLR) problem. Then we prove TMLR is NP-Hard and design an approximation algorithm to solve the TMLR problem. The approximation factor of the proposed algorithm is also analyzed. The simulation results on the SDN platform (Mininet [1]) show that our algorithm can improve the network throughput about 39% on average compared with the existing algorithms. Gongming Zhao, Liusheng Huang, Zhuolong Yu, Hongli Xu 0001, Pengzhan Wang |
ICC | 3 |
| 2017 | Minimizing flow statistics collection cost of SDN using wildcard requestsabstractIn a software defined network (SDN), the control plane needs to frequently collect flow statistics measured at the data plane switches for different applications, such as traffic engineering, flow re-routing, and attack detection. However, existing solutions for flow statistics collection may result in large bandwidth cost in the control channel and long processing delay on switches, which significantly interfere with the basic functions such as packet forwarding and route update. To address this challenge, we propose a Cost-Optimized Flow Statistics Collection (CO-FSC) scheme using wildcard-based requests. We prove that the CO-FSC problem is NP-Hard and present a rounding-based algorithm with an approximation factor f, where f is the maximum number of switches visited by each flow. Moreover, our CO-FSC problem is extended to the general case, in which only a part of flows in a network need to be collected. The extensive simulation results show that the proposed algorithms can reduce the bandwidth overhead by over 41% and switch processing delay by over 45% compared with the existing solutions. Hongli Xu 0001, Zhuolong Yu, Chen Qian 0001, Xiang-Yang Li 0001, Zichun Liu |
INFOCOM | 2 |
| 2017 | Joint Route Selection and Update Scheduling for Low-Latency Update in SDNsabstractDue to flow dynamics, a software defined network (SDN) may need to frequently update its data plane so as to optimize various performance objectives, such as load balancing. Most previous solutions first determine a new route configuration based on the current flow status, and then update the forwarding paths of existing flows. However, due to slow update operations of Ternary Content Addressable Memory-based flow tables, unacceptable update delays may occur, especially in a large or frequently changed network. According to recent studies, most flows have short duration and the workload of the entire network will vary significantly after a long duration. As a result, the new route configuration may be no longer efficient for the workload after the update, if the update duration takes too long. In this paper, we address the real-time route update, which jointly considers the optimization of flow route selection in the control plane and update scheduling in the data plane. We formulate the delay-satisfied route update problem, and prove its NP-hardness. Two algorithms with bounded approximation factors are designed to solve this problem. We implement the proposed methods on our SDN test bed. The experimental results and extensive simulation results show that our method can reduce the route update delay by about 60% compared with previous route update methods while preserving a similar routing performance (with link load ratio increased less than 3%). Hongli Xu 0001, Zhuolong Yu, Xiang-Yang Li 0001, Liusheng Huang, Chen Qian 0001, Taeho Jung |
IEEE/ACM Trans. Netw. | 2 |
| 2017 | Minimizing Flow Statistics Collection Cost Using Wildcard-Based Requests in SDNsabstractIn a software-defined network (SDN), the control plane needs to frequently collect flow statistics measured at the data plane switches for different applications, such as traffic engineering, QoS routing, and attack detection. However, existing solutions for flow statistics collection may result in large bandwidth cost in the control channel and long processing delay on switches, which significantly interfere with the basic functions, such as packet forwarding and route update. To address this challenge, we propose a cost-optimized flow statistics collection (CO-FSC) scheme and a cost-optimized partial flow statistics collection (CO-PFSC) scheme using wildcard-based requests, and prove that both the CO-FSC and CO-PFSC problems are NP-hard. For CO-FSC, we present a rounding-based algorithm with an approximation factor f, where f is the maximum number of switches visited by each flow. For CO-PFSC, we present an approximation algorithm based on randomized rounding for collecting statistics information of a part of flows in a network. Some practical issues are discussed to enhance our algorithms, for example, the applicability of our algorithms. Moreover, we extend CO-FSC to achieve the control link cost optimization FSC problem, and also design an algorithm with an approximation factor f for this problem. We implement our designed flow statistics collection algorithms on the open virtual switch-based SDN platform. The testing and extensive simulation results show that the proposed algorithms can reduce the bandwidth overhead by over 39% and switch processing delay by over 45% compared with the existing solutions. Hongli Xu 0001, Zhuolong Yu, Chen Qian 0001, Xiang-Yang Li 0001, Zichun Liu, Liusheng Huang |
IEEE/ACM Trans. Netw. | 2 |
| 2016 | Real-time update with joint optimization of route selection and update scheduling for SDNsabstractDue to flow dynamics, a software defined network (SDN) may need to frequently update its data plane so as to optimize various performance objectives, such as load balancing. Most previous solutions first determine a new route configuration based on the current flow status, and then update the forwarding paths of existing flows. However, due to slow update operations of Ternary Content Addressable Memory (TCAM) based flow tables, unacceptable update delays may occur, especially in a large or frequently changed network. According to recent studies, most flows have short duration and the workload of the entire network may vary after a long duration. As a result, the new route configuration may be no longer efficient for the workload after the update, if the update duration takes too long. In this paper, we address the real-time route update, which jointly considers the optimization of flow route selection in the control plane and update scheduling in the data plane. We formulate the delay-satisfied route update (DSRU) problem, and prove its NP-Hardness. Two algorithms with bounded approximation factors are designed to solve this problem. We implement the proposed methods on our SDN testbed. The experimental results and extensive simulation results show that our method can reduce the route update delay by about 60% compared with previous route update methods while preserving a similar routing performance (with link load ratio increased less than 3%). Hongli Xu 0001, Zhuolong Yu, Xiang-Yang Li 0001, Chen Qian 0001, Liusheng Huang, Taeho Jung |
ICNP | 2 |
| 2016 | i-Shield: A System to Protect the Security of Your Smartphone
Zhuolong Yu, Liusheng Huang, Hansong Guo, Hongli Xu 0001 |
KSEM | 1 |