Jiawei Liu 0007

dblp:12/8228-7 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0002-5871-9187ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Rethinking Cloud Optimization: Volatility-Driven for Better Outcomes
abstract
Cloud providers commonly employ oversubscription strategies to maximize profitability, leveraging the significant gap between the resources purchased by tenants and those actually consumed by their workloads. However, the temporal volatility of workloads may lead to overload on oversubscribed nodes. To address this issue, existing works typically focus on designing reactive rescheduling mechanisms triggered by overload events or adopt conservative oversubscription strategies to mitigate overload risks. Nonetheless, these solutions compromise either tenant experience or provider profitability. In fact, reducing the temporal volatility of workloads is key to addressing the above challenges. We observe that many workloads exhibit temporal complementarity. Aggregating such workloads can effectively mitigate temporal volatility, thereby improving overall resource utilization. Motivated by this insight, we first design a new metric, called Maximum-based Coefficient of Variation (MCV), to quantify the temporal volatility of workloads. We then propose Hestia, a framework that achieves long-term stable oversubscription through workload aggregation. Specifically, we propose a smoothing-based method to classify workloads suitable for aggregation according to their periodicity. Subsequently, we design an aggregation algorithm to minimize the overall MCV, and treat the aggregated workloads as the units for oversubscription. Experimental results show that, using CPU as a representative example, Hestia reduces MCV by 43.3% and increases oversubscription profit by 66.74%.
Baoqing Wang, Gongming Zhao, Hongli Xu 0001, Shibo Wu, Zhuolong Yu, Jiawei Liu 0007, Junhong Lu, Shaohui Xu, Fanjie Meng
SIGCOMM6
2026 Meteor: High-Performance Control Message Delivery for Large-Scale Clouds
abstract
Virtual private clouds (VPCs) play a critical role in providing secure and isolated network environments for web services. However, with the growing number and size of VPCs, efficiently delivering control messages from the control plane to the data plane has become a major concern for cloud vendors. Existing end-to-end transmission solutions (e.g., RPC) will result in substantial overhead in the control plane, while message-oriented middleware-based solutions (e.g., message queue) will lead to high data plane overhead. To address this issue, we design Meteor, a high-performance control message delivery system for large-scale clouds. Specifically, Meteor combines an RPC path with a message queue (MQ) path and employs an auto dual-path switching mechanism to minimize the message delivery latency. Additionally, we propose a VPC-based message delivery and filtering scheme for the MQ path to reduce data plane overhead. We also design a delivery robustness guarantee mechanism to ensure the reachability and consistency of control messages. Meteor has been thoroughly tested with up to 100k container instances. Evaluation results show that Meteor decreases the message delivery latency by 48.8% and reduces the overhead by about 50% in real-world scenarios, compared with state-of-the-art solutions.
Gongming Zhao, Baoqing Wang, Min Chen 0033, Hongli Xu 0001, Jiawei Liu 0007, Xuwei Yang, Liguang Xie, Yongqiang Yang
WWW5
2026 Scalable High-Fidelity Cloud Network Validation via Hybrid Architecture
abstract
Ensuring reliable operation of cloud networks is critical for cloud service providers to guarantee quality of service for tenants. A promising solution is to design a high-fidelity cloud network validation platform that proactively validates the correctness of all operations before implementing changes to the production network. However, the tight coupling between physical and virtual networks in the cloud poses challenges to achieving high-fidelity cloud network validation. Existing network validation platforms focus primarily on traditional physical networks, while ignoring virtual network validation. Regrettably, neglecting the combined validation of physical and virtual networks will result in inaccurate evaluations. To bridge this gap, we present HifiCNet, a high-fidelity platform that concurrently validates both physical and virtual networks. HifiCNet designs an orchestrator to elegantly coordinate the interaction between physical and virtual networks in the cloud and innovatively adopts an emulator-simulator hybrid architecture to ensure high fidelity and scalability for cloud network validation. Through extensive evaluation based on real topologies and traffic traces, we show that HifiCNet enables high-fidelity validation of cloud network configurations, services, and exceptions. Notably, HifiCNet can leverage 38 servers to establish a physical network comprising 10k hosts, as well as a virtual network consisting of 200k virtual machines.
Jiawei Liu 0007, Ji Qi 0005, Gongming Zhao, Hongli Xu 0001, Baoqing Wang, Chun-Jen Chung, Xuwei Yang
IEEE Trans. Computers1
2026 Accelerating Distributed Training Through In-Network Aggregation and Route Selection
Hongli Xu 0001, Baoqing Wang, Jiawei Liu 0007, Gongming Zhao, Junhong Lu, Chunming Qiao
IEEE Trans. Computers3
2025 Cloud Overbooking Optimization: Reducing Temporal Volatility through Spatial Workload Aggregation
Baoqing Wang, Jiawei Liu 0007, Gongming Zhao, Hongli Xu 0001
APNet4
2025 Fossil: A Cost-Effective and Fault-Tolerant Task Placement Scheme for Geo-Distributed Clouds
Gongming Zhao, Baoqing Wang, Jiawei Liu 0007, Hongli Xu 0001, Gangyi Luo
ICA3PP (2)4
2025 Network Resource-Aware Multi-Job Deployment in Deep Learning Clusters
Ai Zhong, Gongming Zhao, Hongli Xu 0001, Jiawei Liu 0007, Peng Yang 0022
ICCCN5
2025 Multi-Tenant Deployment with Anomaly Isolation in Public Clouds
abstract
Cloud vendors provide network services to tenants through shared service nodes, which may cause the abnormal traffic of one tenant to affect others. Deploying auxiliary systems such as firewalls will reduce the frequency of abnormal traffic occurrences but cannot eliminate them entirely. In practice, proper tenant deployment is a promising method to pursue anomaly isolation. Previous works have explored solutions along this line, such as controlling the impact scope of abnormal traffic to mitigate the influence of anomalies among tenants. However, these solutions cannot ensure full anomaly isolation among all tenants. That is, an anomaly in one tenant may cause complete service disruption for another tenant. To bridge the gap, we study the problem of multi-tenant Deployment with Anomaly Isolation (DAI), which is NP-hard. To address this problem, this paper introduces R-DAI, a rounding-based algorithm that can provide a tenant deployment solution in polynomial time, ensuring anomaly isolation among all tenants and load balancing. We implement our proposed algorithm on a large-scale simulation, and the results demonstrate its superior performance. For example, our algorithm eliminates tenant service disruptions caused by abnormal traffic and reduces the impact scope of a service node failure by 53% compared with other alternatives.
Baoqing Wang, Jiawei Liu 0007, Gongming Zhao, Hongli Xu 0001
IWQoS2
2024 Toward a QoS-Guaranteed Cloud Through Elastic Resource Scaling and Request Updating
Bingchen Shen, Jiawei Liu 0007, Gongming Zhao, Hongli Xu 0001, Jianfeng Bao
ICA3PP (3)2
2024 HifiCNet: High-Fidelity Cloud Network Validation Platform at Scale by Hybrid Architecture
abstract
Ensuring reliable operation of cloud networks is critical for cloud service providers to guarantee quality of service for tenants. A promising solution is to design a high-fidelity cloud network validation platform that proactively validates the correctness of all operations before implementing changes to the production network. However, the tight coupling between physical and virtual networks in the cloud poses challenges to achieving high-fidelity cloud network validation. Existing network validation platforms focus primarily on traditional physical networks, while ignoring virtual network validation. Regrettably, neglecting the combined validation of physical and virtual networks will result in inaccurate evaluations. To bridge this gap, we present HifiCNet, a high-fidelity platform that concurrently validates both physical and virtual networks. HifiCNet designs an orchestrator to elegantly coordinate the interaction between physical and virtual networks in the cloud and innovatively adopts an emulator-simulator hybrid architecture to ensure high fidelity and scalability for cloud network validation. Through extensive evaluation based on real topologies and traffic traces, we show that HifiCNet enables high-fidelity validation of cloud network configurations, services, and exceptions. Notably, HifiCNet can use 38 servers to establish a physical network comprising 10k hosts, and a virtual network consisting of 200 k virtual machines.
Jiawei Liu 0007, Gongming Zhao, Hongli Xu 0001, Baoqing Wang, Peng Yang 0022, Chun-Jen Chung, Min Chen 0033, Xuwei Yang
ICNP1
2024 InArt: In-Network Aggregation with Route Selection for Accelerating Distributed Training
abstract
Deep learning has brought about a revolutionary transformation in network applications, particularly in domains like e-commerce and online advertising. Distributed training (DT), as a critical means to expedite model training, has progressively emerged as a key foundational infrastructure for such applications. However, with the rapid advancement of hardware accelerators, the performance bottleneck in DT has shifted from computation to communication. In-network aggregation (INA) solutions have shown promise in alleviating the communication bottleneck. Regrettably, current INA solutions primarily focus on improving efficiency under the traditional parameter server (PS) architecture and do not fully address the communication bottleneck caused by limited PS ingress bandwidth. To bridge this gap, we propose InArt, the first work to introduce INA with routing selection in a multi-PS architecture. InArt employs a multi-PS architecture to split DT tasks among multiple PSs, and selects appropriate routing schemes to fully harness INA capabilities. To accommodate traffic dynamics, InArt adopts a two-phase approach: splitting the training model among multiple parameter servers and selecting routing paths for INA. We propose Lagrange multiplier and randomized rounding algorithms for these phases, respectively. We implement InArt and evaluate its performance through experiments on physical platforms (Tofino switches) and Mininet emulation (P4 Software Switches). Experimental results show that InArt can reduce communication time by 48%\!\sim57\!% compared with state-of-the-art solutions.
Jiawei Liu 0007, Yutong Zhai, Gongming Zhao, Hongli Xu 0001
WWW1
2024 Toward a Service Availability-Guaranteed Cloud Through VM Placement
abstract
In a multi-tenant cloud, the cloud service provider (CSP) leases physical resources to tenants in the form of virtual machines (VMs) with an agreed service level agreement (SLA). As the most important indicator of SLA, we should guarantee the service availability of tenants when placing the VMs. However, previous works about VM placement mainly concentrate on optimizing the cloud resource utilization, but only a few works consider the service availability by measuring the hardware availability. In fact, abnormal tenants can make the corresponding service unavailable by launching network attacks. That is, both the hardware availability and the tenant uncertainty will affect the service availability of VMs on physical machines (PMs). Without considering this factor, the CSP may fail to meet the tenant’s SLA requirements, leading to a reduction in revenue. To solve such a problem, this paper considers the service availability in terms of both the hardware availability and the tenant uncertainty, and studies the service availability-guaranteed VM placement in multi-tenant clouds (SAG-VMP) problem. This problem is very challenging since the service availability actually changes with the tenants served on the PM. To address this issue, we propose a two-phase approach: PM assignment and VM placement. The first phase determines the availability of each PM through a long-term tenant-PM mapping algorithm and the second phase places each VM on a PM that meets the service availability requirement based on a primal-dual online algorithm. Two algorithms with bounded approximation factors are proposed for these two phases, respectively. Both small-scale experiment results and large-scale simulation results show the superior performance of our proposed algorithms compared with other alternatives.
Jiawei Liu 0007, Gongming Zhao, Hongli Xu 0001, Peng Yang 0022, Baoqing Wang, Chunming Qiao
IEEE/ACM Trans. Netw.1
2024 ALEPH: Accelerating Distributed Training With eBPF-Based Hierarchical Gradient Aggregation
abstract
Distributed training includes two important operations: gradient transmission and gradient aggregation, which will consume massive bandwidth and computing resources. To achieve efficient distributed training, one must overcome two critical challenges: heterogeneity of bandwidth resources and limitation of computing resources among compute nodes. Existing architectures based on Parameter Server (PS) and All-Reduce (AR) fail to cope with these challenges because the PS will aggregate gradients from all workers and suffers from bandwidth bottlenecks, while AR intends to alleviate bandwidth bottlenecks at the PS, but the workers need to process many gradient packets thus can be overloaded. To address these shortcomings, we design a new distributed training system called ALEPH. In the control plane, ALEPH uses an efficient algorithm to group workers into clusters with different sizes so as to fully utilize heterogeneous bandwidth. We show that the proposed algorithm can achieve a good approximation performance. In the data plane, ALEPH leverages, for the first time, extended Berkeley Packet Filter (eBPF) programs to aggregate and forward gradient packets to reduce computation overhead. We show how to overcome several hurdles in using eBPF for distributed training. We implement ALEPH and evaluate its performance on a small-scale testbed and large-scale simulations. Experimental results show that ALEPH reduces training time by 20%-31% and increases bandwidth utilization by 88% compared with state-of-the-art frameworks.
Peng Yang 0022, Hongli Xu 0001, Gongming Zhao, Qianyu Zhang 0001, Jiawei Liu 0007, Chunming Qiao
IEEE/ACM Trans. Netw.5
2023 A Reliability and Robustness-driven Approach for Optimizing VM Placement in Clouds
abstract
Cloud computing plays an increasingly vital role in both commercial and personal services. In multi-tenant clouds, cloud providers encounter challenges such as physical machine failures and malicious tenant attacks. Ensuring the reliability and robustness of cloud remain significant and complex challenges for cloud providers to improve quality of service and profitability. Previous works either fail to strike a balance between these aspects or result in resource waste and increased costs. In this paper, we propose an innovative virtual machine placement solution without additional resource overhead. Specifically, we place tenants’ virtual machines (VMs) on physical machines (PMs) that meet the reliability requirements specified in the service level agreement, while also limiting the number of PMs to mitigate the impact of malicious attacks, thereby enhancing the robustness of the cloud. However, the dynamic nature of tenant traffic exacerbates the complexity of the problem. To tackle this challenge, we present KR-OPD, a two-stage algorithm with superior competitive ratios. Through large-scale simulations and small-scale testbed, KR-OPD outperforms existing state-of-the-art solutions. For example, our algorithm reduces the affected range of tenants by 47%-64% and the packet loss rate of PM nodes by over 70% compared with other alternatives.
Yuheng Zhu, Jiawei Liu 0007, Gongming Zhao, Hongli Xu 0001, Huaqing Tu
ICPADS2
2023 Alleviating the Impact of Abnormal Events Through Multi-Constrained VM Placement
abstract
As a simple and low-cost way to obtain enough computing resources, more and more tenants migrate their tasks to the cloud. However, the frequent occurrence of abnormal events (e.g., malicious tenants and node failures) in the cloud will seriously affect the tenants’ QoS. Conventionally, the cloud vendors reduce the frequency of abnormal events by deploying auxiliary systems, which requires additional costs and increases network complexity. Considering that it is an unrealistic expectation to eliminate the occurrence of abnormal events in clouds, this paper proposes a complementary scheme to alleviate the negative impact scope when an abnormal event occurs through multi-constrained VM placement without consuming additional resources. Specifically, when deploying VMs, we limit the number of pods (or service nodes) each tenant can access and the number of tenants hosted by each pod (or service node). However, the multi-dimensional interaction among numerous system parameters and performance/resource considerations makes the problem of multi-constrained VM placement for alleviating the impact of abnormal events very challenging. To solve this problem, we formulate an integer linear programming and propose a rounding-based algorithm with a logarithmic approximation ratio. We implement our proposed algorithm on a physical testbed. The experimental and simulation results show the high efficiency of the proposed algorithm. For example, our algorithm reduces the impact scope of service node failure by 60%, the impact scope of malicious tenants by 40%, and the tenant task makespan by 25% compared with other alternatives.
Gongming Zhao, Jiawei Liu 0007, Yutong Zhai, Hongli Xu 0001, He Huang 0001
IEEE Trans. Parallel Distributed Syst.2
2021 Towards Robust Multi-Tenant Clouds Through Multi-Constrained VM Placement
abstract
More and more tenants (enterprises and personal users) migrate their tasks to clouds since it is a simple and low-cost way to obtain enough computing resources. However, due to potential node failures and malicious tenants, the modern cloud encounters one critical challenge, i.e., robustness. Conventionally, the cloud vendors deploy auxiliary systems to protect the cloud, which requires additional resource cost and increases the network complexity. To enhance the system robustness, this paper proposes a complementary scheme to improve the cloud robustness through efficient VM placement. Specifically, to alleviate the impact of malicious tenants and node failures on the cloud, when deploying VMs, we limit the number of pods (or service nodes) that each tenant can access, and the number of tenants hosted by each pod (or service node). Though there are a lot of works on VM placement, it is very challenging when the robustness issue is taken into consideration. To solve this problem, we formulate an integer linear programming and propose a rounding-based algorithm with a logarithmic approximation ratio. The simulation results show the high efficiency of the proposed algorithm. For example, our algorithm can improve the network throughput by 150% with other alternatives.
Yutong Zhai, Gongming Zhao, Hongli Xu 0001, Yangming Zhao, Jiawei Liu 0007, Xingpeng Fan
IWQoS5