Jingzhou Wang

dblp:13/2548 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 12 · 7 first-author · 12 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PE-MPINN: A parameters-enhanced multiphysics-informed neural network for data assimilation of seepage-consolidation coupling problems in spatially variable soils
Mingyue Sun, Gang Ma 0002, Tongming Qu, Shaoheng Guan, Jiangzhou Mei, Jingzhou Wang
Adv. Eng. Informatics6
2026 AzuRe: Region- and Availability Zone-Aware Task Placement in Geo-Distributed Clouds
abstract
Large-scale cloud providers build geo-distributed regions to provide a wide range of services for global users. Due to the potentialcrash failuresof hardware components, the availability zone (AZ) architecture has been proposed for fault tolerance and service availability. In practice, a region consists of multiple AZs and each AZ runs independently. Considering the huge traffic volume and service diversity, how to properly deploy tasks for cost saving and QoS guarantee becomes a major challenge for cloud providers. However, existing task placement works often ignore the impact of crash failures, severely affecting users’ QoS. And fault-tolerant task placement works usually overlook other aspects of performance. Moreover, the AZ architecture has not been fully utilized in previous works. To bridge the gap, this paper proposes AzuRe, a joint region- and AZ-aware task placement framework to satisfy various tenants’ requirements in a cost-efficient manner. Due to the features of cloud services, we take a two-step approach:region-level task placement and AZ-level sub-task placement. For region-level task placement, we prove its NP-Hardness and propose an algorithm based on submodular theory with an approximation ratio of (1 − 1/e). For AZ-level sub-task placement, we formalize the problem as a multiple knapsacks problem with conflict graphs. Based on threshold rounding, we propose an algorithm with a bi-criteria approximation ratio of (2,O(log |A|)), where |A| is the number of AZs. Extensive simulation experiments on real datasets show that AzuRe improves the resource utilization ratio by up to 129.42% compared to the state-of-the-art solutions and limits the cost while meeting various kinds of region- and AZ-level requirements.
Jingzhou Wang, Yu-e Sun, He Huang 0001
IEEE Trans. Netw.1
2025 Copo: Joint Cost and Performance Optimization for Task Placement in Geo-Distributed Clouds
abstract
To provide a wide range of services for global users, cloud providers tend to build geo-distributed regions all over the world. With the rapid growth of cloud services, massive workloads and inter-region traffic have been introduced to current cloud networks, resulting in huge expenditure. Therefore, it is essential for a cloud provider to carefully place tasks and transfer traffic among regions to minimize the total operating costs. Existing solutions typically focus on optimizing either placement costs (e.g., computing resources, and electricity) or bandwidth costs, and overlook performance metrics, which leads to increased overall operating costs or lower user QoS. To bridge the gap, this paper proposes Copo, a joint cost and performance optimization framework for tenant task placement in geo-distributed clouds. We first formalize the cost optimization problem as an undetermined multi-commodity flow problem which has never been studied before, and propose a graph transformation algorithm to reduce the complexity. Then we combine the cost optimization with the performance optimization as the final framework. The key idea of Copo is leveraging KKT conditions to transfer the bi-level optimization to a single level. To efficiently acquire the joint task placement and traffic transfer decisions, we leverage McCormick Envelope-based relaxation to design a randomized rounding-based approximation algorithm. Extensive experiments based on real-world data show the superior cost-efficiency and performance of Copo compared with state-of-the-art solutions.
Bichen Wang, Jingzhou Wang, Yu-e Sun, He Huang 0001
ICNP2
2025 Client Selection for Multi-Task Federated Learning: A Lyapunov Optimization Approach
abstract
Federated Learning (FL) has recently garnered considerable attention because it allows multiple clients to collaboratively train machine learning models while keeping their local data private. However, most existing studies focus primarily on optimizing a single FL task, overlooking the dynamic nature of systems where tasks may arrive over time. To address this issue, this paper formulates a novel long-term multi-task FL optimization problem, aiming at balancing the learning quality and the penalties incurred from insufficient client participation in each communication round. To mitigate selection bias, a fairness queue is implemented, and a Lyapunov optimization model is developed to enhance both system stability and learning utility. Furthermore, we derive an upper bound for the objective function, reframing the client selection issue as a minimum weight bipartite matching problem within an auxiliary bipartite graph. The regret of the proposed strategy is theoretically analyzed to quantify the performance gap. Finally, extensive simulations on two real datasets demonstrate the effectiveness of the proposed scheme, highlighting its potential for improving both fairness and efficiency in dynamic, multi-task FL environments.
Jingzhou Wang, Xiumin Wang 0005, Weiwei Wu 0001
ICPADS1
2025 Azure: Achieving Fault Tolerance for Availability Zone-Aware Task Placement in Multi Regions
abstract
Large-scale cloud providers build geo-distributed regions to provide a wide range of services for global users. Due to the potential crash failures of hardware components, the availability zone (AZ) architecture has been proposed for fault tolerance. In practice, a region consists of multiple AZs and each$\mathbf{A Z}$runs independently. Considering the huge traffic volume and service diversity, how to properly deploy tasks for cost saving and QoS guarantee becomes a major challenge for cloud providers. However, existing task placement works often ignore the impact of crash failures, severely affecting users' QoS. And fault tolerance task placement works usually overlook other aspects of performance. Moreover, the AZ architecture has not been fully utilized in previous works. To bridge the gap, this paper proposes Azure, a cost-efficient task placement framework for fault tolerance and cost efficiency at the AZ-level. Due to traffic dynamics, we take a two-step approach: regionlevel task placement and AZ-level sub-task placement. For regionlevel task placement, we prove its NP-Hardness and propose an algorithm based on submodular theory with an approximation ratio of ($1-\frac{1}{e}$). For AZ-level sub-task placement, we formalize the problem as a multiple knapsack problem with conflict graph, which has never been studied before. Based on threshold rounding, we propose an algorithm with an approximation ratio of 2. Extensive simulation experiments on real datasets show that Azure improves resource utilization rate by$\mathbf{1 3 \% - 5 3 \%}$and limits the cost while meeting fault tolerance requirements.
Jingzhou Wang, Yu-e Sun, He Huang 0001
IWQoS2
2025 Talbot: Improving Throughput for Traffic Dynamics in Reconfigurable Datacenters
abstract
Current web applications like social networks and video streaming have been generating magnificent traffic volume, along with intensive traffic dynamics, raising challenges to the fundamental infrastructure of web, i.e., datacenters. However, the traditional electric-based architectures are behind the curve due to the fixed topology and the demand-oblivious nature, failing to guarantee the performance of web applications. Motivated by the new traffic pattern, the reconfigurable technologies, like optical circuit switches (OCSes), are a promising choice to further improve throughput by dynamically adjusting connections when facing traffic dynamics. Currently, static and dynamic updating are two main methods for reconfiguring OCSes. Though static methods can acquire a solution with approximation ratio, it takes long running time and recourse. As comparison, dynamic updating can efficiently acquire a solution, yet existing solutions may degrade along with updates. This paper presents Talbot to further improve throughput with both approximation guarantee and low updating time/recourse. We formulate the throughput maximization problem as a fully dynamic k-weight limited matching problem which is$\mathcal{N} \mathcal{P}$-hard, and we further propose an approximation algorithm based on level and lazy update scheme. To evaluate Talbot, simulations are conducted with both real-world and synthetic datasets. Compared with state-of-the-art works, we show the superior performance of Talbot.
Jingzhou Wang, Yu-e Sun, He Huang 0001, Yang Du 0006
IWQoS1
2025 A cost-efficient traffic engineering framework with various pricing schemes in clouds
Jingzhou Wang, Gongming Zhao, Hongli Xu 0001, Chunming Qiao, He Huang 0001
Comput. Networks1
2024 ContPhy: Continuum Physical Concept Learning and Reasoning from Videos
abstract
We introduce the Continuum Physical Dataset (ContPhy), a novel benchmark for assessing machine physical commonsense. ContPhy complements existing physical reasoning benchmarks by encompassing the inference of diverse physical properties, such as mass and density, across various scenarios and predicting corresponding dynamics. We evaluated a range of AI models and found that they still struggle to achieve satisfactory performance on ContPhy, which shows that current AI models still lack physical commonsense for the continuum, especially soft-bodies, and illustrates the value of the proposed dataset. We also introduce an oracle model (ContPRO) that marries the particle-based physical dynamic models with the recent large language models, which enjoy the advantages of both models, precise dynamic predictions, and interpretable reasoning. ContPhy aims to spur progress in perception and reasoning within diverse physical settings, narrowing the divide between human and machine intelligence in understanding the physical world.
Zhicheng Zheng, Xin Yan 0008, Zhenfang Chen, Jingzhou Wang, Qin Zhi Eddie Lim, Josh Tenenbaum, Chuang Gan 0001
ICML4
2024 An Empirical Study on Low GPU Utilization of Deep Learning Jobs
abstract
Deep learning plays a critical role in numerous intelligent software applications. Enterprise developers submit and run deep learning jobs on shared, multi-tenant platforms to efficiently train and test models. These platforms are typically equipped with a large number of graphics processing units (GPUs) to expedite deep learning computations. However, certain jobs exhibit rather low utilization of the allocated GPUs, resulting in substantial resource waste and reduced development productivity. This paper presents a comprehensive empirical study on low GPU utilization of deep learning jobs, based on 400 real jobs (with an average GPU utilization of 50% or less) collected from Microsoft's internal deep learning platform. We discover 706 low-GPU-utilization issues through meticulous examination of job metadata, execution logs, runtime metrics, scripts, and programs. Furthermore, we identify the common root causes and propose corresponding fixes. Our main findings include: (1) Low GPU utilization of deep learning jobs stems from insufficient GPU computations and interruptions caused by non-GPU tasks; (2) Approximately half (46.03%) of the issues are attributed to data operations; (3) 45.18% of the issues are related to deep learning models and manifest during both model training and evaluation stages; (4) Most (84.99%) low-GPU-utilization issues could be fixed with a small number of code/script modifications. Based on the study results, we propose potential research directions that could help developers utilize GPUs better in cloud-based platforms.
Yanjie Gao, Haoxiang Lin, Yoyo Liang, Hongyu Zhang 0002, Jingzhou Wang, Yonghua Zeng, Keli Gui, Jie Tong, Mao Yang 0004
ICSE9
2024 Leaf: Improving QoS for Reconfigurable Datacenters with Multiple Optical Circuit Switches
abstract
Facing the huge volume of traffic and intensive traffic dynamics, the traditional datacenter architectures are behind the curve due to the fixed topology and the demand-oblivious nature. Driven by the traffic pattern, the reconfigurable technologies, like optical circuit switches (OCSes), are a promising alternative to further improve throughput when facing traffic dynamics, thanks to the high bandwidth and low reconfiguration time. However, existing works on OCSes are relatively elementary. Some of previous works only deploy single OCS which cannot adapt to large-scale datacenters, while others either cannot guarantee near-optimality or overlook practical issues like limited fiber capacity. These solutions may cause low throughput and long reconfiguration time, resulting in poor QoS. In this paper, we present Leaf to maximize throughput thus to further improve QoS, by deploying multiple OCSes to carry traffic in a cooperation manner. Furthermore, we also show its efficient reconfiguration ability and near-optimality. The formulated problem is a new k-weight limited matching problem and is proven to be N P-hard, which can be solved by a new proposed approximation algorithm with bounded approximation ratio. To evaluate our proposed solution, simulation experiments are conducted with both real-world and synthetic datasets. Compared with state-of-the-arts works, Leaf can improve throughput by 42.16%−68.92%, and reduce runnning time by 68.87%−78.72%.
Jingzhou Wang, Gongming Zhao, Hongli Xu 0001, Haibo Wang 0004
IWQoS1
2024 Joint Request Updating and Elastic Resource Provisioning With QoS Guarantee in Clouds
abstract
In a commercial cloud, service providers (e.g., video streaming service provider) rent resources from cloud vendors (e.g., Google Cloud Platform) and provide services to cloud users, making a profit from the price gap. Cloud users acquire services by forwarding their requests to corresponding servers. In practice, as a common scenario, traffic dynamics will cause server overload or load-unbalancing. Existing works mainly deal with the problem by two methods: elastic resource provisioning and request updating. Elastic resource provisioning is a fast and agile solution but may cost too much since service providers need to buy extra resources from cloud vendors. Though request updating is a free solution, it will cause a significant delay, resulting in a bad users’ QoS. In this paper, we present a new scheme, called real-time request updating with elastic resource provisioning (TRUST), to help service providers pay less cost with users’ QoS guarantee in clouds. In addition, we propose an efficient algorithm for TRUST with a bounded approximation factor based on progressive-rounding. Both small-scale experiment results and large-scale simulation results show the superior performance of our proposed algorithm compared with state-of-the-art benchmarks.
Gongming Zhao, Jingzhou Wang, Hongli Xu 0001, Yangming Zhao, Xuwei Yang, He Huang 0001
IEEE/ACM Trans. Netw.2
2024 Segmented Entanglement Establishment With All-Optical Switching in Quantum Networks
abstract
There are two conventional methods to establish an entanglement connection in a Quantum Data Networks (QDN). One is to create single-hop entanglement links first and then connect them with quantum swapping, and the other is forwarding one of the entangled photons from one end to the other via all-optical switching at intermediate nodes to directly establish an entanglement connection. The two methods both have pros and cons. Respectively, the former method has a higher success probability of constructing entanglement link, but it would consume more quantum resources. The latter method, however, has a lower success probability to deliver a photon across multiple quantum links with fewer quantum resources. Accordingly, we are expecting to establish significantly more entanglement connections with limited quantum resources by first creating entanglement segments, each spanning multiple quantum link, using all-optical switching, and then connecting them with quantum swapping. In this paper, we design SEE, a Segmented Entanglement Establishment approach that seamlessly integrates quantum swapping and all-optical switching to maximize quantum network throughput. SEE first creates entanglement segments over one or multiple quantum links with all-optical switching, and then connect them with quantum swapping. Accordingly, SEE can theoretically outperform conventional entanglement link-based approaches. Large scale simulations show that SEE can achieve up to 100.00% larger throughput compared with the state-of-the-art entanglement link-based approaches, e.g., Redundant Entanglement Provisioning and Selection (REPS).
Gongming Zhao, Jingzhou Wang, Yangming Zhao, Hongli Xu 0001, Liusheng Huang, Chunming Qiao
IEEE/ACM Trans. Netw.2
2023 COIN: Cost-Efficient Traffic Engineering with Various Pricing Schemes in Clouds
abstract
The rapid growth of cloud services has brought a significant increase in inter-datacenter traffic. To transfer data among geographically distributed datacenters, cloud providers need to purchase bandwidth from ISPs. The data transferring cost has become one of the major expenses for cloud providers. Therefore, it is essential for a cloud provider to carefully allocate inter-datacenter traffic among the ISPs' links to minimize the costs. Exiting solutions mainly focus on the situations where all links adopt the same pricing scheme. However, in practice, ISPs usually provide multiple pricing schemes for their links due to market competition, which makes the existing solutions nonoptimal. Thus, a new traffic engineering approach that considers various pricing schemes is needed. This paper presents COIN, a new framework for cost-efficient traffic engineering with various pricing schemes. We propose a partition rounding traffic engineering algorithm based on linear independence analysis. The approximation factors and time complexity are formally analyzed. We further conduct large-scale simulations with real- world topologies and datasets. Extensive simulation results show that COIN can save the data transferring cost by up to 54.54% compared with the state-of-the-art solutions.
Gongming Zhao, Jingzhou Wang, Hongli Xu 0001, Zhuolong Yu, Chunming Qiao
INFOCOM2
2022 Segmented Entanglement Establishment for Throughput Maximization in Quantum Networks
abstract
There are two conventional methods to establish an entanglement connection in a Quantum Data Network (QDN). One is to create single-hop entanglement links first and then connect them with quantum swapping, and the other is for-warding one of the entangled photons from one end to the other via all-optical switching at intermediate nodes to directly establish an entanglement connection. Since a photon is easy to be lost during a long distance transmission, all existing works are adopting the former method. However, in a room size network, the success probability of delivering a photon across multiple links via all-optical switching is not that low. In addition, with an all-optical switching technique, we can save quantum memory at the intermediate nodes. Accordingly, we are expecting to establish significantly more entanglement connections with limited quantum resources by first creating entanglement segments, each spanning multiple quantum links, using all-optical switching, and then connecting them with quantum swapping.In this paper, we design SEE, a Segmented Entanglement Establishment approach that seamlessly integrates quantum swapping and all-optical switching to maximize quantum network throughput. SEE first creates entanglement segments over one or multiple quantum links with all-optical switching, and then connect them with quantum swapping. It is clear that an entanglement link is only a special entanglement segment. Accordingly, SEE can theoretically outperform conventional entanglement link based approaches. Large scale simulations show that SEE can achieve up to 100.00% larger throughput compared with the state-of-the-art entanglement link based approach, i.e., REPS.
Gongming Zhao, Jingzhou Wang, Yangming Zhao, Hongli Xu 0001, Chunming Qiao
ICDCS2
2022 TRUST: Real-Time Request Updating with Elastic Resource Provisioning in Clouds
abstract
In a commercial cloud, service providers (e.g., video streaming service provider) rent resources from cloud vendors (e.g., Google Cloud Platform) and provide services to cloud users, making a profit from the price gap. Cloud users acquire services by forwarding their requests to corresponding servers. In practice, as a common scenario, traffic dynamics will cause server overload or load-unbalancing. Existing works mainly deal with the problem by two methods: elastic resource provisioning and request updating. Elastic resource provisioning is a fast and agile solution but may cost too much since service providers need to buy extra resources from cloud vendors. Though request updating is a free solution, it will cause a significant delay, resulting in a bad users’ QoS. In this paper, we present a new scheme, called real-time request updating with elastic resource provisioning (TRUST), to help service providers pay less cost with users’ QoS guarantee in clouds. In addition, we propose an efficient algorithm for TRUST with a bounded approximation factor based on randomized rounding. Both small-scale experiment results and large-scale simulation results show the superior performance of our proposed algorithm compared with state-of-the-art benchmarks.
Jingzhou Wang, Gongming Zhao, Hongli Xu 0001, Yangming Zhao, Xuwei Yang, He Huang 0001
INFOCOM1
2022 A Robust Service Mapping Scheme for Multi-Tenant Clouds
abstract
In a multi-tenant cloud, cloud vendors provide services (e.g., elastic load-balancing, virtual private networks) on service nodes for tenants. Thus, the mapping of tenants’ traffic and service nodes is an important issue in multi-tenant clouds. In practice, unreliability of service nodes and uncertainty/dynamics of tenants’ traffic are two critical challenges that affect the tenants’ QoS. However, previous works often ignore the impact of these two challenges, leading to poor system robustness when encountering system accidents. To bridge the gap, this paper studies the problem of robust service mapping in multi-tenant clouds (RSMP). Due to traffic dynamics, we take a two-step approach:service node assignmentandtenant traffic scheduling. For service node assignment, we prove its NP-Hardness and analyze its problem difficulty. Then, we propose an efficient algorithm with bounded approximation factors based on randomized rounding and knapsack. For tenant traffic scheduling, we design an approximation algorithm based on fully polynomial time approximation scheme (FPTAS). The proposed algorithm achieves the approximation factor of 2+$\epsilon $, where$\epsilon $is an arbitrarily small value. Both small-scale experimental results and large-scale simulation results show the superior performance of our proposed algorithms compared with other alternatives.
Jingzhou Wang, Gongming Zhao, Hongli Xu 0001, Yutong Zhai, Qianyu Zhang 0001, He Huang 0001, Yongqiang Yang
IEEE/ACM Trans. Netw.1
2021 Robust Service Mapping in Multi-Tenant Clouds
abstract
In a multi-tenant cloud, cloud vendors provide services (e.g., elastic load-balancing, virtual private networks) on service nodes for tenants. Thus, the mapping of tenants' traffic and service nodes is an important issue in multi-tenant clouds. In practice, unreliability of service nodes and uncertainty/dynamics of tenants' traffic are two critical challenges that affect the tenants' QoS. However, previous works often ignore the impact of these two challenges, leading to poor system robustness when encountering system accidents. To bridge the gap, this paper studies the problem of robust service mapping in multi-tenant clouds (RSMP). Due to traffic dynamics, we take a two-step approach: service node assignment and tenant traffic scheduling. For service node assignment, we prove its NP-Hardness and analyze its problem difficulty. Then, we propose an efficient algorithm with bounded approximation factors based on randomized rounding and knapsack. For tenant traffic scheduling, we design an approximation algorithm based on fully polynomial time approximation scheme (FPTAS). The proposed algorithm achieves the approximation factor of 2+ ε , where ε is an arbitrarily small value. Both small-scale experimental results and large-scale simulation results show the superior performance of our proposed algorithms compared with other alternatives.
Jingzhou Wang, Gongming Zhao, Hongli Xu 0001, He Huang 0001, Luyao Luo, Yongqiang Yang
INFOCOM1