VLDB 2026 Research / reviewers in the wild / expert
Shengwen Zhou
dblp:324/0015
· DBLP profile ↗
18ranked-venue papers
2as first author
18since 2021 · last 2026
0000-0002-2801-125XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 8 since 2021Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SPML: A Proactive Congestion Control Scheme for Periodic Flows in Distributed Systems
Shengwen Zhou, Jiawei Huang 0001 |
ICDCS | 1 |
| 2026 | Mathematical Modeling and Hybrid GA Optimization for Multifactory Production-Assembly Flexible Job Shop Scheduling With Adaptive Neighborhood SearchabstractIn response to the multifactory collaborative distributed characteristics of building material equipment during the manufacturing process, this paper studies the two‐stage (production–assembly) distributed assembly flexible job shop scheduling problem (DAFJSP) and constructs a corresponding mathematical model. To address this complex NP‐hard problem, a hybrid genetic algorithm with variable neighborhood search (HGA‐VNS) is proposed. The algorithm integrates multiple crossover and mutation strategies, thereby significantly enhancing global search capabilities. Concurrently, the implementation of four neighborhood operations has been demonstrated to enhance local search efficiency. To validate the effectiveness and superiority of the proposed algorithm, comparative experiments with other scheduling algorithms were conducted. The findings indicate that the proposed algorithm demonstrates notable advantages in addressing the DAFJSP. Shengwen Zhou, Baigang Du, Liangyi Nie |
Int. J. Intell. Syst. | 1 |
| 2026 | FAR: Fast and Accurate Rate Control for Lossless Datacenter NetworksabstractIn recent years, end-to-end congestion control algorithms or flow pausing mechanisms are proposed to achieve high throughput and low latency in datacenter networks. However, prior end-to-end congestion control works without complex signals fail to achieve fast convergence to a stable equilibrium state and effectively handle the transient congestion, while existing flow pausing mechanisms are decoupled from congestion control, which leads to long convergence time after transient states and incomplete queue elimination in equilibrium states. To address these issues, we present FAR, a rate control protocol that combines the advantages of flow pausing and congestion control. At its heart, FAR couples the bandwidth-estimation-based congestion control and the end-to-end flow pausing mechanisms. After flow pausing, FAR quickly explore the available bandwidth with a binary-search probe to achieve high throughput and low latency. Meanwhile, FAR employs a probe staggering mechanism to address the queue oscillation issue in high-concurrency scenarios. We implement the prototype of FAR using DPDK. Extensive evaluation results demonstrate that our protocol achieves accurate bandwidth estimation and reduces the tail flow completion time (FCT) by up to 67% compared with the state-of-the-art designs. Jingling Liu, Shengwen Zhou, Yijun Li 0002, Sitan Li, Wanchun Jiang, Jianxin Wang 0001, Ping Zhong 0002, Jiawei Huang 0001 |
IEEE Trans. Netw. | 2 |
| 2025 | Elastic Scheduling for Mix-Flow in Time-Sensitive NetworkingabstractTime-Sensitive Networking (TSN) is the most promising network infrastructure for various time-critical applications in Industry 4.0. However, industry applications generate a mix of time-triggered (TT) and event-triggered (ET) flows. Scheduling such mix-flows is a key challenge for TSN. Though current TSN scheduling mechanisms commonly provide deterministic transmission for TT flows with stringent latency requirements, they cannot flexibly accommodate ET flows, which are usually generated by emergency events. In this paper, we propose Elastic Backoff (EBO), a systematic solution for scheduling mix-flows in an elastic way. Our key insight is that the network resources should be reasonably allocated for ET flows while minimally impacting TT flows. To this end, we incorporate the elasticity into the TSN scheduling to make resource reservations for ET flows without hurting TT flows. We conduct extensive experiments on both testbeds and simulations. The evaluation results show that, compared with the state-of-the-art designs, EBO improves the schedulability of ET flows by up to 7.5×, while still ensuring the deterministic transmission of TT flows. Jiawei Huang 0001, Shengwen Zhou, Hui Li 0120, Yijun Li 0002, Qile Wang, Jishu Tian, Kengchang Chen |
ICDCS | 4 |
| 2025 | Borrow Counter: A Generic Sketch Framework for Non-uniform Flow Estimation
Jiawei Huang 0001, Sitan Li, Zirong Wei, Shengwen Zhou, Wenlu Zhang, Jin Ye 0003 |
ICNP | 5 |
| 2025 | DACC: Data Augmentation for Learning-based Congestion Control
Jiawei Huang 0001, Yijun Li 0002, Shengwen Zhou, Hui Li 0120, Weihe Li, Jingling Liu, Wanchun Jiang |
INFOCOM | 6 |
| 2025 | Accelerating Distributed Graph Learning by Using Collaborative In-Network Multicast and Aggregation
Jiawei Huang 0001, Yijun Li 0002, Jingling Liu, Junxue Zhang 0001, Hui Li 0120, Shengwen Zhou, Xiaojuan Lu, Qichen Su, Jianxin Wang 0001, Chee-Wei Tan 0001, Yong Cui 0001, Kai Chen 0005 |
USENIX ATC | 8 |
| 2025 | Progress-Aware Transmission Protocol for Efficient In-Network Aggregation in Distributed Machine LearningabstractLarge-scale machine learning typically adopts distributed machine learning (DML) techniques to accelerate model training. Due to the large communication overhead, unfortunately, the phase of gradient aggregation has become the performance bottleneck for data-parallel DML. To reduce traffic volume, several in-network aggregation (INA) transmission protocols are proposed to offload gradient aggregation function into the programmable switches. However, since existing INA transmission protocols use synchronous congestion control mechanism to drive each round of gradient aggregation, the straggling workers lead to long iteration time and significant performance degradation. Besides, we reveal that existing INA solutions cannot provide the fairness performance among multiple jobs with varying number of workers. To solve the above problem, we propose PA-ATP, a progress-aware INA transmission protocol, which adopts the progress-aware asynchronous congestion control. PA-ATP adjusts the sending rate in accordance with the transmission progress, allowing the straggling flow to grab more bandwidth than the leading flow and control the asynchronous degree of straggling job. Moreover, to ensure the fair throughput among multiple jobs, we dynamically adjust the aggregator allocation for each job by tuning the number of hash operations. We use a P4 programmable switch and a kernel-bypass protocol stack to implement PA-ATP. The results of testbed and large-scale NS3 simulations show that PA-ATP reduces training time by up to 62% compared to the state-of-the-art INA transmission protocols. Jiawei Huang 0001, Tao Zhang 0019, Shengwen Zhou, Qile Wang, Yijun Li 0002, Jingling Liu, Wanchun Jiang, Jianxin Wang 0001 |
IEEE Trans. Netw. | 4 |
| 2024 | D2T: Dynamic Dual Threshold Policy of Shared-Memory in Data Center SwitchesabstractNowadays the data center switches employ the on-chip shared buffer to absorb bursts and avoid packet loss during transient congestion. However, as the buffer-per-port-per-Gbps in production data centers decreases, it becomes more challenging to provide efficient buffer management to meet the requirements of heterogeneous traffic. We observe that typical shared buffer management policies have two steps: first, they identify short flows arriving at ports and then allocate more buffer room for these ports. Unfortunately, the lack of isolation between long and short flows leads to increased queue buildup and even packet loss of short flows. To address this limitation, we propose D2T, which uses different queue length thresholds for long and short flows. Specifically, we first design a compact data structure to distinguish between long and short flows. Then when two kinds of flows coexist at the same port, the threshold of long flows will decrease to absorb the bursty short flows. We implement D2T at a P4- programmable switch and large-scale simulations. The results demonstrate that D2T reduces both average and tail flow completion times (FCT) of short flows by up to 29% and 62% compared with the state-of-the-art policies, respectively. Jiawei Huang 0001, Hui Li 0120, Jingling Liu, Wenlu Zhang, Yijun Li 0002, Sitan Li, Shengwen Zhou, Ping Zhong 0002, Jianxin Wang 0001, Wanchun Jiang, Yong Cui 0001 |
ICDCS | 10 |
| 2024 | Coupling Congestion Control and Flow Pausing in Data Center NetworkabstractTo achieve high throughput and low latency for data center applications, there are two broad lines of work: end-to-end congestion control algorithms and flow pausing mechanisms. It is challenging for end-to-end congestion control algorithms without complex signals to achieve fast convergence to a stable equilibrium state while effectively handling the transient congestion. Additionally, flow pausing mechanisms are decoupled from congestion control, which leads to long convergence time after transient state and incomplete queue elimination in equilibrium state. We propose a transport protocol that combines the advantages of flow pausing and congestion control, called FAR. The key idea is coupling the bandwidth-estimation based congestion control and the end-to-end flow pausing mechanisms. FAR quickly explores the available bandwidth with binary-search based packet train probe to achieve high throughput and low latency. Extensive evaluation results demonstrate that our protocol achieves accurate bandwidth estimation and reduces the tail flow completion time (FCT) by up to 67 <?TeX $\%$?> Math 1 compared with the state-of-the-art designs. Jiawei Huang 0001, Shengwen Zhou, Yijun Li 0002, Sitan Li, Wanchun Jiang, Jianxin Wang 0007, Ping Zhong 0002 |
ICPP | 2 |
| 2024 | Gsyn: Reducing Staleness and Communication Waiting via Grouping-based Synchronization for Distributed Deep LearningabstractDistributed deep learning has been widely employed to train deep neural network over large-scale dataset. However, the commonly used parameter server architecture suffers from long synchronization time in data-parallel training. Although the existing solutions are proposed to reduce synchronization overhead by breaking the synchronization barriers or limiting the staleness bound, they inevitably experience low convergence efficiency and long synchronization waiting. To address these problems, we propose Gsyn to reduce both synchronization overhead and staleness. Specifically, Gsyn divides workers into multiple groups. The workers in the same group coordinate with each other using the bulk synchronous parallel scheme to achieve high convergence efficiency, and each group communicates with parameter server asynchronously to reduce the synchronization waiting time, consequently increasing the convergence efficiency. Furthermore, we theoretically analyze the optimal number of groups to achieve a good tradeoff between staleness and synchronization waiting. The evaluation test in the realistic cluster with multiple training tasks demonstrates that Gsyn is beneficial and accelerates distributed training by up to 27% over the state-of-the-art solutions. Yijun Li 0002, Jiawei Huang 0001, Jingling Liu, Shengwen Zhou, Wanchun Jiang, Jianxin Wang 0001 |
INFOCOM | 5 |
| 2024 | Wear indicator construction for rolling bearings based on an enhanced and unsupervised stacked auto-encoder
Wenhui Zeng, Lisha Yu, Zhelin Huang, Shengwen Zhou, Shunsheng Guo, Baigang Du |
Soft Comput. | 5 |
| 2024 | Straggler-Aware Gradient Aggregation for Large-Scale Distributed Deep Learning SystemabstractDeep Neural Network (DNN) is a critical component of a wide range of applications. However, with the rapid growth of the training dataset and model size, communication becomes the bottleneck, resulting in low utilization of computing resources. To accelerate communication, recent works propose to aggregate gradients from multiple workers in the programmable switch to reduce the volume of exchanged data. Unfortunately, since using synchronization transmission to aggregate data, current in-network aggregation designs suffer from the straggler problem, which often occurs in shared clusters due to resource contention. To address this issue, we propose a straggler-aware aggregation transport protocol (SA-ATP), which enables the leading worker to leverage the spare computing and storage resources to help the straggling worker. We implement SA-ATP atop clusters using P4-programmable switches. The evaluation results show that SA-ATP reduces the iteration time by up to 57% and accelerates training by up to$1.8\times $in real-world benchmark models. Yijun Li 0002, Jiawei Huang 0001, Jingling Liu, Shengwen Zhou, Tao Zhang 0019, Wanchun Jiang, Jianxin Wang 0001 |
IEEE/ACM Trans. Netw. | 5 |
| 2023 | MEB: an Efficient and Accurate Multicast using Bloom Filter with Customized Hash FunctionabstractMulticast is widely used to support a huge range of applications with one-to-many or many-to-many communication patterns. However, multicast systems do not scale due to considerable state and communication overheads. Some stateful multicast approaches require maintaining the state of each multicast session at switches, thus incurring large memory overhead. Some stateless ones utilize Bloom filter (BF) to encode multicast tree into the packet header to minimize communication overhead, but potentially suffer from the substantial false positive due to the probabilistic nature of Bloom filter. In this paper, we propose a stateless multicast scheme MEB, which uses Bloom filter to achieve large-scale multicast communication with low error, small overhead and high scalability. Specifically, to control the rate of false positive, MEB elaborately selects the hash functions for Bloom filters when constructing the packet header at the sender side, and makes forwarding decision according to packet header at the switch with negligible overhead. We compare MEB against the state-of-the-art multicast system in large-scale simulations. The test results show that MEB reduces the traffic overhead by up to 70% with small error rate. Jiawei Huang 0001, Qile Wang, Jingling Liu, Shengwen Zhou, Zhidong He |
APNet | 6 |
| 2023 | A2TP: Aggregator-aware In-network Aggregation for Multi-tenant LearningabstractDistributed Machine Learning (DML) techniques are widely used to accelerate the training of large-scale machine learning models. However, during training iterations, gradients need to be frequently aggregated across multiple workers, resulting in communication bottleneck. To reduce the communication overhead of DML, several In-Network Aggregation (INA) protocols are proposed to reduce the volume of aggregation traffic by offloading aggregation functions into switches, thus alleviating network bottlenecks. Nevertheless, these protocols couple the congestion control of in-switch aggregator resources and link bandwidth resources, together with the straggler-oblivious manner in aggregator allocation, leading to low aggregation efficiency. Jiawei Huang 0001, Yijun Li 0002, Aikun Xu, Shengwen Zhou, Jingling Liu, Jianxin Wang 0001 |
EuroSys | 5 |
| 2023 | PA-ATP: Progress-Aware Transmission Protocol for In-Network AggregationabstractLarge-scale machine learning typically adopts distributed machine learning (DML) techniques to accelerate model training. Due to the large communication overhead, unfortu-nately, the phase of gradient aggregation has become the performance bottleneck for DML. To reduce traffic volume, several in-network aggregation (INA) transmission protocols are proposed to offload gradient aggregation function into the programmable switches. However, since existing INA transmission protocols use synchronous congestion control mechanism to drive each round of gradient aggregation, the straggling workers lead to long iteration time and significant performance degradation. To solve the above problem, we propose PA-ATP, a progress-aware INA transmission protocol, which adopts the progress-aware asynchronous congestion control. PA-ATP adjusts the sending rate in accordance with the transmission progress, allowing the straggling flow to grab more bandwidth than the leading flow and control the asynchronous degree of straggling job. We use a P4 programmable switch and a kernel-bypass protocol stack to implement PA-ATP. The results of testbed and large-scale NS3 simulations show that PA-ATP reduces training time by up to 62% compared to the state-of-the-art INA transmission protocols. Jiawei Huang 0001, Tao Zhang 0019, Shengwen Zhou, Qile Wang, Yijun Li 0002, Jingling Liu, Wanchun Jiang, Jianxin Wang 0001 |
ICNP | 4 |
| 2023 | Research on DataOps Capability - Practice and DevelopmentabstractThe rapid advancement of modern techniques, including big data, the Internet, cloud computing, communication technology, and artificial intelligence, has resulted in an increasing number of businesses relying on data to manage and monitor their operations. As a result, there is a substantial need for data. Given this immense demand, the secure, efficient, and high-quality delivery of data has become a critical concern for enterprises. This paper proposes a capability framework for the integration of data development and operations (DataOps) to help enterprises build complete DataOps capabilities. Zheng Yin, Shengwen Zhou, Minghui Tian, Musen Lin, Sida Liu |
TrustCom | 2 |
| 2022 | HSP: Hybrid Synchronous Parallelism for Fast Distributed Deep LearningabstractIn the parameter-server-based distributed deep learning system, the workers simultaneously communicate with the parameter server to refine model parameters, easily resulting in severe network contention. To solve this problem, Asynchronous Parallel (ASP) strategy enables each worker to update the parameter independently without synchronization. However, due to the inconsistency of parameters among workers, ASP experiences accuracy loss and slow convergence. In this paper, we propose Hybrid Synchronous Parallelism (HSP), which mitigates the communication contention without excessive degradation of convergence speed. Specifically, the parameter server sequentially pulls gradients from workers to eliminate network congestion and synchronizes all up-to-date parameters after each iteration. Meanwhile, HSP cautiously lets idle workers to compute with out-of-date weights to maximize the utilizations of computing resources. We provide theoretical analysis of convergence efficiency and implement HSP on popular deep learning (DL) framework. The test results show that HSP improves the convergence speedup of three classical deep learning models by up to 67%. Yijun Li 0002, Jiawei Huang 0001, Shengwen Zhou, Wanchun Jiang, Jianxin Wang 0001 |
ICPP | 4 |