Shuyong Zhu

dblp:145/8231 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
13since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 14 · 1 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CoPT: Collaborative Machine-Learning-Based Mechanism for Adaptive Congestion Control Parameter Tuning in Data Center Networks
Ziwen Yang, Ruya Gu, Duokun Xu, Shuyong Zhu, Yanmin Jia, Xiaohe Hu
IWQoS4
2025 Towards Efficient Secure Aggregation based on In-Network Computing
abstract
Privacy-preserving machine learning (PPML) allows collaborative training without exposing raw data but remains susceptible to gradient leakage. Secure aggregation offers pro-tection via masking and secret sharing, yet suffers from high communication overhead. We propose NET-SA, a secure aggregation architecture based on in-network computing. By employing homomorphic pseudorandom generators for local masking and leveraging programmable switches for seed aggregation, NET-SA eliminates key negotiation and secret sharing steps, reducing communication cost and improving dropout tolerance. Experiments on Tofino switches demonstrate up to 77× faster runtime and 2× lower communication overhead compared to existing methods.
Shuyong Zhu, Qingqing Ren, Yujun Zhang 0001
IWQoS2
2025 DCQF: Differentiated Cyclic Queuing and Forwarding in Large-Scale Deterministic Networks
abstract
Emerging time-sensitive applications impose demands on bounded latency and jitter of network transmission, introducing their deterministic scheduling problem. There are mainly two types of methods for it, fixed rate limit based and hyper-period based, which still face challenges in scheduling performance and run time. We also noticed that flows may experience redundant queuing delay at some hops because of current methods' inability to distinguish them with different arrival patterns, increasing their end-to-end latency. To address the above issues, we present a method based on Differentiated Cyclic Queuing and Forwarding for time-sensitive flows' deterministic scheduling in large-scale networks, called DCQF. It provides fast and slow transmission modes at each hop and forwards flows in the earliest feasible sending cycles according to their arrival patterns. Due to employing a relatively simple scheduling model and relaxing flows' constraints by reducing their queuing delay, DCQF demonstrates higher computational efficiency and provides lower latency for flows in experiments, reflecting its feasibility and advantages in solving the deterministic scheduling problem in large-scale networks.
Shuyong Zhu, Yujun Zhang 0001
IWQoS2
2025 NET-SM4: A High-Performance Secure Encryption Mechanism Based on In-Network Computing
abstract
Encryption is crucial for securing critical network infrastructures, including datacenter networks, 5 G networks, and the Internet of Things (IoT). In-network encryption (INE) offers a promising solution by enabling direct encryption of data on the network's data plane during transmission, thereby eliminating the need for host-side hardware encryption. However, existing INE solutions fail to fully leverage the processing capabilities of programmable switches, leading to low throughput, high resource overhead, and limited key flexibility. These limitations hinder their compatibility with other network functions and restrict their real-world deployment. To address these challenges, we introduce NET-SM4, a high-performance secure encryption mechanism based on in-network computing. NET-SM4 offloads the highly secure and pipeline-optimized SM4 encryption algorithm to programmable switches. By employing a hardwarefriendly table lookup approach, NET-SM4 reduces computation dependency chains and supports parallel encryption inherently, thereby achieving high throughput and low resource overhead for in-network encryption. We implement a prototype of NETSM4 on a commercial Tofino switch and evaluate its performance through testbed experiments and a real-world RDMA-based case study. The results demonstrate that NET-SM4 (1) outperforms state-of-the-art in-network encryption solutions in throughput by up to 293.85 %, and (2) ensures link-speed data transmission with less than$20 \mu$s overhead in real-world scenarios.
Shuyong Zhu, Tianyu Zuo, Yujun Zhang 0001
IWQoS2
2025 FastDet: Providing faster deterministic transmission for time-sensitive flows in WAN
Shuyong Zhu
Comput. Networks2
2025 SNS: Smart Node Selection for Scalable Traffic Engineering in Segment Routing Networks
abstract
Segment routing (SR) is an emerging architecture that can benefit traffic engineering (TE). Nowadays, TE in SR networks (SR-TE) is often solved as an optimization problem to optimize network performance such as link utilization. As network size grows rapidly, implementing SR-TE suffers from scalability issues, including long computation time, high control overhead and expensive deployment cost. In this paper, we propose Smart Node Selection (SNS), a scalable SR-TE method with learning-based node selection (NS). NS is a recently proposed technique for reducing computation time of SR-TE. It first selects a subset of nodes as candidate intermediate nodes to route traffic, then builds linear programming (LP) models that can be solved efficiently. However, existing NS methods use simple heuristics and consider only network topology, which may lead to unsatisfying network performance. To address this problem, we for the first time formulates NS as a reinforcement learning task, which learns a selection policy to achieve better trade-offs between TE performance and computation time, considering both topology and traffic. Besides, we extend NS with additional selection policies and a customized training algorithm, making it a unified framework for scalable SR-TE, which reduces not only computation time, but also control overhead and deployment cost. Performance evaluations on various real-world topologies and traffic matrices show that SNS significantly reduces computation time and control overhead of existing LP models while offering good network performance, and can also be used in partially deployed SR networks to reduce deployment cost.
Linghao Wang, Lu Lu 0016, Miao Wang 0007, Shuyong Zhu, Yujun Zhang 0001
IEEE Trans. Netw. Serv. Manag.6
2024 RADD: A Real-Time and Accurate Method for DDoS Detection Based on In-Network Computing
abstract
Distributed Denial-of-Service (DDoS) attacks pose formidable threats to the security and availability of critical Internet infrastructure. In-network computing technology brings new opportunities to address DDoS attacks due to its intrinsic data plane programmability and high performance. However, existing DDoS attacks detection schemes based on in-network computing are difficult to strike a balance between true positive rate and false positive rate, especially in low-rate DDoS attacks scenarios. In response to this challenge, we propose RADD, an entropy-based method to detect DDoS attacks in real time based on in-network computing. RADD measures the distribution of network traffic from the perspective of individual IP address to discern subtle fluctuations within network traffic, hence providing early indications of potential DDoS attacks. We implement a prototype of RADD over programmable switches and results show that our proposed method significantly outperforms the state-of-the-art or has equivalent accuracy in low-rate and highrate DDoS attacks scenarios.
Shuyong Zhu, Lu Lu 0016, Yujun Zhang 0001
ICC2
2024 LoWAR: Enhancing RDMA over Lossy WANs with Transparent Error Correction
abstract
As the increase of geographically distributed applications continues, the demand for high-speed, long-distance data transmission across wide area networks (WANs) has significantly increased. Remote Direct Memory Access (RDMA) is extensively deployed in data center networks (DCNs) for its high throughput, low latency, and reduced CPU utilization, and its extension to WANs is expected to fully leverage these benefits. However, existing RDMA solutions, while demonstrating superior performance in data centers, face a performance gap over WANs due to their reliance on DCNs for optimal performance and lack of optimization for WANs’ high latency and loss rates. To bridge this gap, we introduce Lossy Wide-Area RDMA (LoWAR), a high-goodput, high-reliability RDMA solution for lossy WANs. LoWAR incorporates a forward error correction (FEC) shim layer to protect RDMA messages from packet loss, thus minimizing the inefficiency of retransmissions. It also fully offloads processing to RNICs with minimal computational overhead and storage burden, operating transparently on RNICs without requiring modifications to existing applications and networks. We implement a LoWAR prototype with FPGA and evaluate its performance through testbed experiments. The results demonstrate LoWAR’s enhanced performance in lossy WANs: in WANs with 40ms RTT and 0.001% to 0.01% loss rates, LoWAR increases RDMA goodput by 2.05 to 5.01 times, reduces average flow completion times (FCTs) by 3.5% to 12.2%, and eliminates 99th percentile tail FCTs in most scenarios.
Tianyu Zuo, Tao Sun 0010, Shuyong Zhu, Wenxiao Li 0006, Lu Lu 0016, Zongpeng Du, Yujun Zhang 0001
IWQoS3
2024 HPDNS: A High-Performance Domain Name Service Architecture Based on In-Network Computing
abstract
The Domain Name System (DNS) is a critical infrastructure of the Internet and forms the backbone of many network services. DNS dramatically impairs the user experience if it does not resolve domain names promptly or correctly. As a result, the performance of the DNS are essential. In this paper, We propose HPDNS, a high-performance domain name service architecture based on in-network computing. HPDNS supports responding to the DNS queries in the data plane and dynamically updating hot-spot domain name records in the control plane. We have proposed a data structure Combinatorial Domain Names and an algorithm Segmented Matching, which results in a storage-efficient data plane algorithm. We have developed the prototype of HPDNS based on the programmable switch and evaluated its performance. Results demonstrate that it achieves 75× performance gains for A records and 72× performance gains for AAAA records compared to BIND9.
Shuyong Zhu
LCN2
2023 A Heuristic Online Algorithm for Routing in Large-Scale Deterministic Networks
Shuyong Zhu, Yujun Zhang 0001
APNOMS2
2023 Heuristic Fast Routing in Large-Scale Deterministic Network
abstract
Large-scale Deterministic Network (LDN) is developed to achieve deterministic transmission in large-scale networks, which can provide bounded delay and jitter with the Cycle Specified Queuing and Forwarding (CSQF) and shaping mechanisms. Routing for time-sensitive flows in LDN requires meeting the delay and bandwidth constraints, like the Multi-Constrained Path (MCP) problem. However, there is a discrepancy that the constraints in the MCP problem are definite while in LDN they are uncertain. It is because the flow’s rate can be adjusted by the shaper, which influences the waiting time at the ingress node and reserved bandwidth at the path. We call the routing problem in LDN the Multi Variable Constraints Routing (MVCR) problem. Due to the uncertainty of constraints, existing routing algorithms may encounter overlong runtime or early rejection if applied to the MVCR problem. In this paper, we propose a Heuristic Fast Routing solution to address the MVCR problem in LDN, called HFR-L. Firstly, we establish a pathbook in advance and design a metric for selecting routes from it with taking the variable flow’s rate into consideration. Then, given the possibility of network changes or the absence of feasible paths in pathbook, we design an algorithm to compute routes in real time, which is based on an extended Lagrange Relaxation based Aggregated Cost (LARAC) algorithm and continuously adjusts the flow’s rate to balance the constraints. The experiments show that our HFR-L has excellent routing performance and fast execution speed in both global and online scenarios, confirming its feasibility to be used in LDN.
Shuyong Zhu, Linghao Wang, Wenxiao Li 0006, Yujun Zhang 0001
IPCCC2
2023 NetShield: An in-network architecture against byzantine failures in distributed deep learning
Qingqing Ren, Shuyong Zhu, Lu Lu 0016, Guangyu Zhao, Yujun Zhang 0001
Comput. Networks2
2022 A High-performance FPGA-based Accelerator for Gradient Compression
abstract
Gradient compression technology has attracted much attention in recent years, due to its high effectiveness in alleviating the communication bottleneck of distributed deep learning. However, except for the communication reduction, it also brings in a significant increase of computational overhead, which limits or even eliminates the communication-reduction benefit brought by gradient compression. To solve the high computational overhead problem, we propose an FPGA-based accelerator for gradient compression in this paper. A high-performance and programmable accelerator architecture is developed for accelerating various gradient compression algorithms by offloading compute-intensive compression operations to FPGA. Also, we design and implement the FPGA-based accelerator based on the popular gradient compression algorithm top-k sparsification. Experimental results show that the new accelerator achieves up to hundreds of times faster than the compression algorithm implemented on CPU and GPU. What's more, the stable and controllable performance under different datasets demonstrates that the proposed accelerator is insensitive to data distribution, which is essential for time-sensitive applications.
Qingqing Ren, Shuyong Zhu, Xuying Meng, Yujun Zhang 0001
DCC2
2017 SDPA: Toward a Stateful Data Plane in Software-Defined Networking
abstract
As the prevailing technique of software-defined networking (SDN), open flow introduces significant programmability, granularity, and flexibility for many network applications to effectively manage and process network flows. However, open flow only provides a simple “match-action” paradigm and lacks the functionality of stateful forwarding for the SDN data plane, which limits its ability to support advanced network applications. Heavily relying on SDN controllers for all state maintenance incurs both scalability and performance issues. In this paper, we propose a novel stateful data plane architecture (SDPA) for the SDN data plane. A co-processing unit, forwarding processor (FP), is designed for SDN switches to manage state information through new instructions and state tables. We design and implement an extended open flow protocol to support the communication between the controller and FP. To demonstrate the practicality and feasibility of our approach, we implement both software and hardware prototypes of SDPA switches, and develop a sample network function chain with stateful firewall, domain name system (DNS) reflection defense, and heavy hitter detection applications in one SDPA-based switch. Experimental results show that the SDPA architecture can effectively improve the forwarding efficiency with manageable processing overhead for those applications that need stateful forwarding in SDN-based networks.
Chen Sun 0005, Jun Bi, Haoxian Chen 0001, Hongxin Hu, Zhilong Zheng, Shuyong Zhu, Chenghui Wu
IEEE/ACM Trans. Netw.6
2015 SDPA: Enhancing Stateful Forwarding for Software-Defined Networking
abstract
As the prevailing technique of Software-Defined Networking (SDN), OpenFlow introduces significant programmability, granularity and flexibility for many network applications to effectively manage and process network flows. However, OpenFlow only provides a simple "match-action" paradigm and lacks the function of stateful forwarding for SDN data plane, which limits it to support advanced network applications. Heavily relying on SDN controllers for all state maintenance incurs both scalability and performance issues. In this paper, we propose a novel Stateful Data Plane Architecture (SDPA) for SDN data plane. A co-processing unit, Forwarding Processor (FP), is designed for SDN switches to manage state information through new instructions and state tables. We design and implement an extended OpenFlow protocol to implement the communication between the controller and FP. To demonstrate the practicality and feasibility of our approach, we implement both software and hardware prototypes of SDPA switches, and develop a sample network function chain with stateful firewall, DNS reflection attack defense and NAT applications in one SDPA-based switch. Experimental results show that the SDPA architecture can effectively improve the forwarding efficiency with manageable processing overhead for those applications that need stateful forwarding in SDN-based networks.
Shuyong Zhu, Jun Bi, Chen Sun 0005, Chenhui Wu, Hongxin Hu
ICNP1