EDBT 2026 Demo / reviewers in the wild / expert
Zhuyun Qi
dblp:199/3719
· DBLP profile ↗
13ranked-venue papers
0as first author
11since 2021 · last 2025
0009-0009-0073-3873ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 8 since 2021Systems, architecture and hardware · 2 · 1 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SentinelX: A Lightweight Malicious Traffic Detection System Based on Programmable Switches
Zutao Zhang, Zeyu Luan, Qing Li 0006, Zhuyun Qi, Yong Jiang 0001, Zhenhui Yuan |
INFOCOM | 4 |
| 2025 | DeSync: Proactive Congestion Control via Random Delay Offsets for Large-Scale ML TrainingabstractSynchronization-induced congestion is a critical performance bottleneck in modern distributed machine learning (ML) training, where simultaneous gradient exchanges create bursty traffic patterns. Existing solutions, both reactive and proactive, struggle to balance throughput and latency in the presence of synchronized flows. We propose DeSync, a proactive traffic shaping scheme that introduces structured random delay to de-synchronize communication rounds. Evaluations with DCQCN, HPCC, DCTCP, and TIMELY demonstrate that DeSync significantly improves FCT, job completion times, and congestion metrics, enhancing existing CC mechanisms without specialized hardware. Xingbo Feng, Zhuyun Qi, Yi Wang 0004, Ziyao Huang 0001, Yan Liu 0062, Jiashuo Lin, Chenxi Ling, Weichao Li 0001, Jin Zhang 0001, Jianping Wang 0001 |
IWQoS | 2 |
| 2025 | Effective Phase Alignment: Reducing Queuing Delay in Multi-CQF for Deterministic NetworkingabstractWhile Multi-CQF enables deterministic networking over wide-area networks(WANs) by decoupling transmission and reception, it introduces significant queuing delays. We propose Effective Phase Alignment (EPA), which mitigates queuing delay by adjusting transmission offsets to align the effective phase, defined as the phase difference between the sending and receiving windows. EPA lowers the upper bound of average queuing delay from 2T to 1.5T, and achieves T under perfect alignment. Chenxi Ling, Zhuyun Qi, Shuangping Zhan, Yan Liu 0062, Xingbo Feng, Ruide Cao, Jingbin Feng, Jiashuo Lin, Jian Cheng 0004, Yi Wang 0004 |
IWQoS | 2 |
| 2025 | ReCQF: Enhancing CQF Redundancy with Delay Alignment Scheduling in TSNabstractIntegrating Frame Replication and Elimination for Reliability (FRER) with Cyclic Queuing and Forwarding (CQF) in Time-Sensitive Networks (TSN) encounters redundancy failures and resource reservation inefficiencies due to length disparities across redundant paths. To address these challenges, we propose ReCQF, a Reliability-Enhanced CQF scheduling framework built on Multi-Instance CQF. ReCQF adaptively assigns redundant flows to multiple CQF queue pairs with specific cycles, effectively aligning transmission delays across redundant paths to ensure low delay and inter-path delay differences while significantly reducing resource reservations. Yan Liu 0062, Zhuyun Qi, Xingbo Feng, Shuangping Zhan, Yao Xin, Jiashuo Lin, Chenxi Ling, Ruide Cao, Weichao Li 0001, Yi Wang 0004 |
IWQoS | 2 |
| 2025 | Intelligent In-Network Attack Detection on Programmable Switches With Soterv2abstractTo improve the accuracy of network attack detection, recent work has proposed deep learning (DL) based detectors. Nonetheless, conventional DL-based solutions are computation-intensive and have to be deployed on high-performance x86 servers, which is inefficient for large-scale networks. Unlike x86 servers, current programmable switches (e.g., P4 switches) support a throughput of Tbps and enable programmable logic in networks, indicating a promising alternative. Therefore, we present Soterv2, an intelligent in-network solution deployed on programmable switches. Soterv2 utilizes a two-phase detection manner. In the first phase, we build a P4 program running on the switch's Tofino ASIC to filter malicious packets from the massive traffic. Then, a DL-based inspection is conducted on the switch's CPU, thoroughly detecting the filtered packets. To improve the filtering performance, we propose to embed the rule-based machine learning model, decision tree, in a single match-action table in the P4 program. We also design a lightweight DL model, Branch Convolution Net, running on a multi-core fashion to speed up the thorough detection. Besides, Soterv2 enables the coordination of distributed switches, covering the detection in a large-scale network. Experiments demonstrate that Soterv2 behaves stably in eight network scenarios of different traffic rates (40/100Gbps) and fulfills per-flow detection in 0.03s. Guorui Xie, Qing Li 0006, Chupeng Cui, Ruoyu Li 0003, Lianbo Ma 0004, Zhuyun Qi, Yong Jiang 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2025 | Distributed Multi-Task In-Network Classification on Programmable Switches by Ensemble ModelsabstractOffloading machine learning models for network classification on high-throughput programmable switches is a promising technology, enabling line-speed in-network classification. Existing solutions are centralized, deploying a complete but heavy model on a single switch with limited hardware resources, causing unsatisfactory accuracy, network-wide resource wastage, and non-generic single-task classification. Therefore, we propose In-Forest-M, a general distributed multi-task in-network classification framework. Firstly, we develop a Lightweight Ensemble Generic Optional Model (LEGO), which can be transformed into base models with full functionality. Each switch only needs to deploy lightweight base models rather than complete ensemble models. The significant reduction in resource consumption allows the deployment of larger models with higher accuracy and more models that support diverse tasks. We employ a fine-grained enhancement mechanism to enhance the classification performance of base models. As traffic traverses different switches, In-Forest-M aggregates the classification results of multiple enhanced base models to improve accuracy further. Secondly, we introduce a two-phase resource-aware model allocation strategy that assigns different task-specific enhanced base models to switches under resource constraints and task requirements. To respond to dynamic traffic changes, we design an optimization-driven reinforcement learning algorithm. Moreover, we propose a lightweight update mechanism for flexible model scaling. Comprehensive experiments reveal that, compared with state-of-the-art in-network classification solutions in three real network topologies, In-Forest-M achieves increased accuracy and reduced switch rules while exhibiting great generality in multi-task classification. Qing Li 0006, Jiaye Lin, Guorui Xie, Zhongxu Guan, Zeyu Luan, Zhuyun Qi, Yong Jiang 0001, Zhenhui Yuan |
IEEE Trans. Netw. | 6 |
| 2024 | Poster Abstract: Extending Schedule-Abstraction Graph for Event-Triggered Response-Time AnalysisabstractFor cyber-physical systems, the predictability of their physical behaviors needs to be ensured by the determinism of cyberspace. Response-time analysis (RTA) can theoretically provide this determinism by analyzing the temporal properties of demands. However, the state-space explosion problem makes it challenging to do exact and sustainable RTA for non-preemptive systems where both release jitter and execution time variation exist, particularly when the system has event-triggered (ET) jobs. To address this issue, we propose an ET-enabled RTA based on the schedule-abstraction graph and preliminarily verify its effectiveness and scalability. Ruide Cao, Qinyang He, Yi Wang 0004, Zhuyun Qi |
IPSN | 4 |
| 2023 | Efficient Flow Recording with InheritSketch on Programmable SwitchesabstractSeveral studies have been proposed to deploy the flow recording (i.e., flow size counting and sketching algorithms) on programmable switches for high-speed processing, helping network management tasks like scheduling. Although programmable switches provide a remarkable packet processing speed, they are of compact resources and follow a restrictive pipeline programming. To fit these limitations, current algorithms either sacrifice the recording accuracy or harm the switch throughput. In this paper, we propose InheritSketch for further improvement. InheritSketch utilizes a separation counting fashion, which is memory-efficient for compact switches. It accurately records the more valuable heavy hitters in the large key-value counters (i.e., the primary table), while only sketching non-heavy flows in the small sentinel table. With the recording ongoing, InheritSketch intelligently summarizes the historical recording experience as the basis for flow inheritance. That is, flows with the same IDs as the previous heavy hitters are regarded as new heavy hitters, being recorded in the primary table. To correct some incorrect inheritance, we also propose the flow rebellion, which promotes flows of large sizes but wrongly stored in the sentinel table to the primary table. InheritSketch is also helpful in applications like differentiated scheduling. We compare InheritSketch with six previous recording algorithms on three public traffic datasets, and prototype InheritSketch on a commodity P4 switch. The results demonstrate that InheritSketch reduces the recording errors by at most ∼7×, and that InheritSketch only consumes 10% of hardware resources on the switch. Guorui Xie, Qing Li 0006, Guanglin Duan, Yong Jiang 0001, Zhuyun Qi, Qiaoling Wang |
ICDCS | 5 |
| 2022 | Soter: Deep Learning Enhanced In-Network Attack Detection Based on Programmable SwitchesabstractThough several deep learning (DL) detectors have been proposed for the network attack detection and achieved high accuracy, they are computationally expensive and struggle to satisfy the real-time detection for high-speed networks. Recently, programmable switches exhibit a remarkable throughput efficiency on production networks, indicating a possible deployment of the timely detector. Therefore, we present Soter, a DL enhanced in-network framework for the accurate real-time detection. Soter consists of two phases. One is filtering packets by a rule-based decision tree running on the Tofino ASIC. The other is executing a well-designed lightweight neural network for the thorough inspection of the suspicious packets on the CPU. Experiments on the commodity switch demonstrate that Soter behaves stably in ten network scenarios of different traffic rates and fulfills per-flow detection in 0.03s. Moreover, Soter naturally adapts to the distributed deployment among multiple switches, guaranteeing a higher total throughput for large data centers and cloud networks. Guorui Xie, Qing Li 0006, Chupeng Cui, Peican Zhu, Dan Zhao 0003, Wanxin Shi, Zhuyun Qi, Yong Jiang 0001, Xi Xiao 0001 |
SRDS | 7 |
| 2021 | LAFS: Learning-Based Application-Agnostic Flow Scheduling for DatacentersabstractMany cloud applications in modern datacenters have very demanding latency requirements, making flow completion time (FCT) an important metric for evaluating the network performance. Existing network flow scheduling methods either base on pre-known information or have poor performance. Therefore, we present LAFS, an efficient learning-based flow scheduling approach which minimizes the FCT with estimated information of flows. LAFS combines system call monitoring and learning methods to learn the flow size and implements the Shortest Remaining Processing Time (SRPT) principle with in-network priorities. Moreover, LAFS adopts flowlets to alleviate the packets disorder problem in fine-grained flow scheduling. Our theoretical analysis and extensive simulations show that LAFS is a practical design and significantly outperforms other information-agnostic designs like DCTCP and PIAS under diverse workloads. Feixue Han, Qing Li 0006, Keke Zhu, Jianer Zhou, Yong Jiang 0001, Zhuyun Qi, Fuliang Li |
IPCCC | 6 |
| 2021 | PUFF: A Passive and Universal Learning-based Framework for Intra-domain Failure DetectionabstractThe increasing amount of network devices brings significant improvement to network quality but is inevitably prone to various failures. The frequent occurrence of link failures and node failures in the real-world network, causing packet losses and delay, calls for more accurate and fast detection methods. Existing network failure detection systems focus on probes and end-to-end metrics, but are limited by overhead on bandwidth or storage. Reliance on specific deployment of monitoring systems on devices like hosts also limits the feasibility and compatibility in general network topology, ignoring the potential of transferring monitoring tasks from hosts to switches. In this paper, we propose PUFF, a passive and data-driven network failure detection system based on in-network feature collection in programmable switches and machine learning algorithms. First, PUFF explores the potential use of continuous traffic changes to detect node and link failures instead of end-to-end metrics. Second, PUFF offers a software-based prototype and compares its performance with the latest passive failure detection methods. Evaluation based on simulation on real-world topology shows that PUFF can detect nearly 90% node failures and 80% link failures with less overhead in a shorter time. Lianjin Ye, Qing Li 0006, Xudong Zuo, Jingyu Xiao, Yong Jiang 0001, Zhuyun Qi, Chunsheng Zhu |
IPCCC | 6 |
| 2018 | Reputation-Based Byzantine Fault-Tolerance for Consortium BlockchainabstractThe Practical Byzantine Fault Tolerance algorithm (PBFT)has been highly applied in consortium blockchain systems, however, this kind of consensus algorithm can hardly identify and remove faulty nodes in time, and also vulnerable to many attacks against the primary node of PBFT. The equality of consortium members' discourse rights is inapplicable to some real scenarios where dominating members are likely to have a larger discourse rights in the voting process. To address these problems, this paper presents Reputation-based Byzantine Fault Tolerance (RBFT)algorithm that incorporates a reputation model to evaluate the operations of each node in the consensus process. The faulty nodes will get lower discourse rights in the voting process if any malicious behavior is detected, with their reputation decreased. Furthermore, this paper presents an innovative reputation-based primary change scheme. The node with higher reputation obtains greater opportunities to be a primary to generate new valid blocks, which reduces the security risk of the primary. The experimental results demonstrate that RBFT gains better performance and ensures system security and reliability. Compared with PBFT, it increases the average throughput by 15% and reduces delay by 10%, and the faulty node rate of the system can continue to decrease over time. Kai Lei, Limei Xu, Zhuyun Qi |
ICPADS | 4 |
| 2017 | Statistical Optimal Hash-Based Longest Prefix MatchabstractLongest Prefix Match (LPM) is a basic and important function for current network devices. Hash-based approaches appear to be excellent candidate solutions for LPM with the capability of fast lookup speed and low latency. The number of hash table probes, i.e. the search path of a hash-based LPM algorithm, directly determines the lookup performance. In this paper, we propose Ω-LPM to improve the lookup performance by optimizing the search path of the hash-based LPM. Ω-LPM first reconstructs the forwarding table to support random search [19], then it applies a dynamic programming algorithm to find the shortest search path based on the statistics of the matching probabilities. Ω-LPM concretely reduces the number of hash table probes via searching most of the packets in optimal search paths. Even in the worst case, the upper bound of the average search path of Ω-LPM is 1 + log2(N), here N is the length of the longest prefix in the routing table. The case studies of the name lookup in Named Data Networking and the IP lookup in current Internet demonstrate that Ω-LPM can shorten 61.04% and 86.88% search paths compared with the basic hash-based methods of name lookup [22] and IP lookup [12], respectively, furthermore Ω-LPM reduces 32.3% probes of the name lookup and 73.55% probes of the IP lookup compared with the optimal linear search. The experimental results conducted on extensional name tables and IP tables also show that Ω-LPM has both low memory overhead and excellent scalability. Yi Wang 0004, Zhuyun Qi, Huichen Dai, Hao Wu 0023, Kai Lei, Bin Liu 0001 |
ANCS | 2 |