Guorui Xie

dblp:270/6072 · DBLP profile ↗
← Back
22ranked-venue papers
10as first author
22since 2021 · last 2026
0000-0001-7532-9116ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 11 · 6 first-author · 11 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Security and privacy · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RatioSketch: Towards More Accurate Frequency Estimation in Data Streams via a Lightweight Neural Network
abstract
Sketch-based solutions are widely used to estimate item frequencies in infinite data streams.Traditional hand-crafted sketches face the bottleneck of further eliminating errors because they cannot fully utilize the data stream distribution.Although recent neural sketches represented by MetaSketch and LegoSketch have improved generalization capabilities, they face bottlenecks such as high computational overhead and parameter sensitivity.Meanwhile, they ignore load information, fail to fully utilize the local information in hand-crafted sketches, and do not focus on the frequent items that are usually more important in data streams.In this paper, we propose RatioSketch, a novel lightweight neural network correction framework that synergizes the advantages of hand-crafted sketches and neural sketches in a ``micro-correction'' paradigm.The key idea is to retain the efficient underlying data structure of the hand-crafted sketch and to build a neural correction layer in its output space. We select multiple representative hand-crafted sketches as use cases to study the correction performance of RatioSketch on them.Extensive experimental evaluations on several real-world datasets show that RatioSketch-corrected sketches achieve consistently higher estimation accuracy than their uncorrected counterparts, as well as outperforming neural baselines such as MetaSketch and LegoSketch under identical memory budgets.
Mengbo Wang 0004, Zhuochen Fan, Dayu Wang, Guorui Xie, Qing Li 0006, Zeyu Luan, Yong Jiang 0001, Tong Yang 0003, Mingwei Xu 0001
AAAI4
2026 Defending against Traffic Analysis Attacks with Flexible In-Network Obfuscation
Guorui Xie, Qing Li 0006, Zhenning Shi, Gianni Antichi, Yijia Zhu, Changxing Weng, Sebastiano Miano, Yong Jiang 0001, Mingwei Xu 0001
NSDI1
2025 Reasoning under Uncertainty: Efficient LLM Inference via Unsupervised Confidence Dilution and Convergent Adaptive Sampling
abstract
Zhenning Shi, Yijia Zhu, Yi Xie, Junhan Shi, Guorui Xie, Haotian Zhang, Yong Jiang, Congcong Miao, Qing Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zhenning Shi, Yijia Zhu, Junhan Shi, Guorui Xie, Yong Jiang 0001, Congcong Miao, Qing Li 0006
EMNLP5
2025 Minos : A Lightweight and Dynamic Defense against Traffic Analysis in Programmable Data Planes
Qing Li 0006, Guorui Xie, Dan Zhao 0003, Zhuochen Fan, Lianbo Ma 0004, Yong Jiang 0001
USENIX ATC3
2025 Helios: Learning and Adaptation of Matching Rules for Continual In-Network Malicious Traffic Detection
abstract
Network Intrusion Detection Systems (NIDS) are critical for web security by identifying and blocking malicious traffic. In-network NIDS leverage programmable switches for high-speed traffic processing. However, they are unable to reconcile the fine-grained classification of known classes and the identification of unseen attacks. Moreover, they lack support for incremental updates. In this paper, we propose Helios, an in-network malicious traffic detection system, for continual adaptation in attack-incremental scenarios. First, we design a novel Supervised Mixture Prototypical Learning (SMPL) method combined with clustering initialization to learn prototypes that encapsulate the knowledge, based on the weighted infinity norm distance. SMPL enables known class classification and unseen attack identification through similarity comparison between prototypes and samples. Then, we design boundary calibration and overlap refinement to transform learned prototypes into priority-guided matching rules, ensuring precise and efficient in-network deployment. Additionally, Helios supports incremental prototype learning and rule updates, achieving low-cost hardware reconfiguration. We implement Helios on a Tofino switch and evaluation on three datasets shows that Helios achieves superior performance in classifying known classes (92%+ in ACC and F1) as well as identifying unseen attacks (62% - 98% in TPR). Helios has also reduced resource consumption and reconfiguration time, demonstrating its scalability and efficiency for real-world deployment.
Zhenning Shi, Dan Zhao 0003, Yijia Zhu, Guorui Xie, Qing Li 0006, Yong Jiang 0001
WWW4
2025 DNSGuard: In-Network Defense Against DNS Attacks
abstract
The Domain Name System (DNS) is a growing center of cyber attacks, including both volumetric and non-volumetric attacks. Programmable switches provide a new opportunity for more efficient defense against DNS attacks since they can offer better cost, performance, and flexibility trade-offs compared to traditional defense systems. However, programmable switches have strict limitations on the operations and storage space supported to ensure line-speed packet processing. In this paper, we propose DNSGuard, an intelligent in-network defense framework that can handle volumetric and non-volumetric DNS attacks on programmable switches. We propose a recursive incremental parsing algorithm that can effectively extract variable-length domain names. To achieve real-time and accurate detection against two types of DNS attacks, we design a switch-optimized and resource-efficient algorithm to extract both independent features of each packet and domain-based cumulative features. Then, we propose a multi-phase hybrid model architecture to perform dynamic packet analysis at different time phases of a domain. Further, we design efficient model representation mechanisms to deploy tree-based ensemble models in the data plane. Experimental results show that DNSGuard can defend against diverse DNS attacks at the line rate. In addition, DNSGuard introduces a minimal nanosecond latency to normal traffic in heavily loaded networks.
Guanglin Duan, Qing Li 0006, Dan Zhao 0003, Guorui Xie, Yuan Yang 0001, Zhenhui Yuan, Yong Jiang 0001, Mingwei Xu 0001
IEEE Trans. Dependable Secur. Comput.5
2025 Intelligent In-Network Attack Detection on Programmable Switches With Soterv2
abstract
To improve the accuracy of network attack detection, recent work has proposed deep learning (DL) based detectors. Nonetheless, conventional DL-based solutions are computation-intensive and have to be deployed on high-performance x86 servers, which is inefficient for large-scale networks. Unlike x86 servers, current programmable switches (e.g., P4 switches) support a throughput of Tbps and enable programmable logic in networks, indicating a promising alternative. Therefore, we present Soterv2, an intelligent in-network solution deployed on programmable switches. Soterv2 utilizes a two-phase detection manner. In the first phase, we build a P4 program running on the switch's Tofino ASIC to filter malicious packets from the massive traffic. Then, a DL-based inspection is conducted on the switch's CPU, thoroughly detecting the filtered packets. To improve the filtering performance, we propose to embed the rule-based machine learning model, decision tree, in a single match-action table in the P4 program. We also design a lightweight DL model, Branch Convolution Net, running on a multi-core fashion to speed up the thorough detection. Besides, Soterv2 enables the coordination of distributed switches, covering the detection in a large-scale network. Experiments demonstrate that Soterv2 behaves stably in eight network scenarios of different traffic rates (40/100Gbps) and fulfills per-flow detection in 0.03s.
Guorui Xie, Qing Li 0006, Chupeng Cui, Ruoyu Li 0003, Lianbo Ma 0004, Zhuyun Qi, Yong Jiang 0001
IEEE Trans. Dependable Secur. Comput.1
2025 Distributed Multi-Task In-Network Classification on Programmable Switches by Ensemble Models
abstract
Offloading machine learning models for network classification on high-throughput programmable switches is a promising technology, enabling line-speed in-network classification. Existing solutions are centralized, deploying a complete but heavy model on a single switch with limited hardware resources, causing unsatisfactory accuracy, network-wide resource wastage, and non-generic single-task classification. Therefore, we propose In-Forest-M, a general distributed multi-task in-network classification framework. Firstly, we develop a Lightweight Ensemble Generic Optional Model (LEGO), which can be transformed into base models with full functionality. Each switch only needs to deploy lightweight base models rather than complete ensemble models. The significant reduction in resource consumption allows the deployment of larger models with higher accuracy and more models that support diverse tasks. We employ a fine-grained enhancement mechanism to enhance the classification performance of base models. As traffic traverses different switches, In-Forest-M aggregates the classification results of multiple enhanced base models to improve accuracy further. Secondly, we introduce a two-phase resource-aware model allocation strategy that assigns different task-specific enhanced base models to switches under resource constraints and task requirements. To respond to dynamic traffic changes, we design an optimization-driven reinforcement learning algorithm. Moreover, we propose a lightweight update mechanism for flexible model scaling. Comprehensive experiments reveal that, compared with state-of-the-art in-network classification solutions in three real network topologies, In-Forest-M achieves increased accuracy and reduced switch rules while exhibiting great generality in multi-task classification.
Qing Li 0006, Jiaye Lin, Guorui Xie, Zhongxu Guan, Zeyu Luan, Zhuyun Qi, Yong Jiang 0001, Zhenhui Yuan
IEEE Trans. Netw.3
2024 Proteus: A Difficulty-Aware Deep Learning Framework for Real-Time Malicious Traffic Detection
abstract
Deep learning (DL) has been recently used for malicious traffic detection. However, DL models are often faced with a dilemma between model size and performance: larger models have better accuracy, but suffer from high detection latency, which severely impacts realtime traffic performance, while lightweight models have low detection latencies, but sacrifice accuracy. In this paper, we introduce Proteus, a swift and precise attack detection framework that adaptively adjusts DL models in real-time based on sample detection difficulty. To address diverse detection difficulties in traffic data, we devise a Double Dynamic Convolutional Neural Network (DDCN) with two pivotal modules: the Dynamic Feature Campaign (DFC) and the Tailor Module (TM). DFC enables the model to discern and accentuate the most influential features, while TM autonomously gauges sample difficulty, cropping the overall model. We further design an auxiliary detection module to streamline the detection, especially for network devices like routers lacking GPUs but equipped with multiple CPU cores. Experiments on different network devices show that Proteus completes the detection of each flow within 0.6 ms, and achieves$\mathbf{9 9. 3 4 \%}$detection accuracy, outperforming other solutions.
Chupeng Cui, Qing Li 0006, Guorui Xie, Ruoyu Li 0003, Dan Zhao 0003, Zhenhui Yuan, Yong Jiang 0001
ICNP3
2024 Linc: Enabling Low-Resource in-Network Classification and Incremental Model Update
abstract
Deploying machine learning models within the data plane can facilitate traffic analysis at line rate. Nevertheless, existing schemes suffer from performance degradation when faced with hardware resource constraints. They further cannot accommodate incremental model updates as the underlying traffic distributions change. This paper presents Linc, a novel framework for low-resource and incrementally updatable innetwork classification. To circumvent the complicated P4 program that consumes extensive resources, Linc generalizes explicit deployable classification rules by interpreting a customized neural network. For incremental updates with minimal new training data, we design a divide-and-conquer strategy that decomposes the update task into binary classification tasks within distinct subspaces to simplify the task. We then update the model locally within each subspace to improve accuracy for new categories without compromising the accuracy of existing categories. Experimental results reveal that, compared to state-of-the-art solutions, Linc significantly reduces switch hardware resources by up to$15.6 \times$and achieves up to$11.63 \%$higher classification accuracy. During the incremental model updates, Linc not only improves the accuracy by$26.1 \%$but also decreases the table entries by$6 \times$.
Haolin Yan, Qing Li 0006, Guorui Xie, Gareth Tyson, Yong Jiang 0001
ICNP3
2024 Mitigating Sample Selection Bias with Robust Domain Adaption in Multimedia Recommendation
abstract
Industrial multimedia recommendation systems extensively utilize cascade architectures to deliver personalized content for users, generally consisting of multiple stages like retrieval and ranking. However, retrieval models have long suffered from Sample Selection Bias (SSB) due to the distribution discrepancy between the exposed items used for model training and the candidates (almost unexposed) during inference, affecting recommendation performance. Traditional methods utilize retrieval candidates as augmented training data, indiscriminately treating unexposed data as negative samples, which leads to inaccuracies and noise. Some efforts rely on unbiased datasets, while they are costly to collect and insufficient for industrial models. In this paper, we propose a debiasing framework named DAMCAR, which introduces Domain Adaptation to mitigate SSB in Multimedia CAscade Recommendation systems. Firstly, we sample hard-to-distinguish samples from unexposed data to serve as the target domain, optimizing data quality and resource utilization. Secondly, adversarial domain adaptation is employed to generate pseudo-labels for each sample. To enhance robustness, we utilize Exponential Moving Average (EMA) to create a teacher model that supervises the generation of pseudo-labels via self-distillation. Finally, we obtain a retrieval model that maintains stable performance during inference through a hybrid training mechanism. We conduct offline experiments on two real-world datasets and deploy our approach in the retrieval model of a multimedia video recommendation system for online A/B testing. Comprehensive experimental results demonstrate the effectiveness of DAMCAR in practical applications.
Jiaye Lin, Qing Li 0006, Guorui Xie, Zhongxu Guan, Yong Jiang 0001, Zhong Zhang 0014, Peilin Zhao
ACM Multimedia3
2024 Generating Neural Networks for Diverse Networking Classification Tasks via Hardware-Aware Neural Architecture Search
abstract
Neural networks (NNs) are widely used in classification-based networking analysis to help traffic transmission and system security. However, there are heterogeneous network devices (e.g., switches and routers) in a network. Manually customizing NNs with specific device requirements (e.g., max allowed running latency) can be time-consuming and labor-intensive. Furthermore, the diverse data characteristics of different networking classification tasks add to the burden of NN customization. This paper introduces Loong, a neural architecture search (NAS) based system that automatically generates NNs for various networking tasks and devices. Loong includes a neural operation embedding module, which embeds candidate neural operations into the layer to be designed. Then, the layer-wise training is used to generate a task-specific NN layer by layer. This layer-wise scheme simultaneously trains and selects candidate neural operations using gradient feedback. Finally, only the important operations are selected to form the layer, maximizing accuracy. By incorporating multiple objectives, including deployment memory and running latency of devices, into the training and selection of NNs, Loong is able to customize NNs for heterogeneous network devices. Experiments show that Loong's NNs outperform 13 manual-designed and NAS-based NNs, with a 4.11% improvement in F1-score. Additionally, Loong's NNs achieve faster (7.92X) speeds on commodity devices.
Guorui Xie, Qing Li 0006, Zhenning Shi, Hanbin Fang, Shengpeng Ji, Yong Jiang 0001, Zhenhui Yuan, Lianbo Ma 0004, Mingwei Xu 0001
IEEE Trans. Computers1
2024 DeviceRadar: Online IoT Device Fingerprinting in ISPs Using Programmable Switches
abstract
Device fingerprinting can be used by Internet Service Providers (ISPs) to identify vulnerable IoT devices for early prevention of threats. However, due to the wide deployment of middleboxes in ISP networks, some important data, e.g., 5-tuples and flow statistics, are often obscured, rendering many existing approaches invalid. It is further challenged by the high-speed traffic of hundreds of terabytes per day in ISP networks. This paper proposes DeviceRadar, an online IoT device fingerprinting framework that achieves accurate, real-time processing in ISPs using programmable switches. We innovatively exploit “key packets” as a basis of fingerprints only using packet sizes and directions, which appear periodically while exhibiting differences across different IoT devices. To utilize them, we propose a packet size embedding model to discover the spatial relationships between packets. Meanwhile, we design an algorithm to extract the “key packets” of each device, and propose an approach that jointly considers the spatial relationships and the key packets to produce a neighboring key packet distribution, which can serve as a feature vector for machine learning models for inference. Last, we design a model transformation method and a feature extraction process to deploy the model on a programmable data plane within its constrained arithmetic operations and memory to achieve line-speed processing. Our experiments show that DeviceRadar can achieve state-of-the-art accuracy across 77 IoT devices with 40 Gbps throughput, and requires only 1.3% of the processing time compared to GPU-accelerated approaches.
Ruoyu Li 0003, Qing Li 0006, Qingsong Zou, Dan Zhao 0003, Gareth Tyson, Guorui Xie, Yong Jiang 0001
IEEE/ACM Trans. Netw.8
2024 Empowering In-Network Classification in Programmable Switches by Binary Decision Tree and Knowledge Distillation
abstract
Given the high packet processing efficiency of programmable switches (e.g., P4 switches of Tbps), several works are proposed to offload the decision tree (DT) to P4 switches for in-network classification. Although the DT is suitable for the match-action paradigm in P4 switches, the range match rules used in the DT may not be supported across devices of different P4 standards. Additionally, emerging models including neural networks (NNs) and ensemble models, have shown their superior performance in networking tasks. But their sophisticated operations pose new challenges to the deployment of these models in switches. In this paper, we propose Mousikav2 to address these drawbacks successfully. First, we design a new tree model, i.e., the binary decision tree (BDT). Unlike the DT, our BDT consists of classification rules in the form of bits, which is a good fit for the standard ternary match supported by different hardware/software switches. Second, we introduce a teacher-student knowledge distillation architecture in Mousikav2, which enables the general transfer from other sophisticated models to the BDT. Through this transfer, sophisticated models are indirectly deployed in switches to avoid switch constraints. Finally, a lightweight P4 program is developed to perform classification tasks in switches with the BDT after knowledge distillation. Experiments on three networking tasks and three commodity switches show that Mousikav2 not only improves the classification accuracy by 3.27%, but also reduces the switch stage and memory usage by$2.00\times $and 28.67%, respectively. Code is available athttps://github.com/xgr19/Mousika.
Guorui Xie, Qing Li 0006, Guanglin Duan, Jiaye Lin, Yutao Dong, Yong Jiang 0001, Dan Zhao 0003, Yuan Yang 0001
IEEE/ACM Trans. Netw.1
2023 Efficient Flow Recording with InheritSketch on Programmable Switches
abstract
Several studies have been proposed to deploy the flow recording (i.e., flow size counting and sketching algorithms) on programmable switches for high-speed processing, helping network management tasks like scheduling. Although programmable switches provide a remarkable packet processing speed, they are of compact resources and follow a restrictive pipeline programming. To fit these limitations, current algorithms either sacrifice the recording accuracy or harm the switch throughput. In this paper, we propose InheritSketch for further improvement. InheritSketch utilizes a separation counting fashion, which is memory-efficient for compact switches. It accurately records the more valuable heavy hitters in the large key-value counters (i.e., the primary table), while only sketching non-heavy flows in the small sentinel table. With the recording ongoing, InheritSketch intelligently summarizes the historical recording experience as the basis for flow inheritance. That is, flows with the same IDs as the previous heavy hitters are regarded as new heavy hitters, being recorded in the primary table. To correct some incorrect inheritance, we also propose the flow rebellion, which promotes flows of large sizes but wrongly stored in the sentinel table to the primary table. InheritSketch is also helpful in applications like differentiated scheduling. We compare InheritSketch with six previous recording algorithms on three public traffic datasets, and prototype InheritSketch on a commodity P4 switch. The results demonstrate that InheritSketch reduces the recording errors by at most ∼7×, and that InheritSketch only consumes 10% of hardware resources on the switch.
Guorui Xie, Qing Li 0006, Guanglin Duan, Yong Jiang 0001, Zhuyun Qi, Qiaoling Wang
ICDCS1
2023 In-Forest: Distributed In-Network Classification with Ensemble Models
abstract
A variety of model representation methods have been used in recent works to translate machine learning models into programmable switch rules to address network classification tasks at line-speed, i.e., in-network classification. These works generally deploy a complete but heavy model on a switch with limited hardware resources, causing both network-wide waste of resources and unsatisfactory accuracy. Therefore, we propose In-Forest, a general distributed in-network classification framework. Firstly, to improve accuracy with limited resources, we develop a Lightweight Ensemble Generic Optional Model (LEGO), which can be further enhanced into multiple enhanced base models with full functionality. Each switch only needs to deploy a simple base model, rather than the complete ensemble model. Thus, hardware resources required for both switches and the entire network can be significantly reduced. Secondly, as traffic traverses multiple switches, In-Forest aggregates the classification results from different enhanced base models for higher accuracy. Furthermore, we design a two-phase resource-aware model allocation strategy that assigns enhanced base models to switches under different scenarios. We use stable deep reinforcement learning to respond to dynamic traffic changes. Experimental results show that when compared to SwitchTree, Planter, and Netbeacon in two real network topologies, In-Forest can increase accuracy by up to 19.31%, while reducing the number of switch rules by 89.98%.
Jiaye Lin, Qing Li 0006, Guorui Xie, Yong Jiang 0001, Zhenhui Yuan, Changlin Jiang, Yuan Yang 0001
ICNP3
2023 Efficient Attack Detection with Multi-Latency Neural Models on Heterogeneous Network Devices
abstract
To achieve fast and accurate attack detection, some works manually tailor neural networks (NNs) for deployment on CPUs of gateways, routers, or even programmable switches. However, with such solutions, NNs must be custom-tailored across different devices to meet the heterogeneous settings (e.g., OS and CPU types). Even worse, a model may require frequent adjustments to adapt to the same device's varying traffic rates. In this paper, we present Soteria, an automated multi-latency NN generation and scheduling system for fast and accurate detection against fluctuating traffic rates across heterogeneous hardware. Soteria first uses an evolutionary training algorithm to evolve the Pareto front, i.e., the set of NNs with a good spread on accuracy and model size. Then, for each device, Soteria filters the optimal multi-latency NNs by non-dominating sorting on the NNs' test latency on the device. Finally, to cope with the dynamic traffic rate, we design a heuristic scheduling scheme that adaptively selects NN s to maintain a balance between the detection accuracy and latency.
Guorui Xie, Qing Li 0006, Haolin Yan, Dan Zhao 0003, Gianni Antichi, Yong Jiang 0001
ICNP1
2023 Dryad: Deploying Adaptive Trees on Programmable Switches for Networking Classification
abstract
Decision trees (DT) have been used for high-speed networking classification on programmable switches. Most DT solutions, however, are static and cannot be deployed once the switch resource changes. In this paper, we propose Dryad to fast reprogram tree models when resource budgets change. In Dryad, we first develop a large and accurate “one-training-for-all“ DT (ODT) that can be quickly resized without computational retraining. ODTs are deployed in switches using a progressive search algorithm that searches the adaptations according to their resources. To achieve high accuracy and low packet latency, the adaptation leverages 1) innovative hard and soft pruning methods to compress the ODT rapidly with minimal performance loss; and 2) P4 scaling operations of match-action table arrangement and joint range-ternary match, which allow the switch to accommodate a larger (i.e., more accurate) ODT. Finally, an ODTCompiler is proposed to automatically convert the adapted ODT into a P4 program and then install it. Experimental results on three commodity switches under different resource scenarios show that Dryad achieves a higher classification F1-score (3.78 % higher), and completes the adaptation 161 × faster than other solutions.
Guorui Xie, Qing Li 0006, Jiaye Lin, Gianni Antichi, Dan Zhao 0003, Zhenhui Yuan, Ruoyu Li 0003, Yong Jiang 0001
ICNP1
2023 Pontus: Finding Waves in Data Streams
abstract
The bumps and dips in data streams are valuable patterns for data mining and networking scenarios such as online advertising and botnet detection. In this paper, we define the wave, a data stream pattern with a serious deviation from the stable arrival rate for a period of time. We then propose Pontus, an efficient framework for wave detection and estimation. In Pontus, a lightweight data structure is utilized for the preliminary processing of incoming packets in the data plane to take advantage of its high processing speed; then, the powerful control plane carries out computationally intensive wave detection and estimation. In particular, we propose the Multi-Stage Progressive Tracking strategy which detects waves in stages and removes any disqualified items promptly to save memory. Hash collisions are addressed by a Stage Variance Maximization technique to reduce estimation error. Moreover, we prove the theoretical error bound and establish upper bounds of false positive and false negative. Experiment results show that the software version of Pontus can achieve around 97% F1-Score even under scarce memory when baselines fail. Furthermore, the implemented prototype of Pontus based on P4 achieves 842x higher throughput than the baseline strawman solution.
Qing Li 0006, Guanglin Duan, Dan Zhao 0003, Jingyu Xiao, Guorui Xie, Yong Jiang 0001
Proc. ACM Manag. Data6
2022 Mousika: Enable General In-Network Intelligence in Programmable Switches by Knowledge Distillation
abstract
Given the power efficiency and Tbps throughput of packet processing, several works are proposed to offload the decision tree (DT) to programmable switches, i.e., in-network intelligence. Though the DT is suitable for the switches’ match-action paradigm, it has several limitations. E.g., its range match rules may not be supported well due to the hardware diversity; and its implementation also consumes lots of switch resources (e.g., stages and memory). Moreover, as learning algorithms (particularly deep learning) have shown their superior performance, some more complicated learning models are emerging for networking. However, their high computational complexity and large storage requirement are cause challenges in the deployment on switches. Therefore, we propose Mousika, an in-network intelligence framework that addresses these drawbacks successfully. First, we modify the DT to the Binary Decision Tree (BDT). Compared with the DT, our BDT supports faster training, generates fewer rules, and satisfies switch constraints better. Second, we introduce the teacher-student knowledge distillation in Mousika, which enables the general translation from other learning models to the BDT. Through the translation, we can not only utilize the super learning capabilities of complicated models, but also avoid the computation/memory constraints when deploying them on switches directly for line-speed processing.
Guorui Xie, Qing Li 0006, Yutao Dong, Guanglin Duan, Yong Jiang 0001, Jingpu Duan
INFOCOM1
2022 Soter: Deep Learning Enhanced In-Network Attack Detection Based on Programmable Switches
abstract
Though several deep learning (DL) detectors have been proposed for the network attack detection and achieved high accuracy, they are computationally expensive and struggle to satisfy the real-time detection for high-speed networks. Recently, programmable switches exhibit a remarkable throughput efficiency on production networks, indicating a possible deployment of the timely detector. Therefore, we present Soter, a DL enhanced in-network framework for the accurate real-time detection. Soter consists of two phases. One is filtering packets by a rule-based decision tree running on the Tofino ASIC. The other is executing a well-designed lightweight neural network for the thorough inspection of the suspicious packets on the CPU. Experiments on the commodity switch demonstrate that Soter behaves stably in ten network scenarios of different traffic rates and fulfills per-flow detection in 0.03s. Moreover, Soter naturally adapts to the distributed deployment among multiple switches, guaranteeing a higher total throughput for large data centers and cloud networks.
Guorui Xie, Qing Li 0006, Chupeng Cui, Peican Zhu, Dan Zhao 0003, Wanxin Shi, Zhuyun Qi, Yong Jiang 0001, Xi Xiao 0001
SRDS1
2021 Self-attentive deep learning method for online traffic classification and its interpretability
Guorui Xie, Qing Li 0006, Yong Jiang 0001
Comput. Networks1