EDBT 2026 Demo / reviewers in the wild / expert
Dong Zhang 0010
dblp:68/3245-10
· DBLP profile ↗
82ranked-venue papers
0as first author
74since 2021 · last 2026
0000-0002-6379-0244ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 70 · 64 since 2021Systems, architecture and hardware · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | APTMatch: Empowering Dynamic Rule Updating for Learning-based Packet Classification
Jiashuo Yu, Longlong Zhu, Zongye Lin, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001 |
ICC | 8 |
| 2026 | HYDRA: A Hybrid Synthesizer for Asymmetric Mixture-of-Experts Communication Scheduling
Lida Liao, Qianxun Xu, Xuanwei Si, Hongyan Liu 0001, Jiashuo Yu, Zongye Lin, Qiaoling Hu, Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001 |
ICC | 12 |
| 2026 | MonPlan: Taming Network Measurement with Accurate and Resource-Efficient Sketch-INT Co-Design
Xiang Chen 0017, Linying Zheng, Longlong Zhu, Zedi Chen, Qing Shu, Jialu Tian, Siqi Dong, Qun Huang 0001, Jianshan Zhang, Xuan Liu 0006, Haifeng Zhou, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001 |
INFOCOM | 13 |
| 2026 | LTD: Low-Overhead Topology Discovery using Programmable Data Planes
Dezhang Kong, Minghao Li 0012, Shi Lin, Zhenhua Xu 0004, Longlong Zhu, Linying Zheng, Xiang Chen 0017, Changting Lin, Xuan Liu 0006, Dong Zhang 0010, Chunming Wu 0001 |
INFOCOM | 10 |
| 2026 | BGP-Shadow: Accelerating BGP Verification via Programmable Data Planes
Hongyan Liu 0001, Dong Zhang 0010 |
IWQoS | 3 |
| 2026 | GBNN: In-Network Gradient Boosting Neural Network
Shaowei Xu, Shengrui Lin, Hongyan Liu 0001, Libin Xu, Dong Zhang 0010, Chunming Wu 0001 |
IWQoS | 8 |
| 2026 | SketchPipe: Toward Accurate Sketch-based Network Measurement on Multi-Pipeline Switches with Splitless Sketch Placement
Xiang Chen 0017, Longlong Zhu, Linying Zheng, Hongyang Du 0001, Dong Zhang 0010, Jianshan Zhang, Xuan Liu 0006, Qun Huang 0001, Dusit Niyato, Haifeng Zhou, Chunming Wu 0001, Hongyan Liu 0001, Kui Ren 0001 |
NSDI | 5 |
| 2026 | Proteus: Towards Accurate and Low-overhead In-Network Malicious Traffic DetectionabstractNetwork intrusion detection systems (NIDS) are essential for web security by identifying and dropping malicious traffic. Existing in-network NIDS leverage the Tbps-level packet processing capability of programmable switches to achieve high-speed flow classification. They translate complex trained machine learning models to decision trees (DTs), where DTs are deployed on programmable switches via single-DT or multiple-DT deployment. However, they face a fundamental trade-off: single-DT deployment suffers from low classification accuracy due to over-pruning of trees, while multiple-DT deployment suffers from high overhead due to deploying multiple tree replicas. In this paper, we propose Proteus, an in-network malicious traffic detection system that achieves both high classification accuracy and low overhead. Its key idea is to split the original DT into critical and normal sub-trees, where these sub-trees have different impacts on overall accuracy. More precisely, Proteus first splits a DT into one critical and several normal sub-trees for adapting to the accuracy requirement and switch resource budgets. Second, it minimizes coordination overhead between sub-trees while ensuring full flow coverage via mixed-integer linear programming. Third, it dynamically reallocates or migrates sub-trees to adapt to changing resources by monitoring both classification accuracy and switch resource changes. Testbed experiments with 12.8 Tbps programmable switches show that Proteus improves classification accuracy, reduces switch resource consumption, and reduces classification latency. Longlong Zhu, Linying Zheng, Qing Shu, Zedi Chen, Jiashuo Yu, Shaopeng Zhou, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001, Xiang Chen 0017 |
WWW | 9 |
| 2026 | DeepConfig: A verifiable configuration generation framework for MAN overlays using LLMs
Longlong Zhu, Hongyan Liu 0001, Dong Zhang 0010, Jiashuo Yu, Lida Liao |
Comput. Networks | 4 |
| 2026 | Toward Security-Enhanced In-Band Network Telemetry in Programmable NetworksabstractIn-band Network Telemetry (INT) is a widely used monitoring framework in modern large-scale networks. It provides packet-level visibility into network conditions by inserting telemetry data into packets, enabling unprecedented fine-grained network management. However, this mechanism also introduces new vulnerabilities that malicious attackers can exploit. In this paper, we present eight In-band Network Telemetry Manipulation Attacks that take advantage of INT’s weakness, demonstrating that attackers can cause severe damage with little effort by manipulating INT packets. To address this issue, we designed SecureINT, a security-enhanced INT prototype that provides encryption and integrity verification for INT packets. Specifically, SecureINT deploys Even-Mansour and SipHash for confidentiality and integrity, respectively. It also uses a zero-delay rotation mechanism, which enables administrators to dynamically change the version of the deployed Even-Mansour/SipHash running on programmable switches without the need to re-install new programs. In this way, SecureINT can provide lasting security for INT packets using the limited resources of programmable switches. According to the experiments, SecureINT can be deployed on programmable switches using a single pipeline. Besides, the overhead of the rotation mechanism running on the control plane is still minimal. Dezhang Kong, Xiang Chen 0017, Zhengyan Zhou, Yi Shen 0012, Hongyan Liu 0001, Qiumei Cheng, Xuan Liu 0006, Dong Zhang 0010, Chunming Wu 0001, Muhammad Khurram Khan |
IEEE Trans. Netw. Serv. Manag. | 9 |
| 2025 | Phantom: Virtualizing Switch Register Resources for Accurate Sketch-based Network MeasurementabstractSketches have proven to be useful for measuring traffic. They store measurement results in the registers of data plane switches. However, they suffer from the short of switch register resources, limiting their measurement accuracy. Xiang Chen 0017, Hongyan Liu 0001, Zhengyan Zhou, Wenbin Zhang 0011, Hongyang Du 0001, Dong Zhang 0010, Xuan Liu 0006, Haifeng Zhou, Dusit Niyato, Qun Huang 0001, Chunming Wu 0001, Kui Ren 0001 |
EuroSys | 7 |
| 2025 | DHC: Distributed Homomorphic Compression for Gradient Aggregation in AllreduceabstractDistributed training is critical for efficiently developing deep neural networks (DNNs) on tasks like image classification and natural language processing. However, as model and dataset sizes continue to grow, high communication overhead during gradient exchanges has become a major bottleneck in distributed training. Although existing homomorphic compression frameworks effectively reduce communication overhead, their reliance on centralized architectures makes them unsuitable for the mainstream decentralized AllReduce architecture. To address this, we propose DHC, a framework for homomorphic gradient compression in AllReduce architectures. Its key idea is HG-Sketch, which leverages multi-level index tables for direct in-network aggregation of compressed gradients, thereby eliminating additional computational overhead. Additionally, DHC introduces an index-sharing method to optimize memory usage on programmable switches. Furthermore, we establish an Integer Linear Programming (ILP) model to optimize the deployment strategy of programmable switches, further enhancing in-network aggregation capabilities. Experimental results demonstrate that DHC achieves a$3.8 \times$increase in aggregation speed and a$4.2 \times$improvement in aggregation throughput. Lida Liao, Zhengli Lin, Longlong Zhu, Hongyan Liu 0001, Jiashuo Yu, Dong Zhang 0010, Chunming Wu 0001 |
ICC | 7 |
| 2025 | P4Alex: A Scalable Range Matching Approach for Programmable SwitchesabstractRange matching (RM), a flexible primitive for implementing network applications on programmable switches, is often subject to limited TCAM resources. Consequently, existing RM approaches rely on SRAM/ALU-assisted data structures to extend TCAM capacity. However, they consume significant SRAM, ALU, and pipeline stages, which hinders the implementation of other primitives (e.g., basic forwarding). In this paper, we propose P4Alex, a scalable RM framework for programmable switches. The key idea is leveraging emerging learned index structures to support RM, which replaces the storage by model inference for lightweight and fixed index depth. Unfortunately, the learning index structure cannot be directly implemented on programmable switches due to hardware limitations (e.g., floating-point computation and no-loop operations). In response, we design several optimizations in P4Alex to make it deployable. We successfully implemented P4Alex on Intel Tofino switches. Experimental results show that, compared to existing RM approaches, P4Alex extends RM capabilities from tens of thousands to millions. At the same scales, P4Alex reduces TCAM and SRAM resource consumption by up to 89.9% and 43.7%, respectively, while increasing latency by only$0.48 \mu ~\mathrm{s}$. Jiashuo Yu, Longlong Zhu, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001 |
ICC | 4 |
| 2025 | Carrera: Enabling High-Performance eBPF-based Sketches in Network MeasurementabstractTo achieve dynamic network measurement, trends build sketches on eBPF to avoid service interruptions. However, existing eBPF-based sketches suffer from high CPU consumption, leading to poor throughput and high latency and making them hard to measure high-speed traffic. Optimizing their performance requires users to refactor codes based on each sketch’s characteristics on eBPF, which is highly complex and time-consuming.In this paper, we argue that users should write sketches without concerning low-level eBPF performance optimizations, with the deployment automatically activating cross-sketch performance optimizations. We present Carrera, a library that offers domain-specific optimizations for eBPF-based sketches. Our contributions are (1) systematically analyzing the performance bottlenecks of eBPF-based sketches through microbenchmarks, (2) identifying practical optimizations, including hardware offloading, SIMD-accelerated hashing, traffic-aware flow index caching, prefetched randomization, and active data collection, to address the identified bottlenecks in eBPF-based sketches, (3) evaluating these optimizations with state-of-the-art sketches and demonstrating that Carrera improves throughput by up to 65% and reduces latency by up to 93% via testbed experiments. Xiang Chen 0017, Xin Yao 0008, Longlong Zhu, Linying Zheng, Hongyan Liu 0001, Jianshan Zhang, Dong Zhang 0010, Xuan Liu 0006, Qun Huang 0001, Haifeng Zhou, Chunming Wu 0001 |
ICNP | 8 |
| 2025 | EffiMatch: Enabling Fast and Accurate Learning-based Packet ClassificationabstractLearning-based Packet Classification methods reduce memory overhead by using lightweight Recursive Model Index(RMI) structures to limit the search range, followed by linear matching. However, they face a trade-off: complex RMI structures achieve smaller search ranges but slow down lookup, while simpler ones are faster but require larger scans. In this paper, we propose EffiMatch, a parallel multi-model lookup architecture aimed at resolving the trade-off between RMI complexity and linear search range in learning-based index systems. We propose two key designs: 1) We design a partitioning strategy called Distribution-Distance Partitioning (DDP), which groups data points with similar trends into the same segment. Combined with parallel lookup, this reduces the linear search range while maintaining high lookup speed. 2) We propose a more fine-grained binarization method, Base-Index Representation (BI), which approximates floating-point operations using integers. This method further reduces the search range without increasing model complexity. Experimental results show that EffiMatch reduces the linear search range by 26.84% using lower-complexity RMI models, which improves lookup speed by up to 6× and reduces construction time by up to 4 orders of magnitude compared to state-of-the-art LPC methods. Lida Liao, Jiashuo Yu, Longlong Zhu, Hongyan Liu 0001, Dong Zhang 0010, Xiang Chen 0017, Chunming Wu 0001 |
ICNP | 7 |
| 2025 | TurboCache: Empowering Switch-Accelerated Key-Value Caches with Accurate and Fast Cache UpdatesabstractRecent key-value (KV) caches are offloaded to programmable switches to offer high query processing performance. However, they suffer from both low accuracy in hot key detection and high latency in cache updates due to the strict limitations on switch registers. We propose TurboCache, a switch-accelerated KV cache with accurate hot key detection and fast cache updates. Our key idea is to leverage the switch recirculation capability to build a novel data structure that caches hot KV pairs. With this hardware-compatible cache data structure, TurboCache designs efficient data plane algorithms that accurately detects new hot keys and quickly updates its cache entirely within switch ASIC pipelines. We have implemented TurboCache on a${64}\times {100}$Gbps Tofino switch. Testbed results indicate that TurboCache improves the hot key detection accuracy and decreases the cache update latency of existing KV caches by several orders of magnitude. Xiang Chen 0017, Longlong Zhu, Linying Zheng, Lingfei Cheng, Jianshan Zhang, Xu Yang 0002, Dong Zhang 0010, Xuan Liu 0006, Xiaoming Lu, Xun Yi, Ibrahim Khalil 0001, Albert Y. Zomaya, Haifeng Zhou, Chunming Wu 0001 |
INFOCOM | 7 |
| 2025 | Scaling Learning-based Packet Classification Hardware with NeuTree
Jiashuo Yu, Longlong Zhu, Linying Zheng, Dong Zhang 0010, Xiang Chen 0017 |
INFOCOM | 6 |
| 2025 | 6Map: Enabling Fast Active IPv6 Address Discovery with Programmable Switches
Lin He 0004, Yifan Yang 0009, Xiaoyi Shi, Daguo Cheng, Jinlong E, Ying Liu 0024, Dong Zhang 0010 |
INFOCOM | 8 |
| 2025 | Monica: Towards Scalable Distributed System Verification by Programmable Switch-Based TestingabstractData correctness in distributed systems is ensured by data consistency, where consistency is achieved by consensus algorithms. To safeguard data consistency, current testing tools use stress testing methods to examine consensus algorithms. However, existing tools are unable to simulate the situation under high traffic and suffer from excessive verification time. In this paper, we propose Monica, a scalable and efficient verification framework. Its key idea is to leverage the programmable switch to verify consensus algorithms. Specifically, Monica provides a set of primitives that researchers can invoke. Then, the control server recognizes the primitives and automatically configures the data plane. After that, the programmable switch collaborates with the control server to complete the verification. Experimental results show that Monica can generate traffic at the rate of Tbps level while keeping the computational and memory consumption of the programmable switch under 11.87%. Compared to existing testing tools, Monica increases the verification speed by up to 3.13 times. Further, Monica improved accuracy by 35.71% in high-traffic scenarios over other tools. Jiashuo Yu, Longlong Zhu, Dong Zhang 0010, Lida Liao, Rongbang Wu, Xiang Chen 0017, Chunming Wu 0001 |
IWQoS | 4 |
| 2025 | Polyx: Accelerating Verification of Traffic Migration in Large-Scale BGP NetworksabstractIn BGP networks, traffic migration verification ensures the scalability and reliability of the network during configuration changes. However, previous approaches suffer from low scalability and high computational overhead. In this poster, we propose Polyx, a framework for accelerating verification of traffic migration in large-scale BGP networks. Its key idea is to leverage hardware parallelism with a deterministic serialization algorithm to enhance state machine techniques. We implement the Polyx prototype and evaluate it on our built testbed. The experimental results demonstrate that Polyx achieves up to 46× overall speedup, 36× in state machine construction, and 131× in equivalence verification with minimal FPGA resource usage. Rongbang Wu, Longlong Zhu, Jiashuo Yu, Dong Zhang 0010, Hongyan Liu 0001, Zongye Lin, Lida Liao, Xiang Chen 0017, Chunming Wu 0001 |
IWQoS | 4 |
| 2025 | TBNN: Lookup Tables-Based Optimization for in-Network Binary Neural NetworksabstractBinary Neural Network (BNN) is a meaningful machine learning model on the data plane. However, due to the chip limitations, the scalability, especially the number of hidden layers in one pipeline, is limited. For better inference performance, existing methods reuse the hidden layers through packet recirculations. Recirculations lead to poor processing latency. Additionally, the simplified operations in the in-network BNN model restrict the flexibility of itself, which results in the unarbitrary input length of neurons for more the additional resource consumption than normal BNN model. In this paper, we present TBNN, an optimized in-network BNN model that achieves both scalability and flexibility. This approach eliminates deployment constraints while maximizing hardware utilization, advancing the feasibility of complex BNN models on resourcelimited data planes. By replacing computational bottleneck actions with Lookup Tables (LUTs), TBNN enables at most$4 \times$more neurons per pipeline and reduces per-packet latency by 50% through minimized recirculation. LUT-based implementation supports pruning operations, trading an accuracy loss of$\mathbf{1. 6 9 \%}$for saving about$\mathbf{2 4 \%}$instructions. Shaowei Xu, Shengrui Lin, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001 |
IWQoS | 4 |
| 2025 | Handling Data Plane Program Deployment Dynamics with High-Quality Generative Diffusion ModelsabstractDeploying data plane programs across the network is typically formulated as a mixed-integer programming task, leading to a long execution time. In response, existing studies carefully tailor heuristics for specific task properties such as objectives. However, they suffer from poor solution quality under dynamic task deployment since they overfit specific task properties. Recently, generative diffusion models have been widely adopted in network optimizations due to their strong adaptability and generalization. Accordingly, in this poster, we propose a diffusion model-based framework for data plane program deployment tasks. Our key idea is to leverage the reverse denoising process of diffusion models to react to dynamic task changes at runtime while maintaining high solution quality. Preliminary results on our testbed show that we reduce latency by 66.67% and resource overhead by 58.62% during dynamic deployment. Longlong Zhu, Jiashuo Yu, Xiang Chen 0017, Qing Shu, Zedi Chen, Zhifan Jiang, Qun Huang 0001, Xuan Liu 0006, Dong Zhang 0010, Chunming Wu 0001 |
IWQoS | 9 |
| 2025 | NDIF: A distributed framework for efficient in-network neural network inference
Shengrui Lin, Shaowei Xu, Binjie He, Hongyan Liu 0001, Dezhang Kong, Xiang Chen 0017, Dong Zhang 0010, Chunming Wu 0001, Ming Li 0056, Xuan Liu 0006, Yuqin Wu, Muhammad Khurram Khan |
Comput. Secur. | 7 |
| 2025 | Elastically Scaling Control Channels in Network Measurement With EscalaabstractIn network measurement, data plane switches measure traffic and report events (e.g., heavy hitters) to the control plane via control channels. The control plane makes decisions to process events. However, current network measurement suffers from two problems. First, when traffic bursts occur, massive events are reported in a short time so that the control channels may be overloaded due to limited bandwidth capacity. Second, only a few events are reported in normal cases, making control channels underloaded and wasting network resources. In this paper, we propose$\textsf {Escala}$to provide the elastic scaling of control channels at runtime. The key idea is to dynamically migrate event streams among control channels to regulate the loads of these channels.$\textsf {Escala}$offers two components, including an$\textsf {Escala}$monitor that detects scaling situations based on realtime network statistics, and an optimization framework that makes scaling decisions to eliminate overload and underload situations. We have implemented a prototype of$\textsf {Escala}$on Tofino-based switches. Extensive experiments show that$\textsf {Escala}$achieves timely elastic scaling while preserving high application-level accuracy. Hongyan Liu 0001, Xiang Chen 0017, Qun Huang 0001, Dezhang Kong, Dong Zhang 0010, Chunming Wu 0001, Xuan Liu 0006 |
IEEE Trans. Netw. | 5 |
| 2025 | Toward Secure Inter-Device Coordination in Programmable NetworksabstractIn programmable networks, some networking systems coordinate data plane switches to perform in-network functions (e.g., in-band network telemetry). However, the vulnerabilities associated withinter-device coordinationremain largely unexplored and overlooked, which is highly concerning given the increasing popularity of this paradigm. In this paper, we identify three attack scenarios built upon such vulnerabilities, where attackers mislead the behaviors of networking systems. We implement 20 networking systems on Tofino-based switches and a simulator and test them against the identified attacks. Our experimental results show that our attacks severely disrupt the normal operation of these networking systems, e.g., the cache hit rate of NetCache drops by 38%. However, our analysis reveals that none of existing methods fully mitigate our attacks because they fail to verify the packets for inter-device coordination. To this end, we select characteristics from existing methods while addressing their limitations to design effective mitigation methods. Experimental results indicate that our methods perform well in mitigating our attacks and introduce acceptable overheads. Hongyan Liu 0001, Xiang Chen 0017, Di Wang 0049, Qun Huang 0001, Dong Zhang 0010, Chunming Wu 0001, Xuan Liu 0006 |
IEEE Trans. Netw. | 5 |
| 2024 | X-Match: A Semi-Supervised Framework for Oral Jawbones Segmentation Using Wavelet Transform for Enhanced Consistency LearningabstractDetermining the occlusal position in CBCT images is a critical step in the digital virtual articulator treatment of Anterior Disc Displacement with Reduction (ADDwR). Current oral segmentation methods primarily rely on supervised learning, which typically requires a large dataset. However, acquiring oral datasets is complex and challenging, making semi-supervised learning more suitable. Current semi-supervised segmentation methods have several limitations. The perturbations used in consistency-based semi-supervised methods are often manually designed, which can introduce negative biases detrimental to training. Furthermore, semi-supervised learning often faces an empirical mismatch between labeled and unlabeled data. When these two data types are handled independently or without alignment, significant information derived from labeled data may not be fully utilized. We propose a novel semi-supervised framework X-Match for oral jawbones segmentation. The X-Match utilizes wavelet transforms to extract low-frequency and high-frequency information for consistency training, reducing the learning bias caused by manual perturbations. Furthermore, it combines labeled with unlabeled data bidirectionally during training, allowing unlabeled data to acquire comprehensive shared features from labeled data. Experimental results demonstrate that our method outperforms baselines and achieves superior performance in oral jawbones dataset. Zhengkai Weng, Songwei Zheng, Chunyan Yu, Danhong Zhu, Linghui Jia, Dong Zhang 0010 |
BIBM | 7 |
| 2024 | MAReraser: Metal Artifact Reduction with Image Prior Using CNN and Transformer TogetherabstractThis paper presents a new dual domain network with image prior based on Convolutional Neural Network (CNN) and Transformer simultaneously for CT Metal Artifact Reduction (MAR). Challenges in MAR derive from the following aspects: firstly, the different morphologies of metal artifacts complexify resolving the issue just in a single domain; secondly, albeit many methods excel in quantitative metrics, yet the restored anatomical structures are over-smooth blurring reconstructed CT images; thirdly, MAR demands better performance as a clinical application, but the approaches relying on CNN or Transformer struggle due to CNN’s restricted spatial scope and Transformer’s ignorance to the local details, respectively, that is, CNN focuses on the local information while Transformer emphasizes the global information with higher computational complexity. To address these problems, we put forward MAReraser, a novel dual domain network, to deal with metal artifacts. MAReraser removes metal artifacts in both the projection and image domains, effectively reducing heteromorphic metal artifacts. Moreover, MAReraser introduces the image prior generated by an image prior subnet to refine the quality of reconstructed CT images. The prior subnet is pretrained in an expanded dataset which incorporates CT images corrected by diverse traditional MAR methods, providing extra potential prior knowledge from different perspectives. Further, the network backbone of MAReraser integrates CNN and Transformer, enabling complementary local and global feature extraction and balancing computational complexity. Extensive experiment results demonstrate that our method outperforms several other approaches whether in quantitative metrics or in qualitative visualization results. Songwei Zheng, Dong Zhang 0010, Chunyan Yu, Linghui Jia, Longlong Zhu, Zhanchao Huang, Danhong Zhu |
BIBM | 2 |
| 2024 | N4: Network for N Neural Network TrainingabstractAs the amount of data and complexity of neural network models continue to grow, distributed training has become increasingly crucial for improving training speed. However, the bottleneck of distributed training is the communication overheads among distributed workers. Recent research has shown that performing in-network aggregation using programmable switches is a good way to accelerate distributed training. However, previous work has only targeted specific neural network models and can only be applied in specified network topologies. Administrators may train different models and train them in different network topologies. In order to generalize the approach of using programmable switches to accelerate distributed training, we propose N4, a programmable intra-switch acceleration framework that supports distributed training of multiple neural networks. N4 also realizes the deployment of distributed workers based on any topology. Our experimental results show that N4 ensures high performance and isolation when training numerous neural networks. N4 outperforms state-of-the-art systems, accelerating training for existing methods by up to 3.4×. Shengrui Lin, Hongyan Liu 0001, Pengpai Shi, Longlong Zhu, Dong Zhang 0010 |
ICC | 7 |
| 2024 | SpotMon: Enabling General Hotspot Monitoring in Key-Value StoresabstractKey-value stores are essential to online services such as e-commerce. In key-value stores, a hotspot (i.e., frequently accessed items) may cause severe load imbalances, high response latency, and Service Level Agreement (SLA) violations. However, existing works only focus on specific types of hotspots, thus overlooking other types of hotspots and leading to blind spots. In this paper, we propose SpotMon, a system that enables general hotspot monitoring in key-value stores. Specifically, we (1) systematically identify the generality requirements of hotspot monitoring from existing works, (2) formulate general hotspot monitoring as an arbitrary partial spot query problem, (3) measure the hotness of hotspot candidates with a new vector expression, (4) propose hotspot encoding, filtering, decoding, and querying to support general queries without focusing on specific hotspots, (5) leverage the in-network visibility of programmable switches to identify system-wide hotspots. Our extensive experiments indicate that SpotMon provides high accuracy (e.g., F1 score from 0.88 to 1) and enables efficient hotspot mitigations (e.g., up to$4.03 \times$MQPS). Zhengyan Zhou, Jinhan Zu, Enhao Huang, Haifeng Zhou, Dong Zhang 0010, Xiang Chen 0017, Chunming Wu 0001 |
ICNP | 6 |
| 2024 | OpenINT: Dynamic In-band Network Telemetry with Lightweight Deployment and Flexible PlanningabstractThe normal operation of data center network management tasks relies on accurate measurement of the network status. In-band Network Telemetry (INT) leverages programmable data planes to provide fine-grained and accurate network status. However, existing INT-related works have not considered the telemetry data required for dynamic adjustments of INT under uninterrupted conditions, including additions, deletions, and modifications. To address this issue, this paper proposes OpenINT, a lightweight and flexible In-band Network Telemetry system. The key innovation of OpenINT lies in decoupling telemetry operations in the data plane, using three generic sub-modules to achieve lightweight telemetry. Meanwhile, the control plane utilizes heuristic algorithms for dynamic planning to achieve near-optimal telemetry paths. Additionally, OpenINT provides primitives for defining network measurement tasks, which abstract the underlying telemetry architecture’s details, enabling network operator to conveniently access network status. A prototype of OpenINT is implemented on a programmable switch equipped with the Tofino chip. Experimental results demonstrate that OpenINT achieves highly flexible dynamic telemetry and significantly reduces network overhead. Jiayi Cai, Tingxin Sun, Zhengyan Zhou, Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001 |
INFOCOM | 8 |
| 2024 | TupleRadar: Accelerating Tuple Space Search in Packet Classification by Learned IndexabstractTuple space search(TSS)-based packet classification is the keystone of network system. Previous studies accelerate TSS by partitioning tuples, combining trees and tuples, and merging tuples. However, they do not scale with the number of rules, resulting in a high memory footprint or update time. In this paper, we propose TupleRadar, a framework for accelerating TSS while ensuring low memory footprint and fast rule updates. Our key idea is to construct learned indexes for tuples, which inherently improve the lookup speed but ensure the advantages of TSS. Specifically, TupleRadar builds orderly hash table-based tuples and then constructs the updatable learned index. It provides a bounded memory footprint of the index structure as well. We have evaluated TupleRadar on multiple scales rule-sets. Experimental results show that TupleRadar outperforms previous solutions, reducing 46.66% lookup time and 61.53% memory footprint on average, by up to 86.70% and 88.95%. It also performs a competitive rule update speed. Longlong Zhu, Jiashuo Yu, Kaiwei Huang, Zhengyan Zhou, Dong Zhang 0010, Xiang Chen 0010, Chunming Wu 0001 |
IWQoS | 7 |
| 2024 | CardSketch: Shift Attention for Network-wide Cardinality TelemetryabstractNetwork telemetry is an essential part of network management and infrastructure. Among them, cardinality telemetry provides statistics on network connectivity and distribution. Network-wide cardinality telemetry refers to the deployment of multiple telemetry nodes in network for cardinality estimate. This requires the deployed data structure to be mergeable, enabling the consolidation of data from different nodes. Unfortunately, existing mergeable data structures can’t simultaneously address two important criterions of cardinality telemetry: measurement accuracy and estimation interval. We propose CardSketch, aiming to adjust attention to cardinality telemetry based on changes of the network state. CardSketch incorporates a shift attention mechanism that leverages the randomness of hash functions to achieve unbiased transformations between data structures. This mechanism enables real-time selection of cardinality estimation methods based on the network’s state while preserving the original telemetry information as much as possible during the attention shift. We have implemented prototypes of CardSketch in software and hardware. Through extensive experimentation, the results demonstrate that CardSketch achieves excellent cardinality telemetry with minimal memory overhead. Even with a mere 50KB of memory space, it achieves a measurement precision of 87.75% and a measurement recall of 91.49%. Additionally, CardSketch supports multi-point aggregation and arbitrary partial key queries. Hanze Chen, Zhengyan Zhou, Pengpai Shi, Yanni Wu, Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001 |
LCN | 8 |
| 2024 | DOT: Towards Fast Decision Tree Packet Classification by Optimizing Rule PartitionsabstractPacket classification is a crucial component of modern networks. Existing decision tree-based algorithms alleviate the rule replication problem caused by overlapping rules in the ruleset via rule partitioning. They partition the ruleset into multiple subsets based on rule characteristics to reduce rule overlaps. However, existing algorithms fail to address the overlap between rules in the same set, seriously decreasing speed and memory performance. In this paper, we propose DOT, a framework for optimizing rule partitions before constructing decision trees. Its key idea is to migrate rules in subsets based on rule overlaps and the features of heuristics used to construct trees, as well as reorganize rules aided by tuples. DOT finds out the migrated rule candidates using rule dependency graphs and heuristic features, then transforms the rule migration problem into an integer linear programming problem and solves for the optimal migration strategy. Further, we employ a tuple-assisted approach to accelerate rule matching. Experiments show that DOT enhances existing decision tree-based algorithms, improving lookup speed by 1.69 ×, reducing average 24.85% memory consumption and 31.03% decision tree depth. Longlong Zhu, Jiashuo Yu, Linying Zheng, Dong Zhang 0010, Chunming Wu 0001 |
LCN | 6 |
| 2024 | TransTuple: Toward Fast Packet Classification via Adaptive Tuple ReplacementabstractOpen vSwitch (OVS) is a widely used software switch in virtualized environments and software-defined networks. OVS uses tuple space search (TSS) for packet classification in the datapath, allowing fast network rule updates, but the increasing number of rules poses a classification performance challenge. To address this, existing methods incorporate decision trees with TSS to form a hybrid structure, enhancing classification speed. However, decision trees tend to overfit the initial ruleset, becoming unbalanced after rule updates and leading to a sharp decline in classification performance. In this paper, we propose TransTuple, a framework to optimize hybrid structures for fast packet classification under rule updates. The core idea of TransTuple is to identify bottleneck branches in decision trees that degrade performance and to replace them with lightweight tuples, providing better throughput under rule updates. These tuples maintain rules using hash tables, enabling fast updating and packet matching on bottleneck branches. We use TransTuple to optimize three state-of-the-art hybrid structured methods, i.e., CutTSS, TabTree, and MBitTree, achieving up to a 3.1x improvement in classification speed during rule updates. Jiashuo Yu, Longlong Zhu, Rongbang Wu, Linying Zheng, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001 |
SECON | 6 |
| 2024 | Eagle: Toward Scalable and Near-Optimal Network-Wide Sketch Deployment in Network MeasurementabstractSketches are useful for network measurement thanks to their low resource overheads and theoretically bounded accuracy. However, their network-wide deployment suffers from the trade-off between optimality and scalability: (1) Most solutions rely on mixed integer linear programming (MILP) solvers to provide the optimal decisions. But they are time-consuming and can hardly scale to large-scale deployment scenarios. (2) While heuristics achieve scalability, they deteriorate resource and performance overheads. We propose Eagle, a framework that achieves scalable and near-optimal network-wide sketch deployment. Our key idea is to decompose network-wide sketch deployment into sub-problems. Such decomposition allows Eagle to (1) simultaneously optimize switch resource consumption and end-to-end performance (retaining optimality), and (2) incorporate time-saving techniques into sub-problem solving (achieving scalability). Compared to existing solutions, Eagle improves scalability by up to 255× with negligible loss of optimality. It has also saved administrators in a production network days of efforts and reduced the operation time from O(hour) to O(second). Xiang Chen 0017, Qingjiang Xiao, Hongyan Liu 0001, Qun Huang 0001, Dong Zhang 0010, Xuan Liu 0006, Longbing Hu, Haifeng Zhou, Chunming Wu 0001, Kui Ren 0001 |
SIGCOMM | 5 |
| 2024 | P4Rex: Accelerating regular expression matching with programmable switches
Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001 |
Comput. Networks | 5 |
| 2024 | Efficient service reconfiguration with partial virtual network function migration
Dongquan Liu, Zhengyan Zhou, Dong Zhang 0010, Kaiwei Guo, Yanni Wu, Chunming Wu 0001 |
Comput. Networks | 3 |
| 2024 | Terra: Low-latency and reliable event collection in network measurement
Hongyan Liu 0001, Xiang Chen 0017, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006, Muhammad Khurram Khan |
J. Netw. Comput. Appl. | 5 |
| 2024 | Building a Hierarchical Architecture and Communication Model for the Quantum InternetabstractThe research of architecture has tremendous significance in realizing quantum Internet. Although there is not yet a standard quantum Internet architecture, the distributed architecture is one of the possible solutions, which utilizes quantum repeaters or dedicated entanglement sources in a flat structure for entanglement preparation & distribution. In this paper, we analyze the distributed architecture in detail and demonstrate that it has three limitations: 1) possible high maintenance overhead, 2) possible low-performance entanglement distribution, and 3) unable to support optimal entanglement routing. We design a hierarchical quantum Internet architecture and a communication model to solve the problems above. We also present a W-state Based Centralized Entanglement Preparation & Distribution (W-state Based CEPD) scheme and a Centralized Entanglement Routing (CER) algorithm within our hierarchical architecture and perform an experimental comparison with other entanglement preparation & distribution schemes and entanglement routing algorithms within the distributed architecture. The evaluation results show that the entanglement distribution efficiency of hierarchical architecture is 11.5% higher than that of distributed architecture on average (minimum 3.3%, maximum 37.3%), and the entanglement routing performance of hierarchical architecture is much better than that of a distributed architecture according to the fidelity and throughput. Binjie He, Dong Zhang 0010, Seng W. Loke, Shengrui Lin, Luke Lu |
IEEE J. Sel. Areas Commun. | 2 |
| 2024 | rDefender: A Lightweight and Robust Defense Against Flow Table Overflow Attacks in SDNabstractThe flow table is a critical component of Software-Defined Networking (SDN). However, flow tables’ limited capacity makes them highly vulnerable to flow table overflow attacks (FTOAs). Due to the low attack cost and highly flexible attack forms, it is hard to eradicate FTOAs. This paper addresses three unsolved problems for table security and proposes a robust defense accordingly. First, we reveal that the existing defenses with fixed defense speeds will cause severe packet loss when handling diverse traffic. We prove that deleting multiple rules can efficiently solve this problem and give a rigorous derivation to calculate the suitable deletion number according to the environment. Second, we illustrate that abnormal table occupancy squeezing is a constant characteristic of FTOAs regardless of attack forms. It can be used to identify attacked ports accurately in different scenarios. Third, we mathematically prove that random deletion can guarantee the continuous decrease of malicious flow rules after confirming attacked ports. It achieves fast speed and robust effectiveness in different environments. Based on these findings, we design rDefender, a robust and lightweight defense prototype. We evaluate its effect by designing diverse, powerful attacks and using real-world datasets and topology. The results demonstrate that it achieves the best overall performance compared to six existing mainstream defenses, providing stable security for switch flow tables. Dezhang Kong, Xiang Chen 0017, Chunming Wu 0001, Yi Shen 0012, Zhengyan Zhou, Qiumei Cheng, Xuan Liu 0006, Yubing Qiu, Dong Zhang 0010, Muhammad Khurram Khan |
IEEE Trans. Inf. Forensics Secur. | 10 |
| 2024 | Toward Scalable and Low-Cost Traffic Testing for Evaluating DDoS Defense SolutionsabstractTo date, security researchers evaluate their solutions of mitigating distributed denial-of-service (DDoS) attacks via kernel-based or kernel-bypassing testing tools. However, kernel-based tools exhibit poor scalability in attack traffic generation while kernel-bypassing tools incur unacceptable monetary cost. We propose Excalibur, a scalable and low-cost testing framework for evaluating DDoS defense solutions. The key idea is to leverage the emerging programmable switch to empower testing tasks with Tbps-level scalability and low cost. Specifically, Excalibur offers intent-based primitives to enable academic researchers to customize testing tasks on demand. Moreover, in view of switch resource limitations, Excalibur coordinates both a server and a programmable switch to jointly perform testing tasks. It realizes flexible attack traffic generation, which requires a large number of resources, in the server while using the switch to increase the sending rate of attack traffic to Tbps-level. We have implemented Excalibur on a$64\times 100$Gbps Tofino switch. Our experiments on a$64\times 100$Gbps Tofino switch show that Excalibur achieves orders-of-magnitude higher scalability and lower cost than existing tools. Xiang Chen 0017, Hongyan Liu 0001, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006 |
IEEE/ACM Trans. Netw. | 4 |
| 2024 | Hermes: Low-Overhead Inter-Switch Coordination in Network-Wide Data Plane Program DeploymentabstractNetwork administrators usually realize network functions in data plane programs. They employ the network-wide program deployment that decomposes input programs into match-action tables (MATs) while deploying each MAT on a specific switch. Since MATs may be deployed on different switches, existing solutions propose the inter-switch coordination that uses the per-packet header space to deliver crucial packet processing information among switches. However, such coordination incurs non-trivial per-packet byte overhead, leading to end-to-end performance degradation. We propose, a framework that aims to minimize the per-packet byte overhead. The key idea is to formulate network-wide program deployment as a mixed-integer programming (MIP) problem with the objective of minimizing the per-packet byte overhead. Also, offers a greedy-based heuristic that solves the problem in a near-optimal and timely manner. We have implemented on Tofino switches. Compared to existing frameworks, decreases the per-packet byte overhead by 156 bytes while preserving end-to-end performance in terms of flow completion time and goodput. Xiang Chen 0017, Hongyan Liu 0001, Qingjiang Xiao, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006, Qiang Yang 0004 |
IEEE/ACM Trans. Netw. | 5 |
| 2024 | Toward Full-Coverage and Low-Overhead Profiling of Network-Stack LatencyabstractIn modern data center networks (DCNs), network-stack processing denotes a large portion of the end-to-end latency of TCP flows. So profiling network-stack latency anomalies has been considered as a crucial part in DCN performance diagnosis and troubleshooting. In particular, such profiling requires full coverage (i.e., profiling every TCP packet) and low overhead (i.e., profiling should avoid high CPU consumption in end-hosts). However, existing solutions rely on system calls or tracepoints in end-hosts to implement network-stack latency profiling, leading to either low coverage or high overhead. We propose Torp, a framework that offers full-coverage and low-overhead profiling of network-stack latency. Our key idea is to offload as much of the profiling from costly system calls or tracepoints to the Torp agent built on eBPF modules, and further to include a Torp handler on the ToR switch to accelerate the remaining profiling operations. Torp efficiently coordinates the ToR switch and the Torp agent on end-hosts to jointly execute the entire latency profiling task. We have implemented Torp on$32\times 100$Gbps Tofino switches. Testbed experiments indicate that Torp achieves full coverage and orders of magnitude lower host-side overhead compared to other solutions. Xiang Chen 0017, Hongyan Liu 0001, Wenbin Zhang 0011, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Xuan Liu 0006, Chunming Wu 0001 |
IEEE/ACM Trans. Netw. | 5 |
| 2024 | Resource-Efficient and Timely Packet Header Vector (PHV) Encoding on Programmable SwitchesabstractThe programmable switch offers a limited capacity of packet header vector (PHV) words that store packet header fields and metadata fields defined by network functions. However, existing switch compilers employ inefficient strategies of encoding fields on PHV words. Their encoding wastes scarce PHV words and may result in failures when deploying network functions. In this paper, we propose Melody, a new framework that reuses PHV words for as many fields as possible to achieve resource-efficient PHV encoding. Melody offers a field analyzer and an optimization framework. The analyzer identifies which fields can reuse PHV words while preserving the original packet processing logic. The framework integrates analysis results into its encoding to offer the resource-optimal decisions. Also, to achieve timeliness at runtime, it provides a Greedy-based heuristic, which quickly solves PHV encoding and returns near-optimal results. We evaluate Melody with production-scale network functions. Our results show that Melody reduces the consumption of PHV words by up to 85%. Xiang Chen 0017, Wenbin Zhang 0011, Hongyan Liu 0001, Jianshan Zhang, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Xuan Liu 0006, Chunming Wu 0001 |
IEEE/ACM Trans. Netw. | 7 |
| 2024 | Toward Resource-Efficient and High- Performance Program Deployment in Programmable NetworksabstractProgrammable switches allow administrators to customize packet processing behaviors in data plane programs. However, existing solutions for program deployment fail to achieve resource efficiency and high packet processing performance. In this paper, we propose SPEED, a system that provides resource-efficient and high-performance deployment for data plane programs. For resource efficiency, SPEED merges input data plane programs by reducing program redundancy. Then it abstracts the substrate network into an one big switch (OBS), and deploys the merged program on the OBS while minimizing resource usage. For high performance, SPEED searches for the performance-optimal mapping between the OBS and the substrate network with respect to network-wide constraints. It also maintains program logic among different switches via inter-device packet scheduling. We have implemented SPEED on a Barefoot Tofino switch. The evaluation indicates that SPEED achieves resource-efficient and high-performance deployment for real data plane programs. Hongyan Liu 0001, Xiang Chen 0017, Qun Huang 0001, Peiqiao Wang, Dong Zhang 0010, Chunming Wu 0001, Xuan Liu 0006, Qiang Yang 0004 |
IEEE/ACM Trans. Netw. | 6 |
| 2023 | Aigis: Full-Coverage And Low-Overhead Mitigating Against Amplified Reflection DDoS AttacksabstractIn Internet Service Provider (ISP) networks, Amplified Reflection DDoS (AR-DDoS) attack is one of the main attack categories, which launches gigabytes of traffic with little effort and minimal cost. Thus, the mitigation of AR-DDoS attacks has been considered as a crucial part. In particular, such mitigation requires full coverage (i.e., mitigating AR-DDoS attacks launched from any location) and low overhead (i.e., mitigation should avoid high latency that degrades user experience). However, existing solutions suffer from either limited coverage or high overhead. In this paper, we propose Aigis, a distributed framework that offers full-coverage and low-overhead mitigation of AR-DDoS attacks. Our key idea is to co-design top-of-rack (ToR) switches and end-hosts, which offers line-rate packet processing performance and fine-grained view inherently, to jointly execute endpoint verification. Specifically, Aigis selectively offloads mitigation operations between ToR switches and end-hosts and implements a network-wide epoch synchronization mechanism to guarantee reliable verification. It efficiently coordinates ToR switches and end-hosts to execute the entire mitigation task. We have implemented Aigis on a testbed comprising 32×100 Gbps Tofino switches. Testbed experiments indicate that Aigis achieves complete full coverage and orders of magnitude lower host-side overhead compared to existing solutions. Tingxin Sun, Jiayi Cai, Kaiwei Guo, Dong Zhang 0010, Xiang Chen 0010, Chunming Wu 0001 |
GLOBECOM | 4 |
| 2023 | Vision Transformer with Progressive Tokenization for CT Metal Artifact ReductionabstractHigh-quality Computed Tomography(CT) plays a vital role in clinical diagnosis, but the presence of metallic implants will introduce severe metal artifacts on CT images and obstruct doctors’ decision-making. Many prior researches on Metal Artifact Reduction(MAR) are based on Convolutional Neural Network(CNN). Recently, Transformer has demonstrated phenomenal potential in computer vision. Also, transformer-based methods have been harnessed in CT image denoising. Nevertheless, these methods have been little explored in MAR. To fill the gap, we put forth, to the best of our knowledge, the first transformer-based architecture for MAR. Our method relies on a standard Vision Transformer(ViT). Furthermore, we tap into the progressive tokenization to refrain from the simple tokenization of ViT which gives rise to inability to model the local anatomical information. Additionally, for the sake of facilitating the interaction among tokens, we take advantage of cyclic shift from Swin Transformer. Finally, many experiment results reveal that the transformer-based technique is superior to those on the basis of CNN to some degree. Songwei Zheng, Dong Zhang 0010, Chunyan Yu, Danhong Zhu, Longlong Zhu, Zhongzheng Huang |
ICASSP | 2 |
| 2023 | MINT: Empowering Multiple Flow Definition Query for Network-Wide MeasurementabstractNetwork management tasks rely on precise and fine-grained network information to make correct and appropriate decisions. These tasks (e.g., DDoS detection) require network information with multiple flow definitions to better manage the network. However, the existing works mainly focus on the query of multiple flow definitions on a single switch, without a thoughtful solution for this query in network-wide measurement. In this paper, to address this problem, we overcome several challenges and propose MINT, a system that enables the query for multiple flow definitions in network-wide measurement. The key insights of MINT are: deploying MFSketch to measure multiple flow definitions information on the switch, cutting MFSketch into fixed-size slices, and using in-band telemetry (INT) to carry the slice to the analyzer. Therefore, after the analyzer collects and reorganizes the slices, network operators can query multiple flow definitions information of the whole network for various network management tasks. We implemented a prototype of MINT on a Barefoot Tofino switch. Experimental results show that MINT provides reliable transmission and consistency guarantees while only using switch resources comparable to state-of-the-art works, with less than 1% additional network overhead. Additionally, MFSketch provides accurate measurements for multiple flow definitions query, outperforming other solutions in both accuracy and F1 score. Jiayi Cai, Zhengyan Zhou, Tingxin Sun, Jiashuo Yu, Longlong Zhu, Chengze Li, Dong Zhang 0010, Chunming Wu 0001 |
ICC | 8 |
| 2023 | Halia: Toward Full-Coverage Network Function Offloading in the Data PlaneabstractOffloading network functions (NFs) to data plane switches brings remarkable performance benefits. In such offloading, NFs are required to process all the flows of interest (i.e., full coverage) to preserve the quality of services. However, existing solutions fail to guarantee full coverage for NFs. Thus, NFs may miss some essential flows, leading to accuracy drops. In this paper, we propose Halia, a framework that makes NF offloading decisions while ensuring full coverage for NFs. Specifically, Halia formulates the problem of NF offloading as an optimization problem. It encodes the requirement of full coverage as a constraint. Thus, its decisions activate enough NF instances in the substrate network to achieve full coverage.. We have implemented Halia and conducted experiments under multiple realistic network topologies to evaluate Halia. The experimental results indicate that compared to existing solutions, Halia achieves full coverage and high scalability in large-scale networks. Hongyan Liu 0001, Xiang Chen 0017, Qingjiang Xiao, Kaiwei Guo, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
ICC | 6 |
| 2023 | MiCuts: Combing Bit-Based Cutting and Splitting for Efficient Packet ClassificationabstractPacket classification is a crucial component in computer networking. To achieve high throughput and low memory consumption, existing solutions apply different heuristics in each construction stage to build efficient decision trees. However, previous studies divide the tree construction process based on the scale of rule subsets which is indirect to the performance goal, leading to massive rule replication and high tree depth. In this paper, we propose MiCuts, a fine-grained framework for packet classification with both high speed and low memory footprint. Its key idea is directly utilizing rule replication and tree depth to divide the tree-building process into three stages, each with suitable optimization goals. First, it partitions rules and builds shallow semi-trees without rule replication via selecting effective bits. Second, it transforms the switching problem of heuristics into an ILP problem and aims to minimize memory consumption while ensuring high lookup speed. Third, it merges some nodes to eliminate memory explosion caused by splitting, where MiCuts combines splitting and linear search. Extensive experimental results on ClassBench show that MiCuts outperforms state-of-the-art approaches, improving lookup speed by 1.71× while reducing memory footprint by 74.4% on average. Longlong Zhu, Jiashuo Yu, Linying Zheng, Jinfeng Pan, Zhengyan Zhou, Hanze Chen, Dong Zhang 0010, Xiang Chen 0010, Chunming Wu 0001 |
ICC | 8 |
| 2023 | Excalibur: A Scalable and Low-Cost Traffic Testing Framework for Evaluating DDoS Defense SolutionsabstractTo date, security researchers evaluate their solutions of mitigating denial-of-service (DDoS) attacks via kernel-based or kernel-bypassing testing tools. However, kernel-based tools exhibit poor scalability in attack traffic generation while kernel-bypassing tools result in unacceptable monetary cost. We propose Excalibur, a scalable and low-cost testing framework for DDoS defense solutions. The key idea is to leverage the programmable switch to perform testing tasks with Tbps-level scalability and low cost. Specifically, Excalibur coordinates both a server and a programmable switch to jointly perform testing tasks. It realizes flexible attack traffic generation, which requires a large number of resources, in the server while using the switch to increase the sending rate of attack traffic to Tbps-level. Our experiments on a 64×100Gbps Tofino switch show that Excalibur achieves orders-of-magnitude higher scalability and lower cost than existing tools. Xiang Chen 0017, Hongyan Liu 0001, Tingxin Sun, Qun Huang 0001, Dong Zhang 0010, Xuan Liu 0006, Haifeng Zhou, Chunming Wu 0001 |
INFOCOM | 5 |
| 2023 | Melody: Toward Resource-Efficient Packet Header Vector Encoding on Programmable SwitchesabstractThe programmable switch offers a limited capacity of packet header vector (PHV) words that store packet header fields and metadata fields defined by network functions. However, existing switch compilers employ inefficient strategies of encoding fields on PHV words. Their encoding wastes scarce PHV words and may result in failures when deploying network functions. In this paper, we propose Melody, a new framework that reuses PHV words for as many fields as possible to achieve resource-efficient PHV encoding. Melody offers a field analyzer and an optimization framework. The analyzer identifies which fields can reuse PHV words while preserving the original packet processing logic. The framework integrates analysis results into its encoding to offer the resource-optimal decisions. We evaluate Melody with production-scale network functions. Our results show that Melody reduces the consumption of PHV words by up to 85%. Xiang Chen 0017, Hongyan Liu 0001, Qingjiang Xiao, Jianshan Zhang, Qun Huang 0001, Dong Zhang 0010, Xuan Liu 0006, Chunming Wu 0001 |
INFOCOM | 6 |
| 2023 | A Network Function Virtualization Resource Allocation Model Based on Heterogeneous ComputingabstractWith the continuous increase in the speed and quantity of network traffic, higher performance requirements are put forward for the NFV system. Traditional virtualization technology is limited by slowing down of increase in CPU performance. There has some research on the use of hardware to accelerate network functions. However, existing method only consider use one hardware for acceleration, but hardware itself has limitation. It is difficult to match the network functions and hardware characteristics by considering the combination of each hardware and CPU discretely. In this paper, we propose a resource allocation model of network function virtualization(NFV) based on heterogeneous computing, which can maximize the resource utilization and obtain the global optimal solution. Our experiments prove that the genetic algorithm solution method we propose can take into account both the solution accuracy and the solution speed, and obtain an accurate Pareto curve under the multi-objectives optimization model. Hanze Chen, Lingfei Cheng, Longlong Zhu, Dong Zhang 0010 |
ISCC | 5 |
| 2023 | P4CTM: Compressed Traffic Pattern Matching Based on Programmable Data PlaneabstractPattern matching is an important technology applied to many security applications. Most network service providers choose to compress network traffic for better transmission, which brings the challenges of compressed traffic matching. However, existing works focus on improving the performance of uncompressed traffic matching or only realize the compressed traffic matching on end-host that can not keep pace with the dramatic increase in traffic. In this paper, we present P4CTM, a proof-of-concept method to conduct efficient compressed traffic matching on the programmable data plane. P4CTM uses the two-stage scan scheme to skip some bytes of compressed traffic, the 2-stride DFA combines with the compression algorithm to condense the state space, and the wildcard match to downsize the match action tables in the programmable data plane. The experiment indicates that P4CTM skips 83.10% bytes of compressed traffic, condenses the state space by order of magnitude, and reduces most of the table entries. Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001 |
ISCC | 5 |
| 2023 | DTRadar: Accelerating Search Process of Decision Trees in Packet ClassificationabstractPacket classification is an essential part of computer networks. Existing algorithms propose a partition process to address the memory explosion problem of the decision tree algorithm caused by the huge number of rules with multiple fields. However, the search process requires traversing multiple trees generated by the partition, which reduces the search efficiency. The existing algorithms take simple approaches to optimize the search process, which is low efficiency or high hardware overhead. In this paper, we propose DTRadar, a framework for expediting the decision tree packet lookup process. Its key idea is building an abstract One-Big-Tree(OBT) for multiple decision trees by establishing the middle data structure. DTRadar considers each decision tree as a splittable tree and organizes these subtrees by intermediate data structures. Extensive experiments show that DTRadar benefits existing decision tree-based solutions in classification time by 61.60%, and the memory footprint only increased by 4.21% on average. Jiashuo Yu, Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001 |
ISCC | 4 |
| 2023 | In-band Network Telemetry Manipulation Attacks and Countermeasures in Programmable NetworksabstractIn-band Network Telemetry (INT) is a widely used monitoring framework in modern large-scale networks that provides fine-grained visibility into network conditions by inserting telemetry data into packets. However, this mechanism also introduces new vulnerabilities that malicious attackers can exploit. In this paper, we present four In-band Network Telemetry Manipulation Attacks that take advantage of INT's weakness, demonstrating that attackers can cause severe damage with little effort by manipulating INT packets. To address this issue, we design SecureINT, a novel INT prototype that ensures confidentiality and integrity for INT packets. To meet the stringent computational requirements of programmable switches, we comprehensively analyze possible attacks on the deployed encryption/hash algorithms and modify them accordingly without compromising their security. According to the experiments, SecureINT can be deployed on programmable switches using a single pipeline, providing encryption and integrity verification for INT packets with minimal overhead. Dezhang Kong, Zhengyan Zhou, Yi Shen 0012, Xiang Chen 0017, Qiumei Cheng, Dong Zhang 0010, Chunming Wu 0001 |
IWQoS | 6 |
| 2023 | Vulnerabilities and Attacks of Inter-device Coordination in Programmable NetworksabstractIn programmable networks, some networking systems coordinate data plane switches to realize in-network functions (e.g., in-band network telemetry). However, the vulnerabilities of inter-device coordination are still largely unknown and neglected, which is highly concerning given the increasing popularity of this paradigm. In this paper, we identify three attack scenarios built upon such vulnerabilities, where attackers mislead the behaviors of networking systems that exploit inter-device coordination to execute in-network functions. We implement 20 existing networking systems on Tofino-based switches and a simulator, and attack these systems with the identified attacks. The experimental results indicate that our attacks significantly interfere with the normal operations of the selected networking systems, e.g., the cache hit rate of NetCache drops 38%. Our analysis also demonstrates that none of existing methods can fully mitigate our attacks since they fail to verify the packets for inter-device coordination. Hongyan Liu 0001, Xiang Chen 0017, Yi Shen 0012, Qun Huang 0001, Zhengyan Zhou, Dong Zhang 0010, Chunming Wu 0001 |
IWQoS | 6 |
| 2023 | Optimizing Program Deployment with libopl in Programmable NetworksabstractDeploying data plane programs on programmable switches involves complex optimization problems that make the optimal deployment decisions. However, existing deployment frameworks only focus on deploying programs in specific domains (i.e., supporting fixed optimization requirements), resulting in poor scalability. To this end, our goal is to simplify program deployment through general high-level abstractions that capture optimization requirements. In this paper, we present libopl, a generic library that enables administrators to express various optimization requirements when deploying programs and further calculates the optimal deployment plans. Existing frameworks can also use libopl to extend their functionalities to fit more deployment scenarios. To evaluate libopl, we build a Tofino-based testbed and a simulator. Our experimental results show that libopl exhibits comparable or better scalability than stateof-the-art frameworks and only introduces negligible overhead. Hongyan Liu 0001, Xiang Chen 0017, Yi Shen 0012, Dong Zhang 0010, Chunming Wu 0001 |
SECON | 4 |
| 2023 | Automatic Performance-Optimal Offloading of Network Functions on Programmable SwitchesabstractIn network function virtualization (NFV), network functions (NFs) are chained as a service function chain (SFC) to enhance NF management with low cost and high flexibility. Recent NFV solutions indicate that the packet processing performance of SFCs can be significantly improved by offloading NFs to programmable switches. However, such offloading requires a deep understanding of heterogeneous NF properties (e.g., NF resource consumption and NF performance behaviors) to achieve the maximum SFC performance. Unfortunately, none of existing solutions provide automatic analysis of these NF properties. Thus, network administrators have to manually examine the source codes of NFs and profile various NF properties by hand, which is extremely time-consuming and laborious. In this article, we propose LightNF, a novel system that simplifies NF offloading in programmable networks. LightNF automatically dissects comprehensive NF properties by means of code analysis and performance profiling while eliminating manual efforts. It then leverages its analysis results of NF properties in its SFC placement so as to make the performance-optimal offloading decisions. We have implemented LightNF on Tofino-based hardware programmable switches. We perform extensive experiments to evaluate LightNF with a real-world testbed and large-scale simulation. Our experiments show that LightNF outperforms existing solutions with an orders-of-magnitude reduction in per-packet processing latency and 9.5× improvement in SFC throughput. Xiang Chen 0017, Hongyan Liu 0001, Dong Zhang 0010, Zili Meng, Qun Huang 0001, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006, Qiang Yang 0004 |
IEEE Trans. Cloud Comput. | 3 |
| 2023 | Stalker Attacks: Imperceptibly Dropping Sketch Measurement Accuracy on Programmable SwitchesabstractDue to limited memory usage and provably high accuracy, sketches running on programmable switches have been commonly used by the literature for network measurement. However, their vulnerabilities are still largely unknown and neglected, which is highly concerning given the increasing popularity of network measurement. In this paper, we identify the Stalker attacks, where attackers aim to degrade the accuracy of sketches running on programmable switches. More precisely, attackers tamper with some sketch operations during sketch deployment atop programmable switches. At runtime, the tampered sketch will record highly inaccurate flow data, which degrades measurement accuracy. We implement Stalker attacks on Tofino switches. The results indicate that Stalker attacks significantly drop the accuracy of network management applications, e.g., reducing the F1 score of heavy hitter detection to zero. However, our analysis indicates that none of existing methods can detect Stalker attacks since they can hardly verify the correctness of sketch operations. Finally, we analyze potential defense mechanisms and identify challenges to enable further research in this context. Xiang Chen 0017, Hongyan Liu 0001, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006, Muhammad Khurram Khan |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2023 | Toward Low-Latency and Accurate State Synchronization for Programmable NetworksabstractProgrammable switches empower stateful packet processing, in which incoming packets continuously update states in the data plane, while applications in the control plane read and write states. However, since the data plane and control plane are separated, a consistent view of states in both planes is required for stateful packet processing. Existing approaches suffer from either high latency or low accuracy. In this paper, we propose ApproSync, a framework that offers approximate state synchronization with low latency and high accuracy. To achieve low latency, ApproSync directly transfers states between switch ASICs and the control plane by bypassing switch operating systems. To achieve high accuracy, ApproSync utilizes the resources in the switch ASIC to realize rate control in state synchronization, such that it avoids potential state loss. It also bounds the divergence between the states in the data plane and that in the control plane under limited link capacity. We prototype ApproSync on Barefoot Tofino switches. The experimental results indicate that compared to existing approaches, ApproSync achieves order-of-magnitude latency reduction while maintaining high accuracy of state synchronization. Also, our experiments demonstrate that ApproSync provides significant latency benefits to existing network management applications and well preserves high application-level accuracy. Xiang Chen 0017, Hongyan Liu 0001, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006, Qiang Yang 0004 |
IEEE/ACM Trans. Netw. | 4 |
| 2023 | Eliminating Control Plane Overload via Measurement Task PlacementabstractRecent efforts in network measurement place measurement tasks on programmable switches to measure high-speed traffic. These tasks extract flow data, i.e., events, from packets and send events to the control plane. However, the tasks may generate massive events in a short time. In this context, the links transferring events to the control plane and the control plane servers that handle events may be overloaded, i.e., control plane overload. None of existing solutions can eliminate control plane overload. In this paper, we propose MTP, a framework that eliminates control plane overload via careful measurement task placement. Our key idea is to allocate enough resources for each task during task placement to avoid control plane overload at runtime. For each task, MTP estimates its maximum possible rate of sending events to the control plane. Then its optimization framework addresses the resource restrictions of both switches and the control plane. The experiments on Tofino switches indicate that MTP outperforms existing solutions with higher accuracy in several use cases. Xiang Chen 0017, Hongyan Liu 0001, Dong Zhang 0010, Qun Huang 0001, Haifeng Zhou, Chunming Wu 0001, Qiang Yang 0004 |
IEEE/ACM Trans. Netw. | 3 |
| 2023 | Combination Attacks and Defenses on SDN Topology DiscoveryabstractThe topology discovery service in Software-Defined Networking (SDN) provides the controller with a global view of the substrate network topology, allowing for central management of the entire network. Unfortunately, emerging topology attacks can poison the network topology and result in unforeseeable disasters. Although researchers have made great efforts to mitigate this problem, security hazards still exist. In this paper, we propose Invisible Assailant Attack (IAA), the first combination topology attack capable of injecting and maintaining fake links even when 12 existing defense strategies are deployed simultaneously. IAA consists of 14 attack phases that apply multiple attack strategies. Attackers skillfully disguise the attack traffic in each phase so that it looks like normal network traffic, and perform these phases in a well-planned sequence, thereby bypassing existing defenses step by step. To mitigate this attack, we propose a Route Path Verification (RPV) mechanism that orchestrates multiple defense strategies to identify fake links. According to the experiments, RPV can successfully detect IAA with low overhead: its detection completes within 1 ms while its per-flow storage consumption is only a few KB. Dezhang Kong, Yi Shen 0012, Xiang Chen 0017, Qiumei Cheng, Hongyan Liu 0001, Dong Zhang 0010, Xuan Liu 0006, Shuangxi Chen, Chunming Wu 0001 |
IEEE/ACM Trans. Netw. | 6 |
| 2022 | TableGuard: A Novel Security Mechanism Against Flow Table Overflow Attacks in SDNabstractOne of the most important components of Software-Defined Networking (SDN) is the flow table. It receives flow rules from the controller and uses them to handle network traffic. However, a flow table can only store a few thousand flow rules, which makes it an attractive target for table overflow attacks. These attacks force the controller to populate the flow table with a large number of meaningless flow rules, which prevents normal flows from finding matching rules and therefore having to be reported to the controller. It results in a significant latency overhead, degrading the performance of the whole network. In this paper, we present a key characteristic of table overflow attacks: even though attackers can change some critical attack parameters (e.g., attack speed) to avoid detection, proactive flows from the attacked port always occupy a stable proportion in the flow table regardless of the attack form. In light of this finding, we propose TableGuard, a novel security mechanism that uses the proactive flow rule number as the detection metric and applies a statistical approach to help filter malicious flows. The experiments demonstrate that TableGuard can mitigate both high-rate and low-rate table overflow attacks. Compared with existing defenses, TableGuard has the best mitigation performance and the minimal overhead on normal flows. Dezhang Kong, Chunming Wu 0001, Yi Shen 0012, Xiang Chen 0017, Hongyan Liu 0001, Dong Zhang 0010 |
GLOBECOM | 6 |
| 2022 | KVLB: An In-network Key-Value Load Balancer using Multi-Valued HashabstractToday's Internet service architectures rely extensively on distributed key-value stores (KV-stores) to meet their performance requirements. One of the bottlenecks lies in the un-balanced load among key-value store nodes caused by the skewed workloads. With the flexibility and power of programmable switch ASICs, in-network computing becomes a propeller of application performance. This paper introduces KVLB, a new system that uses the programmable switch to achieve load balancing between key-value store nodes. KVLB uses selective replication of hot items and allocates replica node locations to the hot items through multi-value hash. This allows the switch to reroute the hot item to the replica node through a multi-valued hash calculation and requires fewer hardware resources for programmable switch ASICs. Our experimental results on an initial prototype show that KVLB improves the throughput of KV-stores at various degrees of skew and rely only on a small amount of switch hardware resources. Xikun Zheng, Dong Zhang 0010, Zhengyan Zhou, Jingwen Lv, Chunming Wu 0001 |
GLOBECOM | 2 |
| 2022 | Toward Low-Overhead Inter-Switch Coordination in Network-Wide Data Plane Program DeploymentabstractIn modern networks, administrators realize their desired functions such as network measurement in several data plane programs. They often employ the network-wide program deployment paradigm that decomposes input programs into match-action tables (MATs) while deploying each MAT on a specific programmable switch. Since MATs may be deployed on different switches, existing solutions propose the inter-switch coordination that uses the per-packet header space to deliver crucial packet processing information among switches. However, such coordination introduces non-trivial per-packet byte overhead, leading to significant end-to-end network performance degradation. In this paper, we propose Hermes, a program deployment framework that aims to minimize the per-packet byte overhead. The key idea of Hermes is to formulate the network-wide program deployment as a mixed-integer linear programming (MILP) problem with the objective of minimizing the per-packet byte overhead. In view of the NP hardness of the MILP problem, Hermes further offers a greedy-based heuristic that solves the problem in a near-optimal and timely manner. We have implemented Hermes on Tofino-based switches. Our experiments show that compared to existing frameworks, Hermes decreases the per-packet byte overhead by 156 bytes while preserving end-to-end performance in terms of flow completion time and goodput. Xiang Chen 0017, Hongyan Liu 0001, Qingjiang Xiao, Kaiwei Guo, Tingxin Sun, Xiang Ling 0001, Xuan Liu 0006, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
ICDCS | 9 |
| 2022 | SketchGuide: Reconfiguring Sketch-based Measurement on Programmable SwitchesabstractSketches enable efficient and fine-grained network measurement results with configurable resource-performance trade-offs. While sketch configurations are guided by theories, the current theoretical guidelines are either impractical or deficient for sketch configurations on emerging programmable switches. To better configure sketches on programmable switches, we (1) systematically analyze the limitations of sketch configuration guidelines on programmable hardware switches (i.e., unguided parameters, accuracy profiles, and resource budgets); (2) propose a generic and practical framework called SketchGuide to automate efficient sketch configurations on programmable switches; (3) implement SketchGuide on a Barefoot Tofino switch and compare SketchGuide to the state-of-the-art sketches by conducting extensive experiments. Our evaluations demonstrate that SketchGuide can automatically configure unguided parameters given resource budgets. SketchGuide reduces the hardware resource footprint by 52.92%-99.28% compared with current guidelines without impacting fidelity. Zhengyan Zhou, Jingwen Lv, Lingfei Cheng, Xiang Chen 0017, Tianzhu Zhang 0002, Qun Huang 0001, Jiayu Luo, Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001 |
ICNP | 9 |
| 2022 | Torp: Full-Coverage and Low-Overhead Profiling of Host-Side LatencyabstractIn data center networks (DCNs), host-side packet processing accounts for a large portion of the end-to-end latency of TCP flows. Thus, the profiling of host-side latency anomalies has been considered as a crucial part in DCN performance diagnosis and troubleshooting. In particular, such profiling requires full coverage (i.e., profiling every TCP packet handled by end-hosts) and low overhead (i.e., profiling should avoid high CPU consumption in end-hosts). However, existing solutions fully rely on end-hosts to implement host-side latency profiling, leading to low coverage or high overhead. In this paper, we propose Torp, a framework that offers full-coverage and low-overhead profiling of host-side latency. Our key idea is to offload profiling operations to top-of-rack (ToR) switches, which inherently offer full coverage and line-rate packet processing performance. Specifically, Torp selectively offloads profiling operations to the ToR switch based on switch limitations. It efficiently coordinates the ToR switch and end-hosts to execute the entire latency profiling task. We have implemented Torp on 32×100Gbps Tofino switches. Testbed experiments indicate that Torp achieves full coverage and orders of magnitude lower host-side overhead compared to other solutions. Xiang Chen 0017, Hongyan Liu 0001, Junyi Guo, Qun Huang 0001, Dong Zhang 0010, Chunming Wu 0001, Haifeng Zhou |
INFOCOM | 6 |
| 2022 | Escala: Timely Elastic Scaling of Control Channels in Network MeasurementabstractIn network measurement, data plane switches measure traffic and report events (e.g., heavy hitters) to the control plane via control channels. The control plane makes decisions to process events. However, current network measurement suffers from two problems. First, when traffic bursts occur, massive events are reported in a short time so that the control channels may be overloaded due to limited bandwidth capacity. Second, only a few events are reported in normal cases, making control channels underloaded and wasting network resources. In this paper, we propose Escala to provide the elastic scaling of control channels at runtime. The key idea is to dynamically migrate event streams among control channels to regulate the loads of these channels. Escala offers two components, including an Escala monitor that detects scaling situations based on realtime network statistics, and an optimization framework that makes scaling decisions to eliminate overload and underload situations. We have implemented a prototype of Escala on Tofino-based switches. Extensive experiments show that Escala achieves timely elastic scaling while preserving high application-level accuracy. Hongyan Liu 0001, Xiang Chen 0017, Qun Huang 0001, Dezhang Kong, Jinbo Sun, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
INFOCOM | 6 |
| 2022 | Libra: A Stateful Layer-4 Load Balancer with Fair Load DistributionabstractLayer-4 (L4) load balancers (LBs) are essential for data centers, dispatching incoming connections among thousands of servers. There are two critical requirements for L4 LBs: i) load balancing fairness, i.e., the ability to assign the load to servers in proportion to their capacity; ii) per-connection consistency (PCC), i.e., all the packets belonging to the same connection should be forwarded to the same server. Howbeit, existing LBs at large sacrifice load balancing fairness to mitigate PCC violations, which cannot satisfy both requirements in the meantime. In this paper, we present Libra, a stateful L4 LB that supports fair load distribution, PCC, memory efficiency, and resilience to resource depletion attacks. Libra makes load balancing decisions resorting to the proposed Weighted M-Least-Connection First (WMLCF) mechanism considering the real-time load and available processing capacity of servers, hence enabling load balancing fairness. We prototype Libra in a programmable software switch—BMv2 using P4 language and conduct extensive flow-level simulation to evaluate the performance. The evaluation indicates that Libra significantly improves load balancing fairness (over 95%), fully ensures PCC, and reduces the average flow completion time by 17.27–42.55% compared to existing mechanisms. Xingong Guo, Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001 |
IPCCC | 3 |
| 2022 | FROD: An Efficient Framework for Optimizing Decision Trees in Packet ClassificationabstractTo perform efficient packet classification, decision tree-based methods conduct decision trees via hand-tuned heuristics. Then the performance testing and optimization are executed to ensure an excellent searching speed and space overhead. Specifically, when the performance is below expectation, existing solutions attempt to optimize the algorithms, such as conducting more sophisticated heuristics. However, reconstruction or adjustment for algorithms produces an intolerable time overhead due to the long optimization period, caused by uncertain performance benefits and high pre-processing time. In this paper, we propose FROD, an efficient framework for optimizing the decision trees directly in packet classification. FROD raises a meticulous evaluation to accurately appraise decision trees constructed by different heuristics. It then seeks out the bottleneck components via a lightweight heuristic. After that, FROD searches the optimal division for inferior components considering structural constraints and characteristics of traffic distribution. Evaluation on ClassBench shows that FROD benefits existing decision tree-based solutions in classification time by 41% and memory footprint by 19% on average, and reduces classification time by up to 64%. Longlong Zhu, Jiashuo Yu, Jiayi Cai, Jinfeng Pan, Zhigao Li, Zhengyan Zhou, Dong Zhang 0010, Chunming Wu 0001 |
IWQoS | 7 |
| 2021 | MTP: Avoiding Control Plane Overload with Measurement Task PlacementabstractIn programmable networks, measurement tasks are placed on programmable switches to keep pace with high-speed traffic. At runtime, programmable switches send events to the control plane for further processing. However, existing solutions for task placement overlook the limitations of control plane resources. Thus, excessive events may overload the control plane. In this paper, we propose MTP, a system that eliminates control plane overload via careful task placement. For each task, MTP analyzes its structure to estimate its maximum possible rate of sending events to the control plane. Then it builds an optimization framework that addresses the resource restrictions of both switches and the control plane. We have implemented MTP on Barefoot Tofino switches. The experimental results indicate that MTP outperforms existing solutions with higher accuracy across four real use cases. Xiang Chen 0017, Qun Huang 0001, Peiqiao Wang, Hongyan Liu 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
INFOCOM | 6 |
| 2021 | LightNF: Simplifying Network Function Offloading in Programmable NetworksabstractIn network function virtualization (NFV), network functions (NFs) are chained as a service function chain (SFC) to enhance NF management with high flexibility. Recent solutions indicate that the processing performance of SFCs can be significantly improved by offloading NFs to programmable switches. However, such offloading requires a deep understanding of NF properties to achieve the maximum SFC performance, which brings non-trivial burdens to network administrators. In this paper, we propose LightNF, a novel system that simplifies NF offloading in programmable networks. LightNF automatically dissects comprehensive NF properties (e.g., NF performance behaviors) via code analysis and performance profiling while eliminating manual efforts. It then leverages the analyzed NF properties in its SFC placement so as to produce the performance-optimal offloading. We have implemented a LightNF prototype. Our experiments show that LightNF outperforms state-of-the-art solutions with an orders-of-magnitude reduction in per-packet processing latency and 9.5× improvement in SFC throughput. Xiang Chen 0017, Qun Huang 0001, Peiqiao Wang, Zili Meng, Hongyan Liu 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
IWQoS | 7 |
| 2021 | Machine learning based malicious payload identification in software-defined networkingabstractDeep packet inspection (DPI) has been extensively investigated in software-defined networking (SDN) as complicated attacks may intractably inject malicious payloads in the packets. Existing proprietary pattern-based or port-based third-party DPI tools can suffer from limitations in efficiently processing a large volume of data traffic. In this paper, a novel OpenFlow-enabled deep packet inspection (OFDPI) approach is proposed based on the SDN paradigm to provide adaptive and efficient packet inspection. First, OFDPI prescribes an early detection at the flow-level granularity by checking the IP addresses of each new flow via OpenFlow protocols. Then, OFDPI allows for deep packet inspection at the packet-level granularity: (i) for unencrypted packets, OFDPI extracts the features of accessible payloads, including tri-gram frequency based on Term Frequency and Inverted Document Frequency (TF–IDF) and linguistic features. These features are concatenated into a sparse matrix representation and are then applied to train a binary classifier with logistic regression rather than matching with specific pattern combinations. In order to balance the detection accuracy and performance bottleneck of the SDN controller, OFDPI introduces an adaptive packet sampling window based on the linear prediction; and (ii) for encrypted packets, OFDPI extracts notable features of packets and then trains a binary classifier with a decision tree, instead of decrypting the encrypted traffic to weaken user privacy. A prototype of OFDPI is implemented on the Ryu SDN controller and the Mininet platform. The performance and the overhead of the proposed solution are assessed using the real-world datasets through experiments. The numerical results indicate that OFDPI can provide a significant improvement in detection accuracy with acceptable overheads. Qiumei Cheng, Chunming Wu 0001, Haifeng Zhou, Dezhang Kong, Dong Zhang 0010, Junchi Xing |
J. Netw. Comput. Appl. | 5 |
| 2020 | SRA: Switch Resource Aggregation for Application Offloading in Programmable NetworksabstractProgrammable switches empower network applications with line-rate packet processing performance by allowing the offloading of applications. However, the resource of a programmable switch is extremely limited, which significantly limits the application offloading. Existing solutions to the problem either provide poor efficiency or suffer from accuracy drop. In this paper, we propose SRA, a system that loosens switch resource constraints for application offloading via resource aggregation. SRA provides administrators with an intuitive compiler directive to customize application offloading by resource aggregation. According to compiler directives, it automatically places the program on the substrate network, while maintaining original packet processing logics. We implement a prototype of SRA in P4, and establish an experimental testbed consisting of three 32$\times$100 Gbps Barefoot switches. The experimental results indicate that SRA enhances two real-world applications with sufficient resources while maintaining high performance. Hongyan Liu 0001, Xiang Chen 0017, Qun Huang 0001, Haifeng Zhou, Dong Zhang 0010, Chunming Wu 0001 |
GLOBECOM | 5 |
| 2020 | ApproSync: Approximate State Synchronization for Programmable NetworksabstractProgrammable switches empower stateful packet processing, in which incoming packets continuously update states in the data plane, while applications in the control plane read and write states. However, as the data plane and control plane are separated, a consistent view of states in both planes is required for stateful packet processing. Existing approaches suffer from either high latency or low accuracy. In this paper, we propose ApproSync, a framework that offers approximate state synchronization with low latency and high accuracy. To achieve low latency, ApproSync directly transfers states between switch ASICs and the control plane by bypassing switch operating systems. To achieve high accuracy, ApproSync utilizes the resources in the switch ASIC to realize rate control in state synchronization, such that it avoids potential state loss. It also bounds the divergence between the states in the data plane and that in the control plane under limited link capacity. We prototype ApproSync on Barefoot Tofino switches. The experimental results indicate that compared to existing approaches, ApproSync achieves order-of-magnitude latency reduction while maintaining high accuracy. Xiang Chen 0017, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
ICNP | 3 |
| 2020 | SPEED: Resource-Efficient and High-Performance Deployment for Data Plane ProgramsabstractProgrammable switches allow network administrators to customize packet processing behaviors in data plane programs. However, existing solutions for program deployment fail to achieve resource efficiency and high packet processing performance. In this paper, we propose SPEED, a system that provides resource-efficient and high-performance deployment for data plane programs. For resource efficiency, SPEED merges input data plane programs by reducing program redundancy. Then it abstracts the substrate network into an one big switch (OBS), and deploys the merged program on the OBS while minimizing resource usage. For high performance, SPEED searches for the performance-optimal mapping between the OBS and the substrate network with respect to network-wide constraints. It also maintains program logics among different switches via inter-device packet scheduling. We have implemented SPEED on a Barefoot Tofino switch. The evaluation indicates that SPEED achieves resource-efficient and high-performance deployment for real data plane programs. Xiang Chen 0017, Hongyan Liu 0001, Qun Huang 0001, Peiqiao Wang, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
ICNP | 5 |
| 2020 | Hidden Markov Model-based Load Balancing in Data Center NetworksabstractAbstract Modern data centers provide multiple parallel paths for end-to-end communications. Recent studies have been done on how to allocate rational paths for data flows to increase the throughput of data center networks. A centralized load balancing algorithm can improve the rationality of the path selection by using path bandwidth information. However, to ensure the accuracy of the information, current centralized load balancing algorithms monitor all the link bandwidth information in the path to determine the path bandwidth. Due to the excessive link bandwidth information monitored by the controller, however, much time is consumed, which is unacceptable for modern data centers. This paper proposes an algorithm called hidden Markov Model-based Load Balancing (HMMLB). HMMLB utilizes the hidden Markov Model (HMM) to select paths for data flows with fewer monitored links, less time cost, and approximate the same network throughput rate as a traditional centralized load balancing algorithm. To generate HMMLB, this research first turns the problem of path selection into an HMM problem. Secondly, deploying traditional centralized load balancing algorithms in the data center topology to collect training data. Finally, training the HMM with the collected data. Through simulation experiments, this paper verifies HMMLB’s effectiveness. Binjie He, Dong Zhang 0010 |
Comput. J. | 2 |
| 2019 | P4SC: Towards High-Performance Service Function Chain Implementation on the P4-Capable Device
Xiang Chen 0017, Dong Zhang 0010, Haifeng Zhou |
IM | 2 |
| 2019 | RL-Sketch: Scaling Reinforcement Learning for Adaptive and Automate Anomaly Detection in Network Data StreamsabstractWhen network is undergoing problems, such as DDoS attack, component failures, etc., the detection of heavy flows (e.g. heavy hitters and heavy changers) is much more critical. However, it has been increasing challenging to ensure the accurate detection of heavy flows while dealing with massive network traffic volume, diversified traffic distribution and the stringent memory requirement. Although recent research efforts like LD-Sketch are scalable for diverse network traffic, they depend on excessive memory to maintain high accuracy, such that they fail to work well when the memory is limited. We propose RL-Sketch, a adaptive sketch using reinforcement learning in detecting heavy flows. It predicts potential heavy flows based on the statistics of network traffic, to achieve both high accuracy and scalability with minor memory. Trace-driven evaluation shows that RL-Sketch achieves higher accuracy than state-of-the-art sketch-based technologies with up to 17.79× accuracy gain, while maintaining high robustness in extreme conditions. Zhengyan Zhou, Dong Zhang 0010, Xiaoyan Hong |
LCN | 2 |
| 2018 | MATReduce: Towards High-Performance P4 Pipeline by Reducing Duplicate Match OperationsabstractP4 provides operators with the ability to program the packet processing pipeline of the data plane device. The match-action table (MAT) is a basic component of the P4 pipeline that matches the packet and performs an action on the matched packet. However, different MATs may execute duplicate match operations that decreases the performance of the P4 pipeline. To this end, we present MATReduce, a framework that optimizes the P4 pipeline by reducing duplicate match operations between MATs. MATReduce is composed of two key components, the preprocessor and the runtime manager. By introducing the compound MAT and rewriting the P4 control flow, the preprocessor merges duplicate match operations of the P4 pipeline while maintaining the program semantics. At runtime, the runtime manager converts user rules to actual rules for maintaining the policy consistency. Our preliminary experimental results show that MATReduce provides significant performance improvement, including a 23.90% throughput increase and a 34.37% delay decrease on the software target, and a 45.19% delay decrease on the hardware target. Xiang Chen 0017, Dong Zhang 0010, Haifeng Zhou |
GLOBECOM | 2 |
| 2018 | PAME: Evolutionary membrane computing for virtual network embedding
Chunyan Yu, Qi Lian, Dong Zhang 0010, Chunming Wu 0001 |
J. Parallel Distributed Comput. | 3 |