VLDB 2026 Research / reviewers in the wild / expert
Xiang Chen 0017
dblp:64/3062-17
· DBLP profile ↗
65ranked-venue papers
25as first author
60since 2021 · last 2026
0000-0002-0249-9664ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 49 · 21 first-author · 44 since 2021Security and privacy · 7 · 1 first-author · 7 since 2021Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Potent but Stealthy: Rethink Profile Pollution Against Sequential Recommendation via Bi-Level Constrained Reinforcement ParadigmabstractSequential Recommenders, which exploit dynamic user intents through interaction sequences, are vulnerable to adversarial attacks. While existing attacks primarily rely on data poisoning, they require large-scale user access or fake profiles thus lacking practicality. In this paper, we focus on the Profile Pollution Attack (PPA) that subtly contaminates partial user interactions to induce targeted mispredictions. Previous PPA methods suffer from two limitations, i.e., i) over-reliance on sequence horizon impact restricts fine-grained perturbations on item transitions, and ii) holistic modifications cause detectable distribution shifts. To address these challenges, we propose a constrained reinforcement driven attack CREAT that synergizes a bi-level optimization framework with multi-reward reinforcement learning to balance adversarial efficacy and stealthiness. We first develop a Pattern Balanced Rewarding Policy, which integrates pattern inversion rewards to invert critical patterns and distribution consistency rewards to minimize detectable shifts via unbalanced co-optimal transport. Then we employ a Constrained Group Relative Reinforcement Learning paradigm, enabling step-wise perturbations through dynamic barrier constraints and group-shared experience replay, achieving targeted pollution with minimal detectability. Extensive experiments demonstrate the effectiveness of CREAT. Jiajie Su, Zihan Nan, Yunshan Ma 0002, Xiaobo Xia, Xiaohua Feng 0002, Weiming Liu 0005, Xiang Chen 0017, Chaochao Chen 0001 |
AAAI | 7 |
| 2026 | Turbolearn: Harnessing Accurate and Line-Rate Deep Learning on Programmable Switches
Zhifan Jiang, Longlong Zhu, Jiashuo Yu, Linying Zheng, Chunming Wu 0001, Xiang Chen 0017 |
APNet | 6 |
| 2026 | TurboLearn: Harnessing Accurate and Line-Rate Deep Learning on Programmable SwitchesabstractThe intelligent data plane (IDP) embeds deep learning (DL) models on switches for line-rate traffic analysis, but hardware constraints often force simplified models, reducing accuracy, while complex models like Transformers remain undeployable. We present TurboLearn, which achieves high accuracy and line-rate performance by co-designing inference across the switch ASIC and switch OS on the same switch. TurboLearn uses three key techniques: (1) a hardware fast path for lightweight classification and a software normal path for complex models, (2) confidence-based selective inference that escalates only low-confidence packets, and (3) confidence-calibrated knowledge distillation, where the normal path teaches the fast path. On Intel Tofino2 switches across three real-world traffic tasks, TurboLearn supports models that existing IDPs cannot deploy, improves macro-F1 by up to 31.31%, and keeps over 90% of traffic on the fast path with zero throughput loss. Zhifan Jiang, Longlong Zhu, Jiashuo Yu, Linying Zheng, Chunming Wu 0001, Xiang Chen 0017 |
APNet | 6 |
| 2026 | MonPlan: Taming Network Measurement with Accurate and Resource-Efficient Sketch-INT Co-Design
Xiang Chen 0017, Linying Zheng, Longlong Zhu, Zedi Chen, Qing Shu, Jialu Tian, Siqi Dong, Qun Huang 0001, Jianshan Zhang, Xuan Liu 0006, Haifeng Zhou, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001 |
INFOCOM | 1 |
| 2026 | LTD: Low-Overhead Topology Discovery using Programmable Data Planes
Dezhang Kong, Minghao Li 0012, Shi Lin, Zhenhua Xu 0004, Longlong Zhu, Linying Zheng, Xiang Chen 0017, Changting Lin, Xuan Liu 0006, Dong Zhang 0010, Chunming Wu 0001 |
INFOCOM | 7 |
| 2026 | PSM: Timely and Resource-Efficient Sketch Migration in Network Measurement
Hongyan Liu 0001, Xiang Chen 0017, Zhengyan Zhou, Di Wang 0003, Chunming Wu 0001 |
IWQoS | 3 |
| 2026 | SketchPipe: Toward Accurate Sketch-based Network Measurement on Multi-Pipeline Switches with Splitless Sketch Placement
Xiang Chen 0017, Longlong Zhu, Linying Zheng, Hongyang Du 0001, Dong Zhang 0010, Jianshan Zhang, Xuan Liu 0006, Qun Huang 0001, Dusit Niyato, Haifeng Zhou, Chunming Wu 0001, Hongyan Liu 0001, Kui Ren 0001 |
NSDI | 1 |
| 2026 | Proteus: Towards Accurate and Low-overhead In-Network Malicious Traffic DetectionabstractNetwork intrusion detection systems (NIDS) are essential for web security by identifying and dropping malicious traffic. Existing in-network NIDS leverage the Tbps-level packet processing capability of programmable switches to achieve high-speed flow classification. They translate complex trained machine learning models to decision trees (DTs), where DTs are deployed on programmable switches via single-DT or multiple-DT deployment. However, they face a fundamental trade-off: single-DT deployment suffers from low classification accuracy due to over-pruning of trees, while multiple-DT deployment suffers from high overhead due to deploying multiple tree replicas. In this paper, we propose Proteus, an in-network malicious traffic detection system that achieves both high classification accuracy and low overhead. Its key idea is to split the original DT into critical and normal sub-trees, where these sub-trees have different impacts on overall accuracy. More precisely, Proteus first splits a DT into one critical and several normal sub-trees for adapting to the accuracy requirement and switch resource budgets. Second, it minimizes coordination overhead between sub-trees while ensuring full flow coverage via mixed-integer linear programming. Third, it dynamically reallocates or migrates sub-trees to adapt to changing resources by monitoring both classification accuracy and switch resource changes. Testbed experiments with 12.8 Tbps programmable switches show that Proteus improves classification accuracy, reduces switch resource consumption, and reduces classification latency. Longlong Zhu, Linying Zheng, Qing Shu, Zedi Chen, Jiashuo Yu, Shaopeng Zhou, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001, Xiang Chen 0017 |
WWW | 11 |
| 2026 | ADMM-Based Adversarial False Data Injection Attacks Against Multi-Label Locational DetectionabstractWhile multi-label learning has shown excellent performance in False Data Injection Attack (FDIA) locational detection, it has also exposed some potential security risks and vulnerabilities. However, unlike the image domain, the vulnerabilities of multi-label learning in the field of power grid have just received attention and urgently need to be explored and addressed. In this paper, to achieve a better understanding for the security risks of deep learning-based multi-label FDIA detectors, we propose two Alternating Direction Method of Multipliers (ADMM) based adversarial attacks, which are applicable to two different scenarios. The proposed two ADMM-based attacks aim to reduce additional attack costs while seeking suitable adversarial perturbations, making the attacks more realistic and feasible. The experimental results verify the effectiveness of the proposed ADMM-based attacks, making noteworthy strides in fostering a profound comprehension of the vulnerabilities in the unique field of deep multi-label learning for power systems. Jiwei Tian, Chao Shen 0001, Chenhao Lin, Meng Zhang 0011, Xiaofang Xia, Chao Ren 0006, Peican Zhu, Chunming Wu 0001, Xiang Chen 0017 |
IEEE Trans. Dependable Secur. Comput. | 9 |
| 2026 | Toward Security-Enhanced In-Band Network Telemetry in Programmable NetworksabstractIn-band Network Telemetry (INT) is a widely used monitoring framework in modern large-scale networks. It provides packet-level visibility into network conditions by inserting telemetry data into packets, enabling unprecedented fine-grained network management. However, this mechanism also introduces new vulnerabilities that malicious attackers can exploit. In this paper, we present eight In-band Network Telemetry Manipulation Attacks that take advantage of INT’s weakness, demonstrating that attackers can cause severe damage with little effort by manipulating INT packets. To address this issue, we designed SecureINT, a security-enhanced INT prototype that provides encryption and integrity verification for INT packets. Specifically, SecureINT deploys Even-Mansour and SipHash for confidentiality and integrity, respectively. It also uses a zero-delay rotation mechanism, which enables administrators to dynamically change the version of the deployed Even-Mansour/SipHash running on programmable switches without the need to re-install new programs. In this way, SecureINT can provide lasting security for INT packets using the limited resources of programmable switches. According to the experiments, SecureINT can be deployed on programmable switches using a single pipeline. Besides, the overhead of the rotation mechanism running on the control plane is still minimal. Dezhang Kong, Xiang Chen 0017, Zhengyan Zhou, Yi Shen 0012, Hongyan Liu 0001, Qiumei Cheng, Xuan Liu 0006, Dong Zhang 0010, Chunming Wu 0001, Muhammad Khurram Khan |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2026 | HiMon: Achieving Low-Cost and High-Accuracy Network Monitoring via Hierarchical SketchingabstractAs data centers continue to expand in size and complexity, obtaining global traffic insights necessitates aggregating statistical data from numerous individual nodes, a process critical for effective network management. However, in data centers, existing approaches often rely on querying individual endpoint hosts to gather cluster-wide statistics, which introduces substantial latency and reduces efficiency, particularly in large-scale deployments. To address this issue, we propose HiMon, a cost-efficient and high-accurate distributed monitoring system for optimizing traffic aggregation. HiMon enables distributed nodes to perform real-time, flow-level statistical processing and report the data to a master node with minimal bandwidth consumption. The master node aggregates the collected data to construct a comprehensive global traffic view. To enable high-speed and high-precision perpacket processing on child nodes, we introduce MaxSketch. MaxSketch’s data structure and update strategy allow it to accurately estimate child node traffic with minimal memory and computational overhead. For high-speed aggregation on the master node, we present PolySketch, which significantly boosts aggregation efficiency by delegating most computational tasks to the child nodes. Together, the hierarchical sketch structures of MaxSketch and PolySketch form the HiMon monitoring system. Experimental evaluations demonstrate that HiMon surpasses baseline algorithms, achieving a 17-210× improvement in traffic processing efficiency, a 25-42× reduction in master node bandwidth consumption, and a 3.69-8.97× increase in accuracy. Zhenyu Wen, Shibo He, Xiang Chen 0017, Chaojie Gu, Jiming Chen 0001 |
IEEE Trans. Netw. | 4 |
| 2025 | AIA: Autoregression-Based Injection Attacks Against Text2SQL ModelsabstractTo facilitate understanding of users' diverse queries against the back-end databases in web applications, researchers have introduced Text-to-SQL (Text2SQL) models that can generate well-structured SQL queries from users' query texts in natural language. As the Text2SQL model decouples the user queries with the back-end databases, it inherently mitigates the SQL injection risk posed by inserting users' input into pre-written SQL queries. However, what security risks to web applications may be posed by Text2SQL models remains an open question. In this paper, we present a new attack framework, named Autoregression-based Injection Attacks (AIA), to evaluate the security risks of Text2SQL models. In particular, AIA makes target models generate attack payloads by constructing specific inputs and adjusting the input auto-regressively. Our evaluation demonstrates that AIA can cause Text2SQL models to generate target output by adversarial inputs with success rates of over 70% in most scenarios. The generated adversarial input has certain transferability in target Text2SQL models. Additionally, practice experiments show that AIA can make Text2SQL models extract user lists from databases and even delete data in databases directly. Deyin Li, Xiang Ling 0001, Changjiang Li, Xiang Chen 0017, Chunming Wu 0001 |
AAAI | 4 |
| 2025 | Phantom: Virtualizing Switch Register Resources for Accurate Sketch-based Network MeasurementabstractSketches have proven to be useful for measuring traffic. They store measurement results in the registers of data plane switches. However, they suffer from the short of switch register resources, limiting their measurement accuracy. Xiang Chen 0017, Hongyan Liu 0001, Zhengyan Zhou, Wenbin Zhang 0011, Hongyang Du 0001, Dong Zhang 0010, Xuan Liu 0006, Haifeng Zhou, Dusit Niyato, Qun Huang 0001, Chunming Wu 0001, Kui Ren 0001 |
EuroSys | 1 |
| 2025 | Carrera: Enabling High-Performance eBPF-based Sketches in Network MeasurementabstractTo achieve dynamic network measurement, trends build sketches on eBPF to avoid service interruptions. However, existing eBPF-based sketches suffer from high CPU consumption, leading to poor throughput and high latency and making them hard to measure high-speed traffic. Optimizing their performance requires users to refactor codes based on each sketch’s characteristics on eBPF, which is highly complex and time-consuming.In this paper, we argue that users should write sketches without concerning low-level eBPF performance optimizations, with the deployment automatically activating cross-sketch performance optimizations. We present Carrera, a library that offers domain-specific optimizations for eBPF-based sketches. Our contributions are (1) systematically analyzing the performance bottlenecks of eBPF-based sketches through microbenchmarks, (2) identifying practical optimizations, including hardware offloading, SIMD-accelerated hashing, traffic-aware flow index caching, prefetched randomization, and active data collection, to address the identified bottlenecks in eBPF-based sketches, (3) evaluating these optimizations with state-of-the-art sketches and demonstrating that Carrera improves throughput by up to 65% and reduces latency by up to 93% via testbed experiments. Xiang Chen 0017, Xin Yao 0008, Longlong Zhu, Linying Zheng, Hongyan Liu 0001, Jianshan Zhang, Dong Zhang 0010, Xuan Liu 0006, Qun Huang 0001, Haifeng Zhou, Chunming Wu 0001 |
ICNP | 1 |
| 2025 | EffiMatch: Enabling Fast and Accurate Learning-based Packet ClassificationabstractLearning-based Packet Classification methods reduce memory overhead by using lightweight Recursive Model Index(RMI) structures to limit the search range, followed by linear matching. However, they face a trade-off: complex RMI structures achieve smaller search ranges but slow down lookup, while simpler ones are faster but require larger scans. In this paper, we propose EffiMatch, a parallel multi-model lookup architecture aimed at resolving the trade-off between RMI complexity and linear search range in learning-based index systems. We propose two key designs: 1) We design a partitioning strategy called Distribution-Distance Partitioning (DDP), which groups data points with similar trends into the same segment. Combined with parallel lookup, this reduces the linear search range while maintaining high lookup speed. 2) We propose a more fine-grained binarization method, Base-Index Representation (BI), which approximates floating-point operations using integers. This method further reduces the search range without increasing model complexity. Experimental results show that EffiMatch reduces the linear search range by 26.84% using lower-complexity RMI models, which improves lookup speed by up to 6× and reduces construction time by up to 4 orders of magnitude compared to state-of-the-art LPC methods. Lida Liao, Jiashuo Yu, Longlong Zhu, Hongyan Liu 0001, Dong Zhang 0010, Xiang Chen 0017, Chunming Wu 0001 |
ICNP | 8 |
| 2025 | P4-IDet: A Programmable Switch-Based Framework for Real-Time and High-Accuracy Traffic Anomaly Detection in ICPSsabstractThe rise of Industry 4.0 exposes traditionally isolated Industrial Cyber-Physical Systems (ICPSs) to increasing network attacks, posing serious security threats and potential damage. Traffic anomaly detection is essential for identifying such attacks. Nevertheless, existing work faces a dilemma between high accuracy and real-time performance. In this paper, we resolve this dilemma through P4-IDet, a novel traffic anomaly detection framework based on programmable switches, achieving both high accuracy and real-time performance. P4-IDet first deploys a low-complexity detector in the data plane to stamp timestamps, extract traffic features, and perform line-rate preliminary detection. Only suspicious packets and their features are uploaded to a server for fine-grained analysis by a high-accuracy machine learning model. To further reduce the upload and accelerate detection, a Bayesian optimizer adaptively tunes detection rules based on differences between detection results of the switch and the server. Moreover, P4-IDet can be integrated with existing detection models to enhance accuracy and real-time performance. Finally, we implement the prototype on a Barefoot Tofino 2.0 switch using the P4 language and an x86 server, and validate it on a large-scale ICPS platform with real-world industrial systems. Experiments show 5.6–41.1% accuracy gains, a 28.93% reduction in machine learning model workload, and 8.31–25.90% improvements in real-time performance. Jiayu Luo, Zhengyan Zhou, Qiaoxiong Tang, Ruohan Chen, Xiang Chen 0017, Chao Pei, Qiang Yang 0004, Wenhai Wang, Haifeng Zhou |
IECON | 5 |
| 2025 | TurboCache: Empowering Switch-Accelerated Key-Value Caches with Accurate and Fast Cache UpdatesabstractRecent key-value (KV) caches are offloaded to programmable switches to offer high query processing performance. However, they suffer from both low accuracy in hot key detection and high latency in cache updates due to the strict limitations on switch registers. We propose TurboCache, a switch-accelerated KV cache with accurate hot key detection and fast cache updates. Our key idea is to leverage the switch recirculation capability to build a novel data structure that caches hot KV pairs. With this hardware-compatible cache data structure, TurboCache designs efficient data plane algorithms that accurately detects new hot keys and quickly updates its cache entirely within switch ASIC pipelines. We have implemented TurboCache on a${64}\times {100}$Gbps Tofino switch. Testbed results indicate that TurboCache improves the hot key detection accuracy and decreases the cache update latency of existing KV caches by several orders of magnitude. Xiang Chen 0017, Longlong Zhu, Linying Zheng, Lingfei Cheng, Jianshan Zhang, Xu Yang 0002, Dong Zhang 0010, Xuan Liu 0006, Xiaoming Lu, Xun Yi, Ibrahim Khalil 0001, Albert Y. Zomaya, Haifeng Zhou, Chunming Wu 0001 |
INFOCOM | 1 |
| 2025 | Scaling Learning-based Packet Classification Hardware with NeuTree
Jiashuo Yu, Longlong Zhu, Linying Zheng, Dong Zhang 0010, Xiang Chen 0017 |
INFOCOM | 7 |
| 2025 | Monica: Towards Scalable Distributed System Verification by Programmable Switch-Based TestingabstractData correctness in distributed systems is ensured by data consistency, where consistency is achieved by consensus algorithms. To safeguard data consistency, current testing tools use stress testing methods to examine consensus algorithms. However, existing tools are unable to simulate the situation under high traffic and suffer from excessive verification time. In this paper, we propose Monica, a scalable and efficient verification framework. Its key idea is to leverage the programmable switch to verify consensus algorithms. Specifically, Monica provides a set of primitives that researchers can invoke. Then, the control server recognizes the primitives and automatically configures the data plane. After that, the programmable switch collaborates with the control server to complete the verification. Experimental results show that Monica can generate traffic at the rate of Tbps level while keeping the computational and memory consumption of the programmable switch under 11.87%. Compared to existing testing tools, Monica increases the verification speed by up to 3.13 times. Further, Monica improved accuracy by 35.71% in high-traffic scenarios over other tools. Jiashuo Yu, Longlong Zhu, Dong Zhang 0010, Lida Liao, Rongbang Wu, Xiang Chen 0017, Chunming Wu 0001 |
IWQoS | 8 |
| 2025 | Polyx: Accelerating Verification of Traffic Migration in Large-Scale BGP NetworksabstractIn BGP networks, traffic migration verification ensures the scalability and reliability of the network during configuration changes. However, previous approaches suffer from low scalability and high computational overhead. In this poster, we propose Polyx, a framework for accelerating verification of traffic migration in large-scale BGP networks. Its key idea is to leverage hardware parallelism with a deterministic serialization algorithm to enhance state machine techniques. We implement the Polyx prototype and evaluate it on our built testbed. The experimental results demonstrate that Polyx achieves up to 46× overall speedup, 36× in state machine construction, and 131× in equivalence verification with minimal FPGA resource usage. Rongbang Wu, Longlong Zhu, Jiashuo Yu, Dong Zhang 0010, Hongyan Liu 0001, Zongye Lin, Lida Liao, Xiang Chen 0017, Chunming Wu 0001 |
IWQoS | 10 |
| 2025 | Handling Data Plane Program Deployment Dynamics with High-Quality Generative Diffusion ModelsabstractDeploying data plane programs across the network is typically formulated as a mixed-integer programming task, leading to a long execution time. In response, existing studies carefully tailor heuristics for specific task properties such as objectives. However, they suffer from poor solution quality under dynamic task deployment since they overfit specific task properties. Recently, generative diffusion models have been widely adopted in network optimizations due to their strong adaptability and generalization. Accordingly, in this poster, we propose a diffusion model-based framework for data plane program deployment tasks. Our key idea is to leverage the reverse denoising process of diffusion models to react to dynamic task changes at runtime while maintaining high solution quality. Preliminary results on our testbed show that we reduce latency by 66.67% and resource overhead by 58.62% during dynamic deployment. Longlong Zhu, Jiashuo Yu, Xiang Chen 0017, Qing Shu, Zedi Chen, Zhifan Jiang, Qun Huang 0001, Xuan Liu 0006, Dong Zhang 0010, Chunming Wu 0001 |
IWQoS | 3 |
| 2025 | NDIF: A distributed framework for efficient in-network neural network inference
Shengrui Lin, Shaowei Xu, Binjie He, Hongyan Liu 0001, Dezhang Kong, Xiang Chen 0017, Dong Zhang 0010, Chunming Wu 0001, Ming Li 0056, Xuan Liu 0006, Yuqin Wu, Muhammad Khurram Khan |
Comput. Secur. | 6 |
| 2025 | Elastically Scaling Control Channels in Network Measurement With EscalaabstractIn network measurement, data plane switches measure traffic and report events (e.g., heavy hitters) to the control plane via control channels. The control plane makes decisions to process events. However, current network measurement suffers from two problems. First, when traffic bursts occur, massive events are reported in a short time so that the control channels may be overloaded due to limited bandwidth capacity. Second, only a few events are reported in normal cases, making control channels underloaded and wasting network resources. In this paper, we propose$\textsf {Escala}$to provide the elastic scaling of control channels at runtime. The key idea is to dynamically migrate event streams among control channels to regulate the loads of these channels.$\textsf {Escala}$offers two components, including an$\textsf {Escala}$monitor that detects scaling situations based on realtime network statistics, and an optimization framework that makes scaling decisions to eliminate overload and underload situations. We have implemented a prototype of$\textsf {Escala}$on Tofino-based switches. Extensive experiments show that$\textsf {Escala}$achieves timely elastic scaling while preserving high application-level accuracy. Hongyan Liu 0001, Xiang Chen 0017, Qun Huang 0001, Dezhang Kong, Dong Zhang 0010, Chunming Wu 0001, Xuan Liu 0006 |
IEEE Trans. Netw. | 2 |
| 2025 | Toward Secure Inter-Device Coordination in Programmable NetworksabstractIn programmable networks, some networking systems coordinate data plane switches to perform in-network functions (e.g., in-band network telemetry). However, the vulnerabilities associated withinter-device coordinationremain largely unexplored and overlooked, which is highly concerning given the increasing popularity of this paradigm. In this paper, we identify three attack scenarios built upon such vulnerabilities, where attackers mislead the behaviors of networking systems. We implement 20 networking systems on Tofino-based switches and a simulator and test them against the identified attacks. Our experimental results show that our attacks severely disrupt the normal operation of these networking systems, e.g., the cache hit rate of NetCache drops by 38%. However, our analysis reveals that none of existing methods fully mitigate our attacks because they fail to verify the packets for inter-device coordination. To this end, we select characteristics from existing methods while addressing their limitations to design effective mitigation methods. Experimental results indicate that our methods perform well in mitigating our attacks and introduce acceptable overheads. Hongyan Liu 0001, Xiang Chen 0017, Di Wang 0049, Qun Huang 0001, Dong Zhang 0010, Chunming Wu 0001, Xuan Liu 0006 |
IEEE Trans. Netw. | 2 |
| 2024 | PP-Stream: Toward High-Performance Privacy-Preserving Neural Network Inference via Distributed Stream ProcessingabstractPrivacy preservation is critical for neural network inference, which often involves collaborative execution of different parties to make predictions on sensitive data based on sensitive neural network models. However, the expensive cryptographic operations of privacy preservation also pose performance chal-lenges to neural network inference. We address this performance-security tension by designing PP-Stream, a distributed stream processing system for high-performance privacy-preserving neural network inference. PP-Stream adopts hybrid privacy-preserving mechanisms for linear and non-linear operations of neural network inference. It treats inference data as real-time data streams, and parallelizes the inference operations across multiple pipelined stages that are executed by multiple servers and threads. It also solves the load-balanced resource allocation across servers and threads as an optimization problem. We prototype PP-Stream and show via testbed experiments that it achieves low inference latencies on various neural network models. Qingxiu Liu, Qun Huang 0001, Xiang Chen 0017, Sa Wang, Shujie Han 0001, Patrick P. C. Lee |
ICDE | 3 |
| 2024 | SpotMon: Enabling General Hotspot Monitoring in Key-Value StoresabstractKey-value stores are essential to online services such as e-commerce. In key-value stores, a hotspot (i.e., frequently accessed items) may cause severe load imbalances, high response latency, and Service Level Agreement (SLA) violations. However, existing works only focus on specific types of hotspots, thus overlooking other types of hotspots and leading to blind spots. In this paper, we propose SpotMon, a system that enables general hotspot monitoring in key-value stores. Specifically, we (1) systematically identify the generality requirements of hotspot monitoring from existing works, (2) formulate general hotspot monitoring as an arbitrary partial spot query problem, (3) measure the hotness of hotspot candidates with a new vector expression, (4) propose hotspot encoding, filtering, decoding, and querying to support general queries without focusing on specific hotspots, (5) leverage the in-network visibility of programmable switches to identify system-wide hotspots. Our extensive experiments indicate that SpotMon provides high accuracy (e.g., F1 score from 0.88 to 1) and enables efficient hotspot mitigations (e.g., up to$4.03 \times$MQPS). Zhengyan Zhou, Jinhan Zu, Enhao Huang, Haifeng Zhou, Dong Zhang 0010, Xiang Chen 0017, Chunming Wu 0001 |
ICNP | 7 |
| 2024 | Accelerating Sketch-based End-Host Traffic Measurement with Automatic DPU OffloadingabstractSketch-based traffic measurement is a crucial building block for monitoring traffic statistics and ensuring the quality of services of end-host applications. However, existing approaches for building sketches in end-hosts exhibit poor packet processing performance or high CPU consumption. In this paper, we propose MPU, which automatically offloads sketch-based measurement to the emerging hardware, DPU. MPU consists of a sketch analyzer that profiles sketch resource consumption and an optimization framework that formulates the offloading problem and maximizes sketch performance on DPU. We implement MPU on the NVIDIA BlueField DPU. Our testbed results indicate that MPU achieves 85% lower per-packet processing latency and 47% higher traffic measurement accuracy when compared to existing approaches. Xiang Chen 0017, Wenbin Zhang 0011, Xin Yao 0008, Zizheng Wang, Hongyan Liu 0001, Qun Huang 0001, Xuan Liu 0006, Haifeng Zhou, Chunming Wu 0001 |
INFOCOM | 1 |
| 2024 | Eagle: Toward Scalable and Near-Optimal Network-Wide Sketch Deployment in Network MeasurementabstractSketches are useful for network measurement thanks to their low resource overheads and theoretically bounded accuracy. However, their network-wide deployment suffers from the trade-off between optimality and scalability: (1) Most solutions rely on mixed integer linear programming (MILP) solvers to provide the optimal decisions. But they are time-consuming and can hardly scale to large-scale deployment scenarios. (2) While heuristics achieve scalability, they deteriorate resource and performance overheads. We propose Eagle, a framework that achieves scalable and near-optimal network-wide sketch deployment. Our key idea is to decompose network-wide sketch deployment into sub-problems. Such decomposition allows Eagle to (1) simultaneously optimize switch resource consumption and end-to-end performance (retaining optimality), and (2) incorporate time-saving techniques into sub-problem solving (achieving scalability). Compared to existing solutions, Eagle improves scalability by up to 255× with negligible loss of optimality. It has also saved administrators in a production network days of efforts and reduced the operation time from O(hour) to O(second). Xiang Chen 0017, Qingjiang Xiao, Hongyan Liu 0001, Qun Huang 0001, Dong Zhang 0010, Xuan Liu 0006, Longbing Hu, Haifeng Zhou, Chunming Wu 0001, Kui Ren 0001 |
SIGCOMM | 1 |
| 2024 | Terra: Low-latency and reliable event collection in network measurement
Hongyan Liu 0001, Xiang Chen 0017, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006, Muhammad Khurram Khan |
J. Netw. Comput. Appl. | 3 |
| 2024 | rDefender: A Lightweight and Robust Defense Against Flow Table Overflow Attacks in SDNabstractThe flow table is a critical component of Software-Defined Networking (SDN). However, flow tables’ limited capacity makes them highly vulnerable to flow table overflow attacks (FTOAs). Due to the low attack cost and highly flexible attack forms, it is hard to eradicate FTOAs. This paper addresses three unsolved problems for table security and proposes a robust defense accordingly. First, we reveal that the existing defenses with fixed defense speeds will cause severe packet loss when handling diverse traffic. We prove that deleting multiple rules can efficiently solve this problem and give a rigorous derivation to calculate the suitable deletion number according to the environment. Second, we illustrate that abnormal table occupancy squeezing is a constant characteristic of FTOAs regardless of attack forms. It can be used to identify attacked ports accurately in different scenarios. Third, we mathematically prove that random deletion can guarantee the continuous decrease of malicious flow rules after confirming attacked ports. It achieves fast speed and robust effectiveness in different environments. Based on these findings, we design rDefender, a robust and lightweight defense prototype. We evaluate its effect by designing diverse, powerful attacks and using real-world datasets and topology. The results demonstrate that it achieves the best overall performance compared to six existing mainstream defenses, providing stable security for switch flow tables. Dezhang Kong, Xiang Chen 0017, Chunming Wu 0001, Yi Shen 0012, Zhengyan Zhou, Qiumei Cheng, Xuan Liu 0006, Yubing Qiu, Dong Zhang 0010, Muhammad Khurram Khan |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | AdvSQLi: Generating Adversarial SQL Injections Against Real-World WAF-as-a-ServiceabstractAs the first defensive layer that attacks would hit, the web application firewall (WAF) plays an indispensable role in defending against malicious web attacks like SQL injection (SQLi). With the development of cloud computing, WAF-as-a-service, as one kind of Security-as-a-service, has been proposed to facilitate the deployment, configuration, and update of WAFs in the cloud. Despite its tremendous popularity, the security vulnerabilities of WAF-as-a-service are still largely unknown, which is highly concerning given its massive usage. In this paper, we propose a general and extendable attack framework, namelyAdvSQLi, in which a minimal series of transformations are performed on the hierarchical tree representation of the original SQLi payload, such that the generated SQLi payloads can not only bypass WAF-as-a-service under black-box settings but also keep the same functionality and maliciousness as the original payload. WithAdvSQLi, we make it feasible to inspect and understand the security vulnerabilities of WAFs automatically, helping vendors make products more secure. To evaluate the attack effectiveness and efficiency ofAdvSQLi, we first employ two public datasets to generate adversarial SQLi payloads, leading to a maximum attack success rate of 100% against state-of-the-art ML-based SQLi detectors. Furthermore, to demonstrate the immediate security threats caused byAdvSQLi, we evaluate the attack effectiveness against 7 WAF-as-a-service solutions from mainstream vendors and find all of them are vulnerable toAdvSQLi. For instance,AdvSQLiachieves an attack success rate of over 79% against the F5 WAF. Through in-depth analysis of the evaluation results, we further condense out several general yet severe flaws of these vendors that cannot be easily patched. Zhenqing Qu, Xiang Ling 0001, Ting Wang 0006, Xiang Chen 0017, Shouling Ji, Chunming Wu 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Toward Scalable and Low-Cost Traffic Testing for Evaluating DDoS Defense SolutionsabstractTo date, security researchers evaluate their solutions of mitigating distributed denial-of-service (DDoS) attacks via kernel-based or kernel-bypassing testing tools. However, kernel-based tools exhibit poor scalability in attack traffic generation while kernel-bypassing tools incur unacceptable monetary cost. We propose Excalibur, a scalable and low-cost testing framework for evaluating DDoS defense solutions. The key idea is to leverage the emerging programmable switch to empower testing tasks with Tbps-level scalability and low cost. Specifically, Excalibur offers intent-based primitives to enable academic researchers to customize testing tasks on demand. Moreover, in view of switch resource limitations, Excalibur coordinates both a server and a programmable switch to jointly perform testing tasks. It realizes flexible attack traffic generation, which requires a large number of resources, in the server while using the switch to increase the sending rate of attack traffic to Tbps-level. We have implemented Excalibur on a$64\times 100$Gbps Tofino switch. Our experiments on a$64\times 100$Gbps Tofino switch show that Excalibur achieves orders-of-magnitude higher scalability and lower cost than existing tools. Xiang Chen 0017, Hongyan Liu 0001, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006 |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | Hermes: Low-Overhead Inter-Switch Coordination in Network-Wide Data Plane Program DeploymentabstractNetwork administrators usually realize network functions in data plane programs. They employ the network-wide program deployment that decomposes input programs into match-action tables (MATs) while deploying each MAT on a specific switch. Since MATs may be deployed on different switches, existing solutions propose the inter-switch coordination that uses the per-packet header space to deliver crucial packet processing information among switches. However, such coordination incurs non-trivial per-packet byte overhead, leading to end-to-end performance degradation. We propose, a framework that aims to minimize the per-packet byte overhead. The key idea is to formulate network-wide program deployment as a mixed-integer programming (MIP) problem with the objective of minimizing the per-packet byte overhead. Also, offers a greedy-based heuristic that solves the problem in a near-optimal and timely manner. We have implemented on Tofino switches. Compared to existing frameworks, decreases the per-packet byte overhead by 156 bytes while preserving end-to-end performance in terms of flow completion time and goodput. Xiang Chen 0017, Hongyan Liu 0001, Qingjiang Xiao, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006, Qiang Yang 0004 |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | Toward Full-Coverage and Low-Overhead Profiling of Network-Stack LatencyabstractIn modern data center networks (DCNs), network-stack processing denotes a large portion of the end-to-end latency of TCP flows. So profiling network-stack latency anomalies has been considered as a crucial part in DCN performance diagnosis and troubleshooting. In particular, such profiling requires full coverage (i.e., profiling every TCP packet) and low overhead (i.e., profiling should avoid high CPU consumption in end-hosts). However, existing solutions rely on system calls or tracepoints in end-hosts to implement network-stack latency profiling, leading to either low coverage or high overhead. We propose Torp, a framework that offers full-coverage and low-overhead profiling of network-stack latency. Our key idea is to offload as much of the profiling from costly system calls or tracepoints to the Torp agent built on eBPF modules, and further to include a Torp handler on the ToR switch to accelerate the remaining profiling operations. Torp efficiently coordinates the ToR switch and the Torp agent on end-hosts to jointly execute the entire latency profiling task. We have implemented Torp on$32\times 100$Gbps Tofino switches. Testbed experiments indicate that Torp achieves full coverage and orders of magnitude lower host-side overhead compared to other solutions. Xiang Chen 0017, Hongyan Liu 0001, Wenbin Zhang 0011, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Xuan Liu 0006, Chunming Wu 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | Resource-Efficient and Timely Packet Header Vector (PHV) Encoding on Programmable SwitchesabstractThe programmable switch offers a limited capacity of packet header vector (PHV) words that store packet header fields and metadata fields defined by network functions. However, existing switch compilers employ inefficient strategies of encoding fields on PHV words. Their encoding wastes scarce PHV words and may result in failures when deploying network functions. In this paper, we propose Melody, a new framework that reuses PHV words for as many fields as possible to achieve resource-efficient PHV encoding. Melody offers a field analyzer and an optimization framework. The analyzer identifies which fields can reuse PHV words while preserving the original packet processing logic. The framework integrates analysis results into its encoding to offer the resource-optimal decisions. Also, to achieve timeliness at runtime, it provides a Greedy-based heuristic, which quickly solves PHV encoding and returns near-optimal results. We evaluate Melody with production-scale network functions. Our results show that Melody reduces the consumption of PHV words by up to 85%. Xiang Chen 0017, Wenbin Zhang 0011, Hongyan Liu 0001, Jianshan Zhang, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Xuan Liu 0006, Chunming Wu 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | Toward Resource-Efficient and High- Performance Program Deployment in Programmable NetworksabstractProgrammable switches allow administrators to customize packet processing behaviors in data plane programs. However, existing solutions for program deployment fail to achieve resource efficiency and high packet processing performance. In this paper, we propose SPEED, a system that provides resource-efficient and high-performance deployment for data plane programs. For resource efficiency, SPEED merges input data plane programs by reducing program redundancy. Then it abstracts the substrate network into an one big switch (OBS), and deploys the merged program on the OBS while minimizing resource usage. For high performance, SPEED searches for the performance-optimal mapping between the OBS and the substrate network with respect to network-wide constraints. It also maintains program logic among different switches via inter-device packet scheduling. We have implemented SPEED on a Barefoot Tofino switch. The evaluation indicates that SPEED achieves resource-efficient and high-performance deployment for real data plane programs. Hongyan Liu 0001, Xiang Chen 0017, Qun Huang 0001, Peiqiao Wang, Dong Zhang 0010, Chunming Wu 0001, Xuan Liu 0006, Qiang Yang 0004 |
IEEE/ACM Trans. Netw. | 2 |
| 2023 | Self-Supervised Interest Transfer Network via Prototypical Contrastive Learning for RecommendationabstractCross-domain recommendation has attracted increasing attention from industry and academia recently. However, most existing methods do not exploit the interest invariance between domains, which would yield sub-optimal solutions. In this paper, we propose a cross-domain recommendation method: Self-supervised Interest Transfer Network (SITN), which can effectively transfer invariant knowledge between domains via prototypical contrastive learning. Specifically, we perform two levels of cross-domain contrastive learning: 1) instance-to-instance contrastive learning, 2) instance-to-cluster contrastive learning. Not only that, we also take into account users' multi-granularity and multi-view interests. With this paradigm, SITN can explicitly learn the invariant knowledge of interest clusters between domains and accurately capture users' intents and preferences. We conducted extensive experiments on a public dataset and a large-scale industrial dataset collected from one of the world's leading e-commerce corporations. The experimental results indicate that SITN achieves significant improvements over state-of-the-art recommendation methods. Additionally, SITN has been deployed on a micro-video recommendation platform, and the online A/B testing results further demonstrate its practical value. Supplement is available at: https://github.com/fanqieCoffee/SITN-Supplement. Yibin Shen, Sijin Zhou, Xiang Chen 0017, Hongyan Liu 0001, Chunming Wu 0001, Chenyi Lei, Xianhui Wei, Fei Fang 0002 |
AAAI | 4 |
| 2023 | Halia: Toward Full-Coverage Network Function Offloading in the Data PlaneabstractOffloading network functions (NFs) to data plane switches brings remarkable performance benefits. In such offloading, NFs are required to process all the flows of interest (i.e., full coverage) to preserve the quality of services. However, existing solutions fail to guarantee full coverage for NFs. Thus, NFs may miss some essential flows, leading to accuracy drops. In this paper, we propose Halia, a framework that makes NF offloading decisions while ensuring full coverage for NFs. Specifically, Halia formulates the problem of NF offloading as an optimization problem. It encodes the requirement of full coverage as a constraint. Thus, its decisions activate enough NF instances in the substrate network to achieve full coverage.. We have implemented Halia and conducted experiments under multiple realistic network topologies to evaluate Halia. The experimental results indicate that compared to existing solutions, Halia achieves full coverage and high scalability in large-scale networks. Hongyan Liu 0001, Xiang Chen 0017, Qingjiang Xiao, Kaiwei Guo, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
ICC | 3 |
| 2023 | Excalibur: A Scalable and Low-Cost Traffic Testing Framework for Evaluating DDoS Defense SolutionsabstractTo date, security researchers evaluate their solutions of mitigating denial-of-service (DDoS) attacks via kernel-based or kernel-bypassing testing tools. However, kernel-based tools exhibit poor scalability in attack traffic generation while kernel-bypassing tools result in unacceptable monetary cost. We propose Excalibur, a scalable and low-cost testing framework for DDoS defense solutions. The key idea is to leverage the programmable switch to perform testing tasks with Tbps-level scalability and low cost. Specifically, Excalibur coordinates both a server and a programmable switch to jointly perform testing tasks. It realizes flexible attack traffic generation, which requires a large number of resources, in the server while using the switch to increase the sending rate of attack traffic to Tbps-level. Our experiments on a 64×100Gbps Tofino switch show that Excalibur achieves orders-of-magnitude higher scalability and lower cost than existing tools. Xiang Chen 0017, Hongyan Liu 0001, Tingxin Sun, Qun Huang 0001, Dong Zhang 0010, Xuan Liu 0006, Haifeng Zhou, Chunming Wu 0001 |
INFOCOM | 1 |
| 2023 | Melody: Toward Resource-Efficient Packet Header Vector Encoding on Programmable SwitchesabstractThe programmable switch offers a limited capacity of packet header vector (PHV) words that store packet header fields and metadata fields defined by network functions. However, existing switch compilers employ inefficient strategies of encoding fields on PHV words. Their encoding wastes scarce PHV words and may result in failures when deploying network functions. In this paper, we propose Melody, a new framework that reuses PHV words for as many fields as possible to achieve resource-efficient PHV encoding. Melody offers a field analyzer and an optimization framework. The analyzer identifies which fields can reuse PHV words while preserving the original packet processing logic. The framework integrates analysis results into its encoding to offer the resource-optimal decisions. We evaluate Melody with production-scale network functions. Our results show that Melody reduces the consumption of PHV words by up to 85%. Xiang Chen 0017, Hongyan Liu 0001, Qingjiang Xiao, Jianshan Zhang, Qun Huang 0001, Dong Zhang 0010, Xuan Liu 0006, Chunming Wu 0001 |
INFOCOM | 1 |
| 2023 | In-band Network Telemetry Manipulation Attacks and Countermeasures in Programmable NetworksabstractIn-band Network Telemetry (INT) is a widely used monitoring framework in modern large-scale networks that provides fine-grained visibility into network conditions by inserting telemetry data into packets. However, this mechanism also introduces new vulnerabilities that malicious attackers can exploit. In this paper, we present four In-band Network Telemetry Manipulation Attacks that take advantage of INT's weakness, demonstrating that attackers can cause severe damage with little effort by manipulating INT packets. To address this issue, we design SecureINT, a novel INT prototype that ensures confidentiality and integrity for INT packets. To meet the stringent computational requirements of programmable switches, we comprehensively analyze possible attacks on the deployed encryption/hash algorithms and modify them accordingly without compromising their security. According to the experiments, SecureINT can be deployed on programmable switches using a single pipeline, providing encryption and integrity verification for INT packets with minimal overhead. Dezhang Kong, Zhengyan Zhou, Yi Shen 0012, Xiang Chen 0017, Qiumei Cheng, Dong Zhang 0010, Chunming Wu 0001 |
IWQoS | 4 |
| 2023 | Vulnerabilities and Attacks of Inter-device Coordination in Programmable NetworksabstractIn programmable networks, some networking systems coordinate data plane switches to realize in-network functions (e.g., in-band network telemetry). However, the vulnerabilities of inter-device coordination are still largely unknown and neglected, which is highly concerning given the increasing popularity of this paradigm. In this paper, we identify three attack scenarios built upon such vulnerabilities, where attackers mislead the behaviors of networking systems that exploit inter-device coordination to execute in-network functions. We implement 20 existing networking systems on Tofino-based switches and a simulator, and attack these systems with the identified attacks. The experimental results indicate that our attacks significantly interfere with the normal operations of the selected networking systems, e.g., the cache hit rate of NetCache drops 38%. Our analysis also demonstrates that none of existing methods can fully mitigate our attacks since they fail to verify the packets for inter-device coordination. Hongyan Liu 0001, Xiang Chen 0017, Yi Shen 0012, Qun Huang 0001, Zhengyan Zhou, Dong Zhang 0010, Chunming Wu 0001 |
IWQoS | 2 |
| 2023 | Optimizing Program Deployment with libopl in Programmable NetworksabstractDeploying data plane programs on programmable switches involves complex optimization problems that make the optimal deployment decisions. However, existing deployment frameworks only focus on deploying programs in specific domains (i.e., supporting fixed optimization requirements), resulting in poor scalability. To this end, our goal is to simplify program deployment through general high-level abstractions that capture optimization requirements. In this paper, we present libopl, a generic library that enables administrators to express various optimization requirements when deploying programs and further calculates the optimal deployment plans. Existing frameworks can also use libopl to extend their functionalities to fit more deployment scenarios. To evaluate libopl, we build a Tofino-based testbed and a simulator. Our experimental results show that libopl exhibits comparable or better scalability than stateof-the-art frameworks and only introduces negligible overhead. Hongyan Liu 0001, Xiang Chen 0017, Yi Shen 0012, Dong Zhang 0010, Chunming Wu 0001 |
SECON | 2 |
| 2023 | RFT: Toward Highly Reliable Flow Data Transmission in Network MeasurementabstractHow to satisfy the latency and reliability requirements of flow data transfer is an essential problem. To address this problem, we propose RFT, a framework that aims to satisfy the user-specified latency and reliability requirements of flow data transfer, especially in the situation where the network resources are insufficient. Firstly, we formulate the problem of satisfying the user-specified latency and reliability requirements of data transfer via mixed integer linear programming (MILP), and a heuristic algorithm is then designed to solve it in a polynomialtime. Secondly, to satisfy these requirements under insufficient network resources, we proposed a greedy-based algorithm used to select the minimum number of links added to the network, which can be deployed with low cost, especially in production networks such as data centers. Finally, we have implemented RFT on a 64$\times$100 Gbps Intel Barefoot Tofino switch. Our experimental results indicate that RFT satisfies the user-specified latency and reliability requirements in all test cases at acceptable costs, even when the network resources are insufficient. Xiang Chen 0017, Di Wang 0003, Zhengyan Zhou, Wenhai Wang, Chunming Wu 0001, Haifeng Zhou |
SECON | 2 |
| 2023 | BitSense: Universal and Nearly Zero-Error Optimization for Sketch Counters with Compressive SensingabstractSketch algorithms have been widely deployed for network measurement as they achieve high accuracy with restricted resource usage. They store measurement results compactly in fixed-size counters. However, as sketch counters are skewed towards low values, higher bits in most counters remain zero. Such massive unused bits impair the space efficiency valued by sketch algorithms. Unfortunately, efforts to mitigate the issue either apply to specific algorithms or compromise accuracy. In this paper, we design BitSense, a novel optimization framework that integrates with existing sketch algorithms. The key idea is to regard higher bits in sketch counters as a sparse vector and leverage compressive sensing techniques to compress and restore counters. Further, BitSense provides a programming model to help developers easily realize sketch algorithms without dealing with the details of compression and recovery. Bit-Sense proposes an automatic approach for parameter configuration. It theoretically guarantees nearly zero error under the configuration. We have built a BitSense prototype in P4 and a software platform and integrated it with fourteen sketch solutions. Extensive experiments show that BitSense significantly reduces the memory usage of existing sketch solutions by 25%-80% while incurring little overhead and almost zero accuracy drop, outperforming five state-of-the-art optimization frameworks. Rui Ding 0014, Shibo Yang, Xiang Chen 0017, Qun Huang 0001 |
SIGCOMM | 3 |
| 2023 | Adversarial attacks against Windows PE malware detection: A survey of the state-of-the-art
Xiang Ling 0001, Lingfei Wu 0001, Jiangyu Zhang, Zhenqing Qu, Xiang Chen 0017, Yaguan Qian, Chunming Wu 0001, Shouling Ji, Tianyue Luo, JingZheng Wu |
Comput. Secur. | 6 |
| 2023 | Automatic Performance-Optimal Offloading of Network Functions on Programmable SwitchesabstractIn network function virtualization (NFV), network functions (NFs) are chained as a service function chain (SFC) to enhance NF management with low cost and high flexibility. Recent NFV solutions indicate that the packet processing performance of SFCs can be significantly improved by offloading NFs to programmable switches. However, such offloading requires a deep understanding of heterogeneous NF properties (e.g., NF resource consumption and NF performance behaviors) to achieve the maximum SFC performance. Unfortunately, none of existing solutions provide automatic analysis of these NF properties. Thus, network administrators have to manually examine the source codes of NFs and profile various NF properties by hand, which is extremely time-consuming and laborious. In this article, we propose LightNF, a novel system that simplifies NF offloading in programmable networks. LightNF automatically dissects comprehensive NF properties by means of code analysis and performance profiling while eliminating manual efforts. It then leverages its analysis results of NF properties in its SFC placement so as to make the performance-optimal offloading decisions. We have implemented LightNF on Tofino-based hardware programmable switches. We perform extensive experiments to evaluate LightNF with a real-world testbed and large-scale simulation. Our experiments show that LightNF outperforms existing solutions with an orders-of-magnitude reduction in per-packet processing latency and 9.5× improvement in SFC throughput. Xiang Chen 0017, Hongyan Liu 0001, Dong Zhang 0010, Zili Meng, Qun Huang 0001, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006, Qiang Yang 0004 |
IEEE Trans. Cloud Comput. | 1 |
| 2023 | Stalker Attacks: Imperceptibly Dropping Sketch Measurement Accuracy on Programmable SwitchesabstractDue to limited memory usage and provably high accuracy, sketches running on programmable switches have been commonly used by the literature for network measurement. However, their vulnerabilities are still largely unknown and neglected, which is highly concerning given the increasing popularity of network measurement. In this paper, we identify the Stalker attacks, where attackers aim to degrade the accuracy of sketches running on programmable switches. More precisely, attackers tamper with some sketch operations during sketch deployment atop programmable switches. At runtime, the tampered sketch will record highly inaccurate flow data, which degrades measurement accuracy. We implement Stalker attacks on Tofino switches. The results indicate that Stalker attacks significantly drop the accuracy of network management applications, e.g., reducing the F1 score of heavy hitter detection to zero. However, our analysis indicates that none of existing methods can detect Stalker attacks since they can hardly verify the correctness of sketch operations. Finally, we analyze potential defense mechanisms and identify challenges to enable further research in this context. Xiang Chen 0017, Hongyan Liu 0001, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006, Muhammad Khurram Khan |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2023 | Toward Low-Latency and Accurate State Synchronization for Programmable NetworksabstractProgrammable switches empower stateful packet processing, in which incoming packets continuously update states in the data plane, while applications in the control plane read and write states. However, since the data plane and control plane are separated, a consistent view of states in both planes is required for stateful packet processing. Existing approaches suffer from either high latency or low accuracy. In this paper, we propose ApproSync, a framework that offers approximate state synchronization with low latency and high accuracy. To achieve low latency, ApproSync directly transfers states between switch ASICs and the control plane by bypassing switch operating systems. To achieve high accuracy, ApproSync utilizes the resources in the switch ASIC to realize rate control in state synchronization, such that it avoids potential state loss. It also bounds the divergence between the states in the data plane and that in the control plane under limited link capacity. We prototype ApproSync on Barefoot Tofino switches. The experimental results indicate that compared to existing approaches, ApproSync achieves order-of-magnitude latency reduction while maintaining high accuracy of state synchronization. Also, our experiments demonstrate that ApproSync provides significant latency benefits to existing network management applications and well preserves high application-level accuracy. Xiang Chen 0017, Hongyan Liu 0001, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006, Qiang Yang 0004 |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | Eliminating Control Plane Overload via Measurement Task PlacementabstractRecent efforts in network measurement place measurement tasks on programmable switches to measure high-speed traffic. These tasks extract flow data, i.e., events, from packets and send events to the control plane. However, the tasks may generate massive events in a short time. In this context, the links transferring events to the control plane and the control plane servers that handle events may be overloaded, i.e., control plane overload. None of existing solutions can eliminate control plane overload. In this paper, we propose MTP, a framework that eliminates control plane overload via careful measurement task placement. Our key idea is to allocate enough resources for each task during task placement to avoid control plane overload at runtime. For each task, MTP estimates its maximum possible rate of sending events to the control plane. Then its optimization framework addresses the resource restrictions of both switches and the control plane. The experiments on Tofino switches indicate that MTP outperforms existing solutions with higher accuracy in several use cases. Xiang Chen 0017, Hongyan Liu 0001, Dong Zhang 0010, Qun Huang 0001, Haifeng Zhou, Chunming Wu 0001, Qiang Yang 0004 |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | Combination Attacks and Defenses on SDN Topology DiscoveryabstractThe topology discovery service in Software-Defined Networking (SDN) provides the controller with a global view of the substrate network topology, allowing for central management of the entire network. Unfortunately, emerging topology attacks can poison the network topology and result in unforeseeable disasters. Although researchers have made great efforts to mitigate this problem, security hazards still exist. In this paper, we propose Invisible Assailant Attack (IAA), the first combination topology attack capable of injecting and maintaining fake links even when 12 existing defense strategies are deployed simultaneously. IAA consists of 14 attack phases that apply multiple attack strategies. Attackers skillfully disguise the attack traffic in each phase so that it looks like normal network traffic, and perform these phases in a well-planned sequence, thereby bypassing existing defenses step by step. To mitigate this attack, we propose a Route Path Verification (RPV) mechanism that orchestrates multiple defense strategies to identify fake links. According to the experiments, RPV can successfully detect IAA with low overhead: its detection completes within 1 ms while its per-flow storage consumption is only a few KB. Dezhang Kong, Yi Shen 0012, Xiang Chen 0017, Qiumei Cheng, Hongyan Liu 0001, Dong Zhang 0010, Xuan Liu 0006, Shuangxi Chen, Chunming Wu 0001 |
IEEE/ACM Trans. Netw. | 3 |
| 2022 | TableGuard: A Novel Security Mechanism Against Flow Table Overflow Attacks in SDNabstractOne of the most important components of Software-Defined Networking (SDN) is the flow table. It receives flow rules from the controller and uses them to handle network traffic. However, a flow table can only store a few thousand flow rules, which makes it an attractive target for table overflow attacks. These attacks force the controller to populate the flow table with a large number of meaningless flow rules, which prevents normal flows from finding matching rules and therefore having to be reported to the controller. It results in a significant latency overhead, degrading the performance of the whole network. In this paper, we present a key characteristic of table overflow attacks: even though attackers can change some critical attack parameters (e.g., attack speed) to avoid detection, proactive flows from the attacked port always occupy a stable proportion in the flow table regardless of the attack form. In light of this finding, we propose TableGuard, a novel security mechanism that uses the proactive flow rule number as the detection metric and applies a statistical approach to help filter malicious flows. The experiments demonstrate that TableGuard can mitigate both high-rate and low-rate table overflow attacks. Compared with existing defenses, TableGuard has the best mitigation performance and the minimal overhead on normal flows. Dezhang Kong, Chunming Wu 0001, Yi Shen 0012, Xiang Chen 0017, Hongyan Liu 0001, Dong Zhang 0010 |
GLOBECOM | 4 |
| 2022 | Toward Low-Overhead Inter-Switch Coordination in Network-Wide Data Plane Program DeploymentabstractIn modern networks, administrators realize their desired functions such as network measurement in several data plane programs. They often employ the network-wide program deployment paradigm that decomposes input programs into match-action tables (MATs) while deploying each MAT on a specific programmable switch. Since MATs may be deployed on different switches, existing solutions propose the inter-switch coordination that uses the per-packet header space to deliver crucial packet processing information among switches. However, such coordination introduces non-trivial per-packet byte overhead, leading to significant end-to-end network performance degradation. In this paper, we propose Hermes, a program deployment framework that aims to minimize the per-packet byte overhead. The key idea of Hermes is to formulate the network-wide program deployment as a mixed-integer linear programming (MILP) problem with the objective of minimizing the per-packet byte overhead. In view of the NP hardness of the MILP problem, Hermes further offers a greedy-based heuristic that solves the problem in a near-optimal and timely manner. We have implemented Hermes on Tofino-based switches. Our experiments show that compared to existing frameworks, Hermes decreases the per-packet byte overhead by 156 bytes while preserving end-to-end performance in terms of flow completion time and goodput. Xiang Chen 0017, Hongyan Liu 0001, Qingjiang Xiao, Kaiwei Guo, Tingxin Sun, Xiang Ling 0001, Xuan Liu 0006, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
ICDCS | 1 |
| 2022 | SketchGuide: Reconfiguring Sketch-based Measurement on Programmable SwitchesabstractSketches enable efficient and fine-grained network measurement results with configurable resource-performance trade-offs. While sketch configurations are guided by theories, the current theoretical guidelines are either impractical or deficient for sketch configurations on emerging programmable switches. To better configure sketches on programmable switches, we (1) systematically analyze the limitations of sketch configuration guidelines on programmable hardware switches (i.e., unguided parameters, accuracy profiles, and resource budgets); (2) propose a generic and practical framework called SketchGuide to automate efficient sketch configurations on programmable switches; (3) implement SketchGuide on a Barefoot Tofino switch and compare SketchGuide to the state-of-the-art sketches by conducting extensive experiments. Our evaluations demonstrate that SketchGuide can automatically configure unguided parameters given resource budgets. SketchGuide reduces the hardware resource footprint by 52.92%-99.28% compared with current guidelines without impacting fidelity. Zhengyan Zhou, Jingwen Lv, Lingfei Cheng, Xiang Chen 0017, Tianzhu Zhang 0002, Qun Huang 0001, Jiayu Luo, Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001 |
ICNP | 4 |
| 2022 | Torp: Full-Coverage and Low-Overhead Profiling of Host-Side LatencyabstractIn data center networks (DCNs), host-side packet processing accounts for a large portion of the end-to-end latency of TCP flows. Thus, the profiling of host-side latency anomalies has been considered as a crucial part in DCN performance diagnosis and troubleshooting. In particular, such profiling requires full coverage (i.e., profiling every TCP packet handled by end-hosts) and low overhead (i.e., profiling should avoid high CPU consumption in end-hosts). However, existing solutions fully rely on end-hosts to implement host-side latency profiling, leading to low coverage or high overhead. In this paper, we propose Torp, a framework that offers full-coverage and low-overhead profiling of host-side latency. Our key idea is to offload profiling operations to top-of-rack (ToR) switches, which inherently offer full coverage and line-rate packet processing performance. Specifically, Torp selectively offloads profiling operations to the ToR switch based on switch limitations. It efficiently coordinates the ToR switch and end-hosts to execute the entire latency profiling task. We have implemented Torp on 32×100Gbps Tofino switches. Testbed experiments indicate that Torp achieves full coverage and orders of magnitude lower host-side overhead compared to other solutions. Xiang Chen 0017, Hongyan Liu 0001, Junyi Guo, Qun Huang 0001, Dong Zhang 0010, Chunming Wu 0001, Haifeng Zhou |
INFOCOM | 1 |
| 2022 | Escala: Timely Elastic Scaling of Control Channels in Network MeasurementabstractIn network measurement, data plane switches measure traffic and report events (e.g., heavy hitters) to the control plane via control channels. The control plane makes decisions to process events. However, current network measurement suffers from two problems. First, when traffic bursts occur, massive events are reported in a short time so that the control channels may be overloaded due to limited bandwidth capacity. Second, only a few events are reported in normal cases, making control channels underloaded and wasting network resources. In this paper, we propose Escala to provide the elastic scaling of control channels at runtime. The key idea is to dynamically migrate event streams among control channels to regulate the loads of these channels. Escala offers two components, including an Escala monitor that detects scaling situations based on realtime network statistics, and an optimization framework that makes scaling decisions to eliminate overload and underload situations. We have implemented a prototype of Escala on Tofino-based switches. Extensive experiments show that Escala achieves timely elastic scaling while preserving high application-level accuracy. Hongyan Liu 0001, Xiang Chen 0017, Qun Huang 0001, Dezhang Kong, Jinbo Sun, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
INFOCOM | 2 |
| 2021 | MTP: Avoiding Control Plane Overload with Measurement Task PlacementabstractIn programmable networks, measurement tasks are placed on programmable switches to keep pace with high-speed traffic. At runtime, programmable switches send events to the control plane for further processing. However, existing solutions for task placement overlook the limitations of control plane resources. Thus, excessive events may overload the control plane. In this paper, we propose MTP, a system that eliminates control plane overload via careful task placement. For each task, MTP analyzes its structure to estimate its maximum possible rate of sending events to the control plane. Then it builds an optimization framework that addresses the resource restrictions of both switches and the control plane. We have implemented MTP on Barefoot Tofino switches. The experimental results indicate that MTP outperforms existing solutions with higher accuracy across four real use cases. Xiang Chen 0017, Qun Huang 0001, Peiqiao Wang, Hongyan Liu 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
INFOCOM | 1 |
| 2021 | LightNF: Simplifying Network Function Offloading in Programmable NetworksabstractIn network function virtualization (NFV), network functions (NFs) are chained as a service function chain (SFC) to enhance NF management with high flexibility. Recent solutions indicate that the processing performance of SFCs can be significantly improved by offloading NFs to programmable switches. However, such offloading requires a deep understanding of NF properties to achieve the maximum SFC performance, which brings non-trivial burdens to network administrators. In this paper, we propose LightNF, a novel system that simplifies NF offloading in programmable networks. LightNF automatically dissects comprehensive NF properties (e.g., NF performance behaviors) via code analysis and performance profiling while eliminating manual efforts. It then leverages the analyzed NF properties in its SFC placement so as to produce the performance-optimal offloading. We have implemented a LightNF prototype. Our experiments show that LightNF outperforms state-of-the-art solutions with an orders-of-magnitude reduction in per-packet processing latency and 9.5× improvement in SFC throughput. Xiang Chen 0017, Qun Huang 0001, Peiqiao Wang, Zili Meng, Hongyan Liu 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
IWQoS | 1 |
| 2021 | Toward Nearly-Zero-Error Sketching via Compressive Sensing
Qun Huang 0001, Siyuan Sheng, Xiang Chen 0017, Yungang Bao, Yanwei Xu 0004, Gong Zhang 0001 |
NSDI | 3 |
| 2021 | Conditional Variational Auto-Encoder and Extreme Value Theory Aided Two-Stage Learning Approach for Intelligent Fine-Grained Known/Unknown Intrusion DetectionabstractPromptly discovering unknown network attacks is critical for reducing the risk of major loss imposed on organizations and information infrastructure. This paper aims at developing an intelligent intrusion detection system capable of classifying known attacks as well as inferring unknown ones. To achieve this, we formulate the problem of fine-grained known/unknown intrusion detection as a two-stage minimization problem, where the first stage is to seek a score measure for minimizing the empirical risk of misclassifying the known attacks, while the second stage is to find another score measure for minimizing the identification risk of inferring unknown attacks. The hierarchical nature of problem formulation allows us to employ the class conditioned auto-encoders to construct a hierarchical intrusion detection framework. Since the reconstruction errors of unknown attacks are generally higher than that of the known attacks, we further employ extreme value theory in the second stage to model the distribution of reconstruction errors for differentiating known/unknown attack. To further reduce the false positive rate, we add a benign clustering module for learning the multimodal distribution of benign traffic. We conduct an experiment on two widely used datasets for assessing intrusion detection. The results show that the proposed method improves the detection rate of unknown attacks while keeping a low false positive rate. Jian Yang 0014, Xiang Chen 0017, Shuangwu Chen, Xiaofeng Jiang, Xiaobin Tan |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | SRA: Switch Resource Aggregation for Application Offloading in Programmable NetworksabstractProgrammable switches empower network applications with line-rate packet processing performance by allowing the offloading of applications. However, the resource of a programmable switch is extremely limited, which significantly limits the application offloading. Existing solutions to the problem either provide poor efficiency or suffer from accuracy drop. In this paper, we propose SRA, a system that loosens switch resource constraints for application offloading via resource aggregation. SRA provides administrators with an intuitive compiler directive to customize application offloading by resource aggregation. According to compiler directives, it automatically places the program on the substrate network, while maintaining original packet processing logics. We implement a prototype of SRA in P4, and establish an experimental testbed consisting of three 32$\times$100 Gbps Barefoot switches. The experimental results indicate that SRA enhances two real-world applications with sufficient resources while maintaining high performance. Hongyan Liu 0001, Xiang Chen 0017, Qun Huang 0001, Haifeng Zhou, Dong Zhang 0010, Chunming Wu 0001 |
GLOBECOM | 2 |
| 2020 | ApproSync: Approximate State Synchronization for Programmable NetworksabstractProgrammable switches empower stateful packet processing, in which incoming packets continuously update states in the data plane, while applications in the control plane read and write states. However, as the data plane and control plane are separated, a consistent view of states in both planes is required for stateful packet processing. Existing approaches suffer from either high latency or low accuracy. In this paper, we propose ApproSync, a framework that offers approximate state synchronization with low latency and high accuracy. To achieve low latency, ApproSync directly transfers states between switch ASICs and the control plane by bypassing switch operating systems. To achieve high accuracy, ApproSync utilizes the resources in the switch ASIC to realize rate control in state synchronization, such that it avoids potential state loss. It also bounds the divergence between the states in the data plane and that in the control plane under limited link capacity. We prototype ApproSync on Barefoot Tofino switches. The experimental results indicate that compared to existing approaches, ApproSync achieves order-of-magnitude latency reduction while maintaining high accuracy. Xiang Chen 0017, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
ICNP | 1 |
| 2020 | SPEED: Resource-Efficient and High-Performance Deployment for Data Plane ProgramsabstractProgrammable switches allow network administrators to customize packet processing behaviors in data plane programs. However, existing solutions for program deployment fail to achieve resource efficiency and high packet processing performance. In this paper, we propose SPEED, a system that provides resource-efficient and high-performance deployment for data plane programs. For resource efficiency, SPEED merges input data plane programs by reducing program redundancy. Then it abstracts the substrate network into an one big switch (OBS), and deploys the merged program on the OBS while minimizing resource usage. For high performance, SPEED searches for the performance-optimal mapping between the OBS and the substrate network with respect to network-wide constraints. It also maintains program logics among different switches via inter-device packet scheduling. We have implemented SPEED on a Barefoot Tofino switch. The evaluation indicates that SPEED achieves resource-efficient and high-performance deployment for real data plane programs. Xiang Chen 0017, Hongyan Liu 0001, Qun Huang 0001, Peiqiao Wang, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
ICNP | 1 |
| 2019 | P4SC: Towards High-Performance Service Function Chain Implementation on the P4-Capable Device
Xiang Chen 0017, Dong Zhang 0010, Haifeng Zhou |
IM | 1 |
| 2018 | MATReduce: Towards High-Performance P4 Pipeline by Reducing Duplicate Match OperationsabstractP4 provides operators with the ability to program the packet processing pipeline of the data plane device. The match-action table (MAT) is a basic component of the P4 pipeline that matches the packet and performs an action on the matched packet. However, different MATs may execute duplicate match operations that decreases the performance of the P4 pipeline. To this end, we present MATReduce, a framework that optimizes the P4 pipeline by reducing duplicate match operations between MATs. MATReduce is composed of two key components, the preprocessor and the runtime manager. By introducing the compound MAT and rewriting the P4 control flow, the preprocessor merges duplicate match operations of the P4 pipeline while maintaining the program semantics. At runtime, the runtime manager converts user rules to actual rules for maintaining the policy consistency. Our preliminary experimental results show that MATReduce provides significant performance improvement, including a 23.90% throughput increase and a 34.37% delay decrease on the software target, and a 45.19% delay decrease on the hardware target. Xiang Chen 0017, Dong Zhang 0010, Haifeng Zhou |
GLOBECOM | 1 |