Chunming Wu 0001

dblp:97/8329-1 · also Chun-Ming Wu 0001 · DBLP profile ↗
← Back
150ranked-venue papers
0as first author
116since 2021 · last 2026
0000-0001-7958-9687ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 95 · 76 since 2021Security and privacy · 18 · 15 since 2021Artificial intelligence and machine learning · 10 · 8 since 2021Systems, architecture and hardware · 10 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Turbolearn: Harnessing Accurate and Line-Rate Deep Learning on Programmable Switches
Zhifan Jiang, Longlong Zhu, Jiashuo Yu, Linying Zheng, Chunming Wu 0001, Xiang Chen 0017
APNet5
2026 TurboLearn: Harnessing Accurate and Line-Rate Deep Learning on Programmable Switches
abstract
The intelligent data plane (IDP) embeds deep learning (DL) models on switches for line-rate traffic analysis, but hardware constraints often force simplified models, reducing accuracy, while complex models like Transformers remain undeployable. We present TurboLearn, which achieves high accuracy and line-rate performance by co-designing inference across the switch ASIC and switch OS on the same switch. TurboLearn uses three key techniques: (1) a hardware fast path for lightweight classification and a software normal path for complex models, (2) confidence-based selective inference that escalates only low-confidence packets, and (3) confidence-calibrated knowledge distillation, where the normal path teaches the fast path. On Intel Tofino2 switches across three real-world traffic tasks, TurboLearn supports models that existing IDPs cannot deploy, improves macro-F1 by up to 31.31%, and keeps over 90% of traffic on the fast path with zero throughput loss.
Zhifan Jiang, Longlong Zhu, Jiashuo Yu, Linying Zheng, Chunming Wu 0001, Xiang Chen 0017
APNet5
2026 APTMatch: Empowering Dynamic Rule Updating for Learning-based Packet Classification
Jiashuo Yu, Longlong Zhu, Zongye Lin, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001
ICC9
2026 HYDRA: A Hybrid Synthesizer for Asymmetric Mixture-of-Experts Communication Scheduling
Lida Liao, Qianxun Xu, Xuanwei Si, Hongyan Liu 0001, Jiashuo Yu, Zongye Lin, Qiaoling Hu, Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001
ICC13
2026 DCAD: Dual-Condition Adaptive Consistency Model for IoT Privacy-Preserving Video Anomaly Detection
Shaopeng Zhou, Chaohao Li, Haonan Yan, Longlong Zhu, Chunming Wu 0001, Bin Wang 0062
ICIC (2)6
2026 MonPlan: Taming Network Measurement with Accurate and Resource-Efficient Sketch-INT Co-Design
Xiang Chen 0017, Linying Zheng, Longlong Zhu, Zedi Chen, Qing Shu, Jialu Tian, Siqi Dong, Qun Huang 0001, Jianshan Zhang, Xuan Liu 0006, Haifeng Zhou, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001
INFOCOM14
2026 LTD: Low-Overhead Topology Discovery using Programmable Data Planes
Dezhang Kong, Minghao Li 0012, Shi Lin, Zhenhua Xu 0004, Longlong Zhu, Linying Zheng, Xiang Chen 0017, Changting Lin, Xuan Liu 0006, Dong Zhang 0010, Chunming Wu 0001
INFOCOM11
2026 Holm: A DPU-Based Robust Host-Side Latency Monitoring with Low-Overhead and Selective Full-Coverage
Haifeng Zhou, Di Wang 0003, Wenbin Zhang 0011, Dianxing Tang, Zhengyan Zhou, Chunming Wu 0001
INFOCOM8
2026 Achieving Precise Host Congestion Mitigation by DPU Offloading
Haifeng Zhou, Di Wang 0003, Dianxing Tang, Zhengyan Zhou, Chunming Wu 0001
INFOCOM7
2026 GBNN: In-Network Gradient Boosting Neural Network
Shaowei Xu, Shengrui Lin, Hongyan Liu 0001, Libin Xu, Dong Zhang 0010, Chunming Wu 0001
IWQoS9
2026 PSM: Timely and Resource-Efficient Sketch Migration in Network Measurement
Hongyan Liu 0001, Xiang Chen 0017, Zhengyan Zhou, Di Wang 0003, Chunming Wu 0001
IWQoS7
2026 Breaking Isolation: A New Perspective on Hypervisor Exploitation via Cross-Domain Attacks
Yiming Tao, Qinying Wang, Chunming Wu 0001, Mingde Hu, Yizhi Ren, Shouling Ji
NDSS4
2026 SketchPipe: Toward Accurate Sketch-based Network Measurement on Multi-Pipeline Switches with Splitless Sketch Placement
Xiang Chen 0017, Longlong Zhu, Linying Zheng, Hongyang Du 0001, Dong Zhang 0010, Jianshan Zhang, Xuan Liu 0006, Qun Huang 0001, Dusit Niyato, Haifeng Zhou, Chunming Wu 0001, Hongyan Liu 0001, Kui Ren 0001
NSDI11
2026 Proteus: Towards Accurate and Low-overhead In-Network Malicious Traffic Detection
abstract
Network intrusion detection systems (NIDS) are essential for web security by identifying and dropping malicious traffic. Existing in-network NIDS leverage the Tbps-level packet processing capability of programmable switches to achieve high-speed flow classification. They translate complex trained machine learning models to decision trees (DTs), where DTs are deployed on programmable switches via single-DT or multiple-DT deployment. However, they face a fundamental trade-off: single-DT deployment suffers from low classification accuracy due to over-pruning of trees, while multiple-DT deployment suffers from high overhead due to deploying multiple tree replicas. In this paper, we propose Proteus, an in-network malicious traffic detection system that achieves both high classification accuracy and low overhead. Its key idea is to split the original DT into critical and normal sub-trees, where these sub-trees have different impacts on overall accuracy. More precisely, Proteus first splits a DT into one critical and several normal sub-trees for adapting to the accuracy requirement and switch resource budgets. Second, it minimizes coordination overhead between sub-trees while ensuring full flow coverage via mixed-integer linear programming. Third, it dynamically reallocates or migrates sub-trees to adapt to changing resources by monitoring both classification accuracy and switch resource changes. Testbed experiments with 12.8 Tbps programmable switches show that Proteus improves classification accuracy, reduces switch resource consumption, and reduces classification latency.
Longlong Zhu, Linying Zheng, Qing Shu, Zedi Chen, Jiashuo Yu, Shaopeng Zhou, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001, Xiang Chen 0017
WWW10
2026 DFQ+: Dynamic queuing for approximate fairness in programmable shared memory switches
Minghui Chang, Yunqi Gao, Bing Hu 0002, Pei Xiao 0001, Chunming Wu 0001, Liyan Li
Comput. Networks5
2026 Attention-Biased Reinforcement Learning Framework for Adaptive and Scalable Flocking of UAV Swarms
abstract
With the widespread development of multiple unmanned aerial vehicle (multi-UAV) systems, flocking motion has become a common but essential application in UAV swarms. However, most existing approaches are rule-based and only valid for specific scenarios, which are limited in adaptability and scalability. In this paper, we present a novel framework termed attention-biased deep deterministic policy gradient (ABDDPG), which combines the multi-agent deep deterministic policy gradient (MADDPG) algorithm and attention-biased Transformer (ABTransformer). First, we encode the comprehensive environmental features observed by each UAV as input. Then, we introduce ABTransformer into the actor-critic network architecture. This allows the model to fully utilize each drone’s positional relationship with surrounding objects (including obstacles and other drones), endowing it with the ability to allocate attention reasonably. Furthermore, we utilize LoRA for pretraining and fine-tuning to further improve performance and training efficiency. This attention-biased reinforcement learning model enables each UAV in the swarm to learn and allocate local attention autonomously, guaranteeing internal communication stability and flocking motion task completion. The experimental results demonstrate that ABDDPG outperforms previous methods in terms of arrival rate, training speed, and inference efficiency, and is robust to different scenarios.
Lan Mu, Tong Duan, Chunming Wu 0001
IEEE Trans Autom. Sci. Eng.3
2026 ADMM-Based Adversarial False Data Injection Attacks Against Multi-Label Locational Detection
abstract
While multi-label learning has shown excellent performance in False Data Injection Attack (FDIA) locational detection, it has also exposed some potential security risks and vulnerabilities. However, unlike the image domain, the vulnerabilities of multi-label learning in the field of power grid have just received attention and urgently need to be explored and addressed. In this paper, to achieve a better understanding for the security risks of deep learning-based multi-label FDIA detectors, we propose two Alternating Direction Method of Multipliers (ADMM) based adversarial attacks, which are applicable to two different scenarios. The proposed two ADMM-based attacks aim to reduce additional attack costs while seeking suitable adversarial perturbations, making the attacks more realistic and feasible. The experimental results verify the effectiveness of the proposed ADMM-based attacks, making noteworthy strides in fostering a profound comprehension of the vulnerabilities in the unique field of deep multi-label learning for power systems.
Jiwei Tian, Chao Shen 0001, Chenhao Lin, Meng Zhang 0011, Xiaofang Xia, Chao Ren 0006, Peican Zhu, Chunming Wu 0001, Xiang Chen 0017
IEEE Trans. Dependable Secur. Comput.8
2026 Toward Security-Enhanced In-Band Network Telemetry in Programmable Networks
abstract
In-band Network Telemetry (INT) is a widely used monitoring framework in modern large-scale networks. It provides packet-level visibility into network conditions by inserting telemetry data into packets, enabling unprecedented fine-grained network management. However, this mechanism also introduces new vulnerabilities that malicious attackers can exploit. In this paper, we present eight In-band Network Telemetry Manipulation Attacks that take advantage of INT’s weakness, demonstrating that attackers can cause severe damage with little effort by manipulating INT packets. To address this issue, we designed SecureINT, a security-enhanced INT prototype that provides encryption and integrity verification for INT packets. Specifically, SecureINT deploys Even-Mansour and SipHash for confidentiality and integrity, respectively. It also uses a zero-delay rotation mechanism, which enables administrators to dynamically change the version of the deployed Even-Mansour/SipHash running on programmable switches without the need to re-install new programs. In this way, SecureINT can provide lasting security for INT packets using the limited resources of programmable switches. According to the experiments, SecureINT can be deployed on programmable switches using a single pipeline. Besides, the overhead of the rotation mechanism running on the control plane is still minimal.
Dezhang Kong, Xiang Chen 0017, Zhengyan Zhou, Yi Shen 0012, Hongyan Liu 0001, Qiumei Cheng, Xuan Liu 0006, Dong Zhang 0010, Chunming Wu 0001, Muhammad Khurram Khan
IEEE Trans. Netw. Serv. Manag.10
2026 Code Language Models for Security Patch Management: How Far are We?
abstract
The rapid expansion of open-source software has also brought significant security challenges to cloud infrastructure, particularly introducing and propagating vulnerabilities. In response, effective security patch management establishes a continuous, structured pipeline by systematically identifying, testing, and deploying security patches to fix vulnerabilities. However, manually managing a large number of security patches (i.e., any update is approved and installed by hand) is time-consuming, leading to a great motivation for automating this process. Although Code Language Models (CodeLMs) have shown potential in various code-centric tasks, there remains an open question as to how well CodeLMs perform within the context of security patch management. To bridge this gap, we performed the first comprehensive empirical study on fine-tuning or prompting nine state-of-the-art CodeLMs for three security-patch-related downstream tasks, including silent patch identification (distinguishing security patches from normal commits), record-patch linking (connecting authoritative vulnerability records, e.g., CVE, to the corresponding fixing commits), and vulnerability description generation (providing a piece of text summarizing the vulnerability fixed by the patch), covering classification, ranking, and generation problems. Our findings reveal that there is no “one-size-fits-all” model that can always perform the best. Furthermore, due to the lack of task-specific knowledge, naively prompting LLMs with the basic strategies is not consistently reliable and may even underperform smaller PTMs. Additionally, existing automated evaluation metrics cannot fully reflect the capability of LLMs in considered tasks. These findings underscore the considerable gap between current capabilities and the practical requirements for deploying CodeLMs in automating security patch management.
Xingwei Lin, Sicong Cao, Le Yu 0002, Xiaobing Sun 0001, Fu Xiao 0001, Lei Xue 0001, Chunming Wu 0001, Kui Ren 0001, David Lo 0001
IEEE Trans. Serv. Comput.7
2025 AIA: Autoregression-Based Injection Attacks Against Text2SQL Models
abstract
To facilitate understanding of users' diverse queries against the back-end databases in web applications, researchers have introduced Text-to-SQL (Text2SQL) models that can generate well-structured SQL queries from users' query texts in natural language. As the Text2SQL model decouples the user queries with the back-end databases, it inherently mitigates the SQL injection risk posed by inserting users' input into pre-written SQL queries. However, what security risks to web applications may be posed by Text2SQL models remains an open question. In this paper, we present a new attack framework, named Autoregression-based Injection Attacks (AIA), to evaluate the security risks of Text2SQL models. In particular, AIA makes target models generate attack payloads by constructing specific inputs and adjusting the input auto-regressively. Our evaluation demonstrates that AIA can cause Text2SQL models to generate target output by adversarial inputs with success rates of over 70% in most scenarios. The generated adversarial input has certain transferability in target Text2SQL models. Additionally, practice experiments show that AIA can make Text2SQL models extract user lists from databases and even delete data in databases directly.
Deyin Li, Xiang Ling 0001, Changjiang Li, Xiang Chen 0017, Chunming Wu 0001
AAAI5
2025 Phantom: Virtualizing Switch Register Resources for Accurate Sketch-based Network Measurement
abstract
Sketches have proven to be useful for measuring traffic. They store measurement results in the registers of data plane switches. However, they suffer from the short of switch register resources, limiting their measurement accuracy.
Xiang Chen 0017, Hongyan Liu 0001, Zhengyan Zhou, Wenbin Zhang 0011, Hongyang Du 0001, Dong Zhang 0010, Xuan Liu 0006, Haifeng Zhou, Dusit Niyato, Qun Huang 0001, Chunming Wu 0001, Kui Ren 0001
EuroSys12
2025 DHC: Distributed Homomorphic Compression for Gradient Aggregation in Allreduce
abstract
Distributed training is critical for efficiently developing deep neural networks (DNNs) on tasks like image classification and natural language processing. However, as model and dataset sizes continue to grow, high communication overhead during gradient exchanges has become a major bottleneck in distributed training. Although existing homomorphic compression frameworks effectively reduce communication overhead, their reliance on centralized architectures makes them unsuitable for the mainstream decentralized AllReduce architecture. To address this, we propose DHC, a framework for homomorphic gradient compression in AllReduce architectures. Its key idea is HG-Sketch, which leverages multi-level index tables for direct in-network aggregation of compressed gradients, thereby eliminating additional computational overhead. Additionally, DHC introduces an index-sharing method to optimize memory usage on programmable switches. Furthermore, we establish an Integer Linear Programming (ILP) model to optimize the deployment strategy of programmable switches, further enhancing in-network aggregation capabilities. Experimental results demonstrate that DHC achieves a$3.8 \times$increase in aggregation speed and a$4.2 \times$improvement in aggregation throughput.
Lida Liao, Zhengli Lin, Longlong Zhu, Hongyan Liu 0001, Jiashuo Yu, Dong Zhang 0010, Chunming Wu 0001
ICC8
2025 P4Alex: A Scalable Range Matching Approach for Programmable Switches
abstract
Range matching (RM), a flexible primitive for implementing network applications on programmable switches, is often subject to limited TCAM resources. Consequently, existing RM approaches rely on SRAM/ALU-assisted data structures to extend TCAM capacity. However, they consume significant SRAM, ALU, and pipeline stages, which hinders the implementation of other primitives (e.g., basic forwarding). In this paper, we propose P4Alex, a scalable RM framework for programmable switches. The key idea is leveraging emerging learned index structures to support RM, which replaces the storage by model inference for lightweight and fixed index depth. Unfortunately, the learning index structure cannot be directly implemented on programmable switches due to hardware limitations (e.g., floating-point computation and no-loop operations). In response, we design several optimizations in P4Alex to make it deployable. We successfully implemented P4Alex on Intel Tofino switches. Experimental results show that, compared to existing RM approaches, P4Alex extends RM capabilities from tens of thousands to millions. At the same scales, P4Alex reduces TCAM and SRAM resource consumption by up to 89.9% and 43.7%, respectively, while increasing latency by only$0.48 \mu ~\mathrm{s}$.
Jiashuo Yu, Longlong Zhu, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001
ICC5
2025 Carrera: Enabling High-Performance eBPF-based Sketches in Network Measurement
abstract
To achieve dynamic network measurement, trends build sketches on eBPF to avoid service interruptions. However, existing eBPF-based sketches suffer from high CPU consumption, leading to poor throughput and high latency and making them hard to measure high-speed traffic. Optimizing their performance requires users to refactor codes based on each sketch’s characteristics on eBPF, which is highly complex and time-consuming.In this paper, we argue that users should write sketches without concerning low-level eBPF performance optimizations, with the deployment automatically activating cross-sketch performance optimizations. We present Carrera, a library that offers domain-specific optimizations for eBPF-based sketches. Our contributions are (1) systematically analyzing the performance bottlenecks of eBPF-based sketches through microbenchmarks, (2) identifying practical optimizations, including hardware offloading, SIMD-accelerated hashing, traffic-aware flow index caching, prefetched randomization, and active data collection, to address the identified bottlenecks in eBPF-based sketches, (3) evaluating these optimizations with state-of-the-art sketches and demonstrating that Carrera improves throughput by up to 65% and reduces latency by up to 93% via testbed experiments.
Xiang Chen 0017, Xin Yao 0008, Longlong Zhu, Linying Zheng, Hongyan Liu 0001, Jianshan Zhang, Dong Zhang 0010, Xuan Liu 0006, Qun Huang 0001, Haifeng Zhou, Chunming Wu 0001
ICNP12
2025 EffiMatch: Enabling Fast and Accurate Learning-based Packet Classification
abstract
Learning-based Packet Classification methods reduce memory overhead by using lightweight Recursive Model Index(RMI) structures to limit the search range, followed by linear matching. However, they face a trade-off: complex RMI structures achieve smaller search ranges but slow down lookup, while simpler ones are faster but require larger scans. In this paper, we propose EffiMatch, a parallel multi-model lookup architecture aimed at resolving the trade-off between RMI complexity and linear search range in learning-based index systems. We propose two key designs: 1) We design a partitioning strategy called Distribution-Distance Partitioning (DDP), which groups data points with similar trends into the same segment. Combined with parallel lookup, this reduces the linear search range while maintaining high lookup speed. 2) We propose a more fine-grained binarization method, Base-Index Representation (BI), which approximates floating-point operations using integers. This method further reduces the search range without increasing model complexity. Experimental results show that EffiMatch reduces the linear search range by 26.84% using lower-complexity RMI models, which improves lookup speed by up to 6× and reduces construction time by up to 4 orders of magnitude compared to state-of-the-art LPC methods.
Lida Liao, Jiashuo Yu, Longlong Zhu, Hongyan Liu 0001, Dong Zhang 0010, Xiang Chen 0017, Chunming Wu 0001
ICNP9
2025 SkewTide: Bridging Efficiency and Tail Latency in Key-Value Stores via Kernel Re-Architecture
abstract
Key-value stores are the key building block of online services such as e-commerce. However, highly skewed workloads (i.e., skewed access frequency and request size) may cause severe load imbalance and head-of-line blocking, resulting in significant performance penalty (e.g., low throughput and high latency). Existing works mitigate skewed workloads, but often struggle to balance CPU efficiency with low tail latency or require specialized hardware. In this paper, we present SkewTide, an in-kernel architecture that breaks this trade-off through workload-aware request pre-processing and bypassing unnecessary network stack operations. Moreover, SkewTide carefully orchestrates size-aware parsing, sharding, caching, and queueing in the kernel. Both designs enable efficient CPU multiplexing and preserve low tail latency without specialized hardware. We implement SkewTide as an out-of-the-box framework using eBPF, making it readily deployable in existing key-value store infrastructure. Evaluation with YCSB traces shows that SkewTide achieves up to 8.1× higher throughput, 37% lower 99th-percentile latency, and 32% lower CPU usage compared to existing systems.
Jinghan Zu, Zhengyan Zhou, Lingfei Cheng, Zhongfeng Jin, Haifeng Zhou, Chunming Wu 0001
ICNP6
2025 Planning ECMP Paths with Minimal Overlap for Efficient Cross-Host Collective Communications
abstract
In data center networks, cross-host collective communications (CC) for LLM training often suffer from ECMP's hash-based randomness which funnels flows onto overlapping spine-leaf links, creating hotspots, rank stragglers, and degraded CC efficiency. A promising yet underexplored approach is to plan cross-host paths ahead during CC initialization. This leverages host-side steering, exploiting ECMP hash linearity via lightweight packet-header modification. Assigning cross-host paths to minimize link overlap and balance load is NP-complete for large-scale networks. To address this, we propose PathPlanner, a centralized service that heuristically selects near-optimal paths with minimal spine-leaf overlaps, generates multiple valid source ports via the host-side steering, and distributes them to workers. By cycling through these ports, each flow traverses the intended path without modifying software logics. High-fidelity SimAI simulations with realistic LLM workloads demonstrate that PathPlanner significantly reduces link overlap and straggler effects, cutting CC primitive flow completion times by up to 44.7 % and execution times by up to 21.35 %. In 32-rank Mixtral training, it shortens per-iteration runtimes by 1.6–2.1s, yielding estimated cumulative savings of 2.65-3.52 days over a complete training run.
Chunming Wu 0001, Qiang Yang 0004, Bing Hu 0002
ICPADS2
2025 Insvdf: Interface-State-Aware Virtual Device Fuzzing
abstract
Hypervisor is the core technology of virtualization for emulating independent hardware resources for each virtual machine. Virtual devices serve as the main interface of the hypervisor, making the security of virtual devices crucial, as any vulnerabilities can impact the entire virtualization environment and pose a threat to the host machine's security. Direct Memory Access (DMA) is the interface of virtual devices, enabling communication with the host machine. Recently, many efforts have focused on fuzzing against DMA to discover the hypervisor's vulnerabilities. However, the lack of sensitivity to the DMA state causes these efforts to be hindered in efficiency during fuzzing. Specifically, there are two main issues: the uncertain interaction moment and the unclear interaction depth. In this paper, we introduce InSVDF, a DMA interface stateaware fuzzing engine. InSVDF first models the intra-interface state of the DMA interface and incorporates an asynchronyaware state snapshot mechanism along with a depth-aware seed preservation mechanism. To validate our approach, we compare InSVDF with a state-of-the-art fuzzer. The results demonstrate that InSVDF significantly enhances vulnerability discovery speed, with improvements of up to 24.2 x in the best case. Furthermore, InSVDF has identified 2 new vulnerabilities, one of which has been assigned a CVE ID.
Zexiang Zhang, Yiming Tao, Zulie Pan, Cheng Tu, Min Zhang 0054, Yang Li 0215, Yi Shen 0012, Chunming Wu 0001
ICSE10
2025 TurboCache: Empowering Switch-Accelerated Key-Value Caches with Accurate and Fast Cache Updates
abstract
Recent key-value (KV) caches are offloaded to programmable switches to offer high query processing performance. However, they suffer from both low accuracy in hot key detection and high latency in cache updates due to the strict limitations on switch registers. We propose TurboCache, a switch-accelerated KV cache with accurate hot key detection and fast cache updates. Our key idea is to leverage the switch recirculation capability to build a novel data structure that caches hot KV pairs. With this hardware-compatible cache data structure, TurboCache designs efficient data plane algorithms that accurately detects new hot keys and quickly updates its cache entirely within switch ASIC pipelines. We have implemented TurboCache on a${64}\times {100}$Gbps Tofino switch. Testbed results indicate that TurboCache improves the hot key detection accuracy and decreases the cache update latency of existing KV caches by several orders of magnitude.
Xiang Chen 0017, Longlong Zhu, Linying Zheng, Lingfei Cheng, Jianshan Zhang, Xu Yang 0002, Dong Zhang 0010, Xuan Liu 0006, Xiaoming Lu, Xun Yi, Ibrahim Khalil 0001, Albert Y. Zomaya, Haifeng Zhou, Chunming Wu 0001
INFOCOM14
2025 Monica: Towards Scalable Distributed System Verification by Programmable Switch-Based Testing
abstract
Data correctness in distributed systems is ensured by data consistency, where consistency is achieved by consensus algorithms. To safeguard data consistency, current testing tools use stress testing methods to examine consensus algorithms. However, existing tools are unable to simulate the situation under high traffic and suffer from excessive verification time. In this paper, we propose Monica, a scalable and efficient verification framework. Its key idea is to leverage the programmable switch to verify consensus algorithms. Specifically, Monica provides a set of primitives that researchers can invoke. Then, the control server recognizes the primitives and automatically configures the data plane. After that, the programmable switch collaborates with the control server to complete the verification. Experimental results show that Monica can generate traffic at the rate of Tbps level while keeping the computational and memory consumption of the programmable switch under 11.87%. Compared to existing testing tools, Monica increases the verification speed by up to 3.13 times. Further, Monica improved accuracy by 35.71% in high-traffic scenarios over other tools.
Jiashuo Yu, Longlong Zhu, Dong Zhang 0010, Lida Liao, Rongbang Wu, Xiang Chen 0017, Chunming Wu 0001
IWQoS9
2025 Polyx: Accelerating Verification of Traffic Migration in Large-Scale BGP Networks
abstract
In BGP networks, traffic migration verification ensures the scalability and reliability of the network during configuration changes. However, previous approaches suffer from low scalability and high computational overhead. In this poster, we propose Polyx, a framework for accelerating verification of traffic migration in large-scale BGP networks. Its key idea is to leverage hardware parallelism with a deterministic serialization algorithm to enhance state machine techniques. We implement the Polyx prototype and evaluate it on our built testbed. The experimental results demonstrate that Polyx achieves up to 46× overall speedup, 36× in state machine construction, and 131× in equivalence verification with minimal FPGA resource usage.
Rongbang Wu, Longlong Zhu, Jiashuo Yu, Dong Zhang 0010, Hongyan Liu 0001, Zongye Lin, Lida Liao, Xiang Chen 0017, Chunming Wu 0001
IWQoS11
2025 TBNN: Lookup Tables-Based Optimization for in-Network Binary Neural Networks
abstract
Binary Neural Network (BNN) is a meaningful machine learning model on the data plane. However, due to the chip limitations, the scalability, especially the number of hidden layers in one pipeline, is limited. For better inference performance, existing methods reuse the hidden layers through packet recirculations. Recirculations lead to poor processing latency. Additionally, the simplified operations in the in-network BNN model restrict the flexibility of itself, which results in the unarbitrary input length of neurons for more the additional resource consumption than normal BNN model. In this paper, we present TBNN, an optimized in-network BNN model that achieves both scalability and flexibility. This approach eliminates deployment constraints while maximizing hardware utilization, advancing the feasibility of complex BNN models on resourcelimited data planes. By replacing computational bottleneck actions with Lookup Tables (LUTs), TBNN enables at most$4 \times$more neurons per pipeline and reduces per-packet latency by 50% through minimized recirculation. LUT-based implementation supports pruning operations, trading an accuracy loss of$\mathbf{1. 6 9 \%}$for saving about$\mathbf{2 4 \%}$instructions.
Shaowei Xu, Shengrui Lin, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001
IWQoS6
2025 Handling Data Plane Program Deployment Dynamics with High-Quality Generative Diffusion Models
abstract
Deploying data plane programs across the network is typically formulated as a mixed-integer programming task, leading to a long execution time. In response, existing studies carefully tailor heuristics for specific task properties such as objectives. However, they suffer from poor solution quality under dynamic task deployment since they overfit specific task properties. Recently, generative diffusion models have been widely adopted in network optimizations due to their strong adaptability and generalization. Accordingly, in this poster, we propose a diffusion model-based framework for data plane program deployment tasks. Our key idea is to leverage the reverse denoising process of diffusion models to react to dynamic task changes at runtime while maintaining high solution quality. Preliminary results on our testbed show that we reduce latency by 66.67% and resource overhead by 58.62% during dynamic deployment.
Longlong Zhu, Jiashuo Yu, Xiang Chen 0017, Qing Shu, Zedi Chen, Zhifan Jiang, Qun Huang 0001, Xuan Liu 0006, Dong Zhang 0010, Chunming Wu 0001
IWQoS10
2025 NDIF: A distributed framework for efficient in-network neural network inference
Shengrui Lin, Shaowei Xu, Binjie He, Hongyan Liu 0001, Dezhang Kong, Xiang Chen 0017, Dong Zhang 0010, Chunming Wu 0001, Ming Li 0056, Xuan Liu 0006, Yuqin Wu, Muhammad Khurram Khan
Comput. Secur.8
2025 FlowTracker: A refined and versatile data plane measurement approach
Chunming Wu 0001, Zhengyan Zhou, Di Wang 0003, Dezhang Kong, Muhammad Khurram Khan, Xuan Liu 0006
J. Netw. Comput. Appl.2
2025 Elastically Scaling Control Channels in Network Measurement With Escala
abstract
In network measurement, data plane switches measure traffic and report events (e.g., heavy hitters) to the control plane via control channels. The control plane makes decisions to process events. However, current network measurement suffers from two problems. First, when traffic bursts occur, massive events are reported in a short time so that the control channels may be overloaded due to limited bandwidth capacity. Second, only a few events are reported in normal cases, making control channels underloaded and wasting network resources. In this paper, we propose$\textsf {Escala}$to provide the elastic scaling of control channels at runtime. The key idea is to dynamically migrate event streams among control channels to regulate the loads of these channels.$\textsf {Escala}$offers two components, including an$\textsf {Escala}$monitor that detects scaling situations based on realtime network statistics, and an optimization framework that makes scaling decisions to eliminate overload and underload situations. We have implemented a prototype of$\textsf {Escala}$on Tofino-based switches. Extensive experiments show that$\textsf {Escala}$achieves timely elastic scaling while preserving high application-level accuracy.
Hongyan Liu 0001, Xiang Chen 0017, Qun Huang 0001, Dezhang Kong, Dong Zhang 0010, Chunming Wu 0001, Xuan Liu 0006
IEEE Trans. Netw.6
2025 Toward Secure Inter-Device Coordination in Programmable Networks
abstract
In programmable networks, some networking systems coordinate data plane switches to perform in-network functions (e.g., in-band network telemetry). However, the vulnerabilities associated withinter-device coordinationremain largely unexplored and overlooked, which is highly concerning given the increasing popularity of this paradigm. In this paper, we identify three attack scenarios built upon such vulnerabilities, where attackers mislead the behaviors of networking systems. We implement 20 networking systems on Tofino-based switches and a simulator and test them against the identified attacks. Our experimental results show that our attacks severely disrupt the normal operation of these networking systems, e.g., the cache hit rate of NetCache drops by 38%. However, our analysis reveals that none of existing methods fully mitigate our attacks because they fail to verify the packets for inter-device coordination. To this end, we select characteristics from existing methods while addressing their limitations to design effective mitigation methods. Experimental results indicate that our methods perform well in mitigating our attacks and introduce acceptable overheads.
Hongyan Liu 0001, Xiang Chen 0017, Di Wang 0049, Qun Huang 0001, Dong Zhang 0010, Chunming Wu 0001, Xuan Liu 0006
IEEE Trans. Netw.6
2024 FlexPDD: Enabling Proportional Delay Differentiation Service on Programmable Switches
abstract
Quality-of-Service (QoS) guarantees are crucial for meeting the diverse performance requirements of applications in packet networks. The Proportional Delay Differentiation (PDD) model offers relative service differentiation based on the delay requirements of different traffic classes. However, implementing PDD on current hardware switches faces challenges due to the lack of inherent queuing behavior description in switch ASICs. This paper introduces FlexPDD, a dynamic and adaptive packet prioritization mechanism designed to implement the PDD model on programmable switches. FlexPDD leverages the flexibility of programmable switch to adjust the mapping between packet classes and output queues dynamically, ensuring precise control over delay differentiation. Our implementation of FlexPDD on a Barefoot Tofino switch and an NS3 simulator demonstrates its feasibility and effectiveness. The results indicate that FlexPDD successfully maintains approximate delay differentiation among service classes proportional to their delay weights, highlighting its potential as a practical solution for achieving advanced service differentiation in modern network infrastructures.
Dezhang Kong, Zhengyan Zhou, Di Wang 0003, Shuangxi Chen, Chunming Wu 0001
GLOBECOM6
2024 Enabling Source Hosts to Precisely Select Paths via ECMP Hash Linearity in Data Center Networks
abstract
In data center networks (DCNs) with high-density computing power, the path selection managed by the equal-cost multi-path (ECMP) hashing function often causes path overlap and overuse of switch ports, undermining performance. While recent approaches allow hosts to influence egress port selection via single-bit changes in packet headers, they offer limited port coverage and are restricted to determining a single hop’s egress port. To address this, we propose a host-based path selector (HPS) that enables source hosts to precisely select paths via multiple hops by making targeted multi-bit header changes. HPS is based on two key principles: (a) exploiting the predictable relationship between the relative changes to hash values to select switch egress ports effectively, and (b) applying the criteria for ensuring modified headers accurately direct packets along the desired path. HPS iteratively adjusts packet headers at each switch along the path, ensuring efficient path selection with polynomial time and constant space complexity, making it scalable for large networks. We evaluated HPS in two and three-layer DCN topologies, testing up to 1,000 paths. The results show that HPS enables precise path selection, greatly reducing path overlap compared to traditional ECMP and state-of-the-art RePaC port influence method, resulting in substantial performance improvements.
Chunming Wu 0001, Qiang Yang 0004
HPCC2
2024 SpotMon: Enabling General Hotspot Monitoring in Key-Value Stores
abstract
Key-value stores are essential to online services such as e-commerce. In key-value stores, a hotspot (i.e., frequently accessed items) may cause severe load imbalances, high response latency, and Service Level Agreement (SLA) violations. However, existing works only focus on specific types of hotspots, thus overlooking other types of hotspots and leading to blind spots. In this paper, we propose SpotMon, a system that enables general hotspot monitoring in key-value stores. Specifically, we (1) systematically identify the generality requirements of hotspot monitoring from existing works, (2) formulate general hotspot monitoring as an arbitrary partial spot query problem, (3) measure the hotness of hotspot candidates with a new vector expression, (4) propose hotspot encoding, filtering, decoding, and querying to support general queries without focusing on specific hotspots, (5) leverage the in-network visibility of programmable switches to identify system-wide hotspots. Our extensive experiments indicate that SpotMon provides high accuracy (e.g., F1 score from 0.88 to 1) and enables efficient hotspot mitigations (e.g., up to$4.03 \times$MQPS).
Zhengyan Zhou, Jinhan Zu, Enhao Huang, Haifeng Zhou, Dong Zhang 0010, Xiang Chen 0017, Chunming Wu 0001
ICNP8
2024 OpenINT: Dynamic In-band Network Telemetry with Lightweight Deployment and Flexible Planning
abstract
The normal operation of data center network management tasks relies on accurate measurement of the network status. In-band Network Telemetry (INT) leverages programmable data planes to provide fine-grained and accurate network status. However, existing INT-related works have not considered the telemetry data required for dynamic adjustments of INT under uninterrupted conditions, including additions, deletions, and modifications. To address this issue, this paper proposes OpenINT, a lightweight and flexible In-band Network Telemetry system. The key innovation of OpenINT lies in decoupling telemetry operations in the data plane, using three generic sub-modules to achieve lightweight telemetry. Meanwhile, the control plane utilizes heuristic algorithms for dynamic planning to achieve near-optimal telemetry paths. Additionally, OpenINT provides primitives for defining network measurement tasks, which abstract the underlying telemetry architecture’s details, enabling network operator to conveniently access network status. A prototype of OpenINT is implemented on a programmable switch equipped with the Tofino chip. Experimental results demonstrate that OpenINT achieves highly flexible dynamic telemetry and significantly reduces network overhead.
Jiayi Cai, Tingxin Sun, Zhengyan Zhou, Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001
INFOCOM9
2024 Accelerating Sketch-based End-Host Traffic Measurement with Automatic DPU Offloading
abstract
Sketch-based traffic measurement is a crucial building block for monitoring traffic statistics and ensuring the quality of services of end-host applications. However, existing approaches for building sketches in end-hosts exhibit poor packet processing performance or high CPU consumption. In this paper, we propose MPU, which automatically offloads sketch-based measurement to the emerging hardware, DPU. MPU consists of a sketch analyzer that profiles sketch resource consumption and an optimization framework that formulates the offloading problem and maximizes sketch performance on DPU. We implement MPU on the NVIDIA BlueField DPU. Our testbed results indicate that MPU achieves 85% lower per-packet processing latency and 47% higher traffic measurement accuracy when compared to existing approaches.
Xiang Chen 0017, Wenbin Zhang 0011, Xin Yao 0008, Zizheng Wang, Hongyan Liu 0001, Qun Huang 0001, Xuan Liu 0006, Haifeng Zhou, Chunming Wu 0001
INFOCOM11
2024 TupleRadar: Accelerating Tuple Space Search in Packet Classification by Learned Index
abstract
Tuple space search(TSS)-based packet classification is the keystone of network system. Previous studies accelerate TSS by partitioning tuples, combining trees and tuples, and merging tuples. However, they do not scale with the number of rules, resulting in a high memory footprint or update time. In this paper, we propose TupleRadar, a framework for accelerating TSS while ensuring low memory footprint and fast rule updates. Our key idea is to construct learned indexes for tuples, which inherently improve the lookup speed but ensure the advantages of TSS. Specifically, TupleRadar builds orderly hash table-based tuples and then constructs the updatable learned index. It provides a bounded memory footprint of the index structure as well. We have evaluated TupleRadar on multiple scales rule-sets. Experimental results show that TupleRadar outperforms previous solutions, reducing 46.66% lookup time and 61.53% memory footprint on average, by up to 86.70% and 88.95%. It also performs a competitive rule update speed.
Longlong Zhu, Jiashuo Yu, Kaiwei Huang, Zhengyan Zhou, Dong Zhang 0010, Xiang Chen 0010, Chunming Wu 0001
IWQoS9
2024 CardSketch: Shift Attention for Network-wide Cardinality Telemetry
abstract
Network telemetry is an essential part of network management and infrastructure. Among them, cardinality telemetry provides statistics on network connectivity and distribution. Network-wide cardinality telemetry refers to the deployment of multiple telemetry nodes in network for cardinality estimate. This requires the deployed data structure to be mergeable, enabling the consolidation of data from different nodes. Unfortunately, existing mergeable data structures can’t simultaneously address two important criterions of cardinality telemetry: measurement accuracy and estimation interval. We propose CardSketch, aiming to adjust attention to cardinality telemetry based on changes of the network state. CardSketch incorporates a shift attention mechanism that leverages the randomness of hash functions to achieve unbiased transformations between data structures. This mechanism enables real-time selection of cardinality estimation methods based on the network’s state while preserving the original telemetry information as much as possible during the attention shift. We have implemented prototypes of CardSketch in software and hardware. Through extensive experimentation, the results demonstrate that CardSketch achieves excellent cardinality telemetry with minimal memory overhead. Even with a mere 50KB of memory space, it achieves a measurement precision of 87.75% and a measurement recall of 91.49%. Additionally, CardSketch supports multi-point aggregation and arbitrary partial key queries.
Hanze Chen, Zhengyan Zhou, Pengpai Shi, Yanni Wu, Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001
LCN9
2024 DOT: Towards Fast Decision Tree Packet Classification by Optimizing Rule Partitions
abstract
Packet classification is a crucial component of modern networks. Existing decision tree-based algorithms alleviate the rule replication problem caused by overlapping rules in the ruleset via rule partitioning. They partition the ruleset into multiple subsets based on rule characteristics to reduce rule overlaps. However, existing algorithms fail to address the overlap between rules in the same set, seriously decreasing speed and memory performance. In this paper, we propose DOT, a framework for optimizing rule partitions before constructing decision trees. Its key idea is to migrate rules in subsets based on rule overlaps and the features of heuristics used to construct trees, as well as reorganize rules aided by tuples. DOT finds out the migrated rule candidates using rule dependency graphs and heuristic features, then transforms the rule migration problem into an integer linear programming problem and solves for the optimal migration strategy. Further, we employ a tuple-assisted approach to accelerate rule matching. Experiments show that DOT enhances existing decision tree-based algorithms, improving lookup speed by 1.69 ×, reducing average 24.85% memory consumption and 31.03% decision tree depth.
Longlong Zhu, Jiashuo Yu, Linying Zheng, Dong Zhang 0010, Chunming Wu 0001
LCN7
2024 TransTuple: Toward Fast Packet Classification via Adaptive Tuple Replacement
abstract
Open vSwitch (OVS) is a widely used software switch in virtualized environments and software-defined networks. OVS uses tuple space search (TSS) for packet classification in the datapath, allowing fast network rule updates, but the increasing number of rules poses a classification performance challenge. To address this, existing methods incorporate decision trees with TSS to form a hybrid structure, enhancing classification speed. However, decision trees tend to overfit the initial ruleset, becoming unbalanced after rule updates and leading to a sharp decline in classification performance. In this paper, we propose TransTuple, a framework to optimize hybrid structures for fast packet classification under rule updates. The core idea of TransTuple is to identify bottleneck branches in decision trees that degrade performance and to replace them with lightweight tuples, providing better throughput under rule updates. These tuples maintain rules using hash tables, enabling fast updating and packet matching on bottleneck branches. We use TransTuple to optimize three state-of-the-art hybrid structured methods, i.e., CutTSS, TabTree, and MBitTree, achieving up to a 3.1x improvement in classification speed during rule updates.
Jiashuo Yu, Longlong Zhu, Rongbang Wu, Linying Zheng, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001
SECON7
2024 Eagle: Toward Scalable and Near-Optimal Network-Wide Sketch Deployment in Network Measurement
abstract
Sketches are useful for network measurement thanks to their low resource overheads and theoretically bounded accuracy. However, their network-wide deployment suffers from the trade-off between optimality and scalability: (1) Most solutions rely on mixed integer linear programming (MILP) solvers to provide the optimal decisions. But they are time-consuming and can hardly scale to large-scale deployment scenarios. (2) While heuristics achieve scalability, they deteriorate resource and performance overheads. We propose Eagle, a framework that achieves scalable and near-optimal network-wide sketch deployment. Our key idea is to decompose network-wide sketch deployment into sub-problems. Such decomposition allows Eagle to (1) simultaneously optimize switch resource consumption and end-to-end performance (retaining optimality), and (2) incorporate time-saving techniques into sub-problem solving (achieving scalability). Compared to existing solutions, Eagle improves scalability by up to 255× with negligible loss of optimality. It has also saved administrators in a production network days of efforts and reduced the operation time from O(hour) to O(second).
Xiang Chen 0017, Qingjiang Xiao, Hongyan Liu 0001, Qun Huang 0001, Dong Zhang 0010, Xuan Liu 0006, Longbing Hu, Haifeng Zhou, Chunming Wu 0001, Kui Ren 0001
SIGCOMM9
2024 P4Rex: Accelerating regular expression matching with programmable switches
Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001
Comput. Networks6
2024 Efficient service reconfiguration with partial virtual network function migration
Dongquan Liu, Zhengyan Zhou, Dong Zhang 0010, Kaiwei Guo, Yanni Wu, Chunming Wu 0001
Comput. Networks6
2024 Adversarial examples: A survey of attacks and defenses in deep learning-enabled cybersecurity systems
Mayra Alexandra Macas Carrasco, Chunming Wu 0001, Walter Fuertes
Expert Syst. Appl.2
2024 Resilient Sensor Data Dissemination to Mitigate Link Faults in IoT Networks With Long-Haul Optical Wires for Power Transmission Grids
abstract
In today’s power transmission grids, Internet-of-Things networks employ long-haul optical wires for regular sensor data dissemination to a server. Ensuring resilience against link faults is paramount to observe the grid states accurately via a process known as state estimation (SE). The accuracy is achieved by minimizing the end-to-end failure rate in packet delivery (EEFR). Current approaches focus on hop-by-hop retransmission control with in-path caching. Notably, the disruption-resilient transport protocol (DRTP) stands out for achieving the lowest EEFR. DRTP employs robust hop-by-hop retransmission and a recursive collaboration process guided by arrival timeouts. However, challenges arise in maintaining recursiveness with timeouts, leading to increased EEFR due to cache mismatch. These intensify when a hop triggers arrival timeouts, spawning retransmission instances in an unexpected sequence, which can experience an unprotected parallel race condition. To address this, we propose RSDD, a resilient mechanism for sensor data dissemination for implementing DRTP in the correct and fully verified manner. RSDD orchestrates concurrent retransmission instances, ensuring exclusive execution for the same lost packet, precisely scheduled based on timeouts. We evaluated the performance of RSDD in a simulated network that combines SE and a grid, using ndnSIM, MATPOWER, and RTDS. The results validate RSDD as a correct DRTP implementation, highlighting its exclusiveness and quality-of-service performance. RSDD achieves an EEFR of 2.44% and an average end-to-end packet delivery time (EEDT) of 2.7 ms during full path disruption with a 20% link loss rate in packets. Moreover, RSDD excels in enabling SE to maintain the grid observability and accuracy.
Chunming Wu 0001, Qiang Yang 0004, Yaguan Qian, Yinghui Nie
IEEE Internet Things J.2
2024 Terra: Low-latency and reliable event collection in network measurement
Hongyan Liu 0001, Xiang Chen 0017, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006, Muhammad Khurram Khan
J. Netw. Comput. Appl.7
2024 G-Fuzz: A Directed Fuzzing Framework for gVisor
abstract
gVisor is a Google-published application-level kernel for containers. As gVisor is lightweight and has sound isolation, it has been widely used in many IT enterprises [1],[2],[3]. When a new vulnerability of the upstream gVisor is found, it is important for the downstream developers to test the corresponding code to maintain the security. To achieve this aim, directed fuzzing is promising. Nevertheless, there are many challenges in applying existing directed fuzzing methods for gVisor. The core reason is that existing directed fuzzers are mainly for general C/C++ applications, while gVisor is an OS kernel written in the Go language. To address the above challenges, we propose G-Fuzz, a directed fuzzing framework for gVisor. There are three core methods in G-Fuzz, including lightweight and fine-grained distance calculation, target related syscall inference and utilization, and exploration and exploitation dynamic switch. Note that the methods of G-Fuzz are general and can be transferred to other OS kernels. We conduct extensive experiments to evaluate the performance of G-Fuzz. Compared to Syzkaller, the state-of-the-art kernel fuzzer, G-Fuzz outperforms it significantly. Furthermore, we have rigorously evaluated the importance for each core method of G-Fuzz. G-Fuzz has been deployed in industry and has detected multiple serious vulnerabilities.
Yuwei Li 0002, Shouling Ji, Xuhong Zhang 0002, Guanglu Yan, Alex X. Liu, Chunming Wu 0001, Zulie Pan
IEEE Trans. Dependable Secur. Comput.7
2024 rDefender: A Lightweight and Robust Defense Against Flow Table Overflow Attacks in SDN
abstract
The flow table is a critical component of Software-Defined Networking (SDN). However, flow tables’ limited capacity makes them highly vulnerable to flow table overflow attacks (FTOAs). Due to the low attack cost and highly flexible attack forms, it is hard to eradicate FTOAs. This paper addresses three unsolved problems for table security and proposes a robust defense accordingly. First, we reveal that the existing defenses with fixed defense speeds will cause severe packet loss when handling diverse traffic. We prove that deleting multiple rules can efficiently solve this problem and give a rigorous derivation to calculate the suitable deletion number according to the environment. Second, we illustrate that abnormal table occupancy squeezing is a constant characteristic of FTOAs regardless of attack forms. It can be used to identify attacked ports accurately in different scenarios. Third, we mathematically prove that random deletion can guarantee the continuous decrease of malicious flow rules after confirming attacked ports. It achieves fast speed and robust effectiveness in different environments. Based on these findings, we design rDefender, a robust and lightweight defense prototype. We evaluate its effect by designing diverse, powerful attacks and using real-world datasets and topology. The results demonstrate that it achieves the best overall performance compared to six existing mainstream defenses, providing stable security for switch flow tables.
Dezhang Kong, Xiang Chen 0017, Chunming Wu 0001, Yi Shen 0012, Zhengyan Zhou, Qiumei Cheng, Xuan Liu 0006, Yubing Qiu, Dong Zhang 0010, Muhammad Khurram Khan
IEEE Trans. Inf. Forensics Secur.3
2024 AdvSQLi: Generating Adversarial SQL Injections Against Real-World WAF-as-a-Service
abstract
As the first defensive layer that attacks would hit, the web application firewall (WAF) plays an indispensable role in defending against malicious web attacks like SQL injection (SQLi). With the development of cloud computing, WAF-as-a-service, as one kind of Security-as-a-service, has been proposed to facilitate the deployment, configuration, and update of WAFs in the cloud. Despite its tremendous popularity, the security vulnerabilities of WAF-as-a-service are still largely unknown, which is highly concerning given its massive usage. In this paper, we propose a general and extendable attack framework, namelyAdvSQLi, in which a minimal series of transformations are performed on the hierarchical tree representation of the original SQLi payload, such that the generated SQLi payloads can not only bypass WAF-as-a-service under black-box settings but also keep the same functionality and maliciousness as the original payload. WithAdvSQLi, we make it feasible to inspect and understand the security vulnerabilities of WAFs automatically, helping vendors make products more secure. To evaluate the attack effectiveness and efficiency ofAdvSQLi, we first employ two public datasets to generate adversarial SQLi payloads, leading to a maximum attack success rate of 100% against state-of-the-art ML-based SQLi detectors. Furthermore, to demonstrate the immediate security threats caused byAdvSQLi, we evaluate the attack effectiveness against 7 WAF-as-a-service solutions from mainstream vendors and find all of them are vulnerable toAdvSQLi. For instance,AdvSQLiachieves an attack success rate of over 79% against the F5 WAF. Through in-depth analysis of the evaluation results, we further condense out several general yet severe flaws of these vendors that cannot be easily patched.
Zhenqing Qu, Xiang Ling 0001, Ting Wang 0006, Xiang Chen 0017, Shouling Ji, Chunming Wu 0001
IEEE Trans. Inf. Forensics Secur.6
2024 Toward Scalable and Low-Cost Traffic Testing for Evaluating DDoS Defense Solutions
abstract
To date, security researchers evaluate their solutions of mitigating distributed denial-of-service (DDoS) attacks via kernel-based or kernel-bypassing testing tools. However, kernel-based tools exhibit poor scalability in attack traffic generation while kernel-bypassing tools incur unacceptable monetary cost. We propose Excalibur, a scalable and low-cost testing framework for evaluating DDoS defense solutions. The key idea is to leverage the emerging programmable switch to empower testing tasks with Tbps-level scalability and low cost. Specifically, Excalibur offers intent-based primitives to enable academic researchers to customize testing tasks on demand. Moreover, in view of switch resource limitations, Excalibur coordinates both a server and a programmable switch to jointly perform testing tasks. It realizes flexible attack traffic generation, which requires a large number of resources, in the server while using the switch to increase the sending rate of attack traffic to Tbps-level. We have implemented Excalibur on a$64\times 100$Gbps Tofino switch. Our experiments on a$64\times 100$Gbps Tofino switch show that Excalibur achieves orders-of-magnitude higher scalability and lower cost than existing tools.
Xiang Chen 0017, Hongyan Liu 0001, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006
IEEE/ACM Trans. Netw.6
2024 Hermes: Low-Overhead Inter-Switch Coordination in Network-Wide Data Plane Program Deployment
abstract
Network administrators usually realize network functions in data plane programs. They employ the network-wide program deployment that decomposes input programs into match-action tables (MATs) while deploying each MAT on a specific switch. Since MATs may be deployed on different switches, existing solutions propose the inter-switch coordination that uses the per-packet header space to deliver crucial packet processing information among switches. However, such coordination incurs non-trivial per-packet byte overhead, leading to end-to-end performance degradation. We propose, a framework that aims to minimize the per-packet byte overhead. The key idea is to formulate network-wide program deployment as a mixed-integer programming (MIP) problem with the objective of minimizing the per-packet byte overhead. Also, offers a greedy-based heuristic that solves the problem in a near-optimal and timely manner. We have implemented on Tofino switches. Compared to existing frameworks, decreases the per-packet byte overhead by 156 bytes while preserving end-to-end performance in terms of flow completion time and goodput.
Xiang Chen 0017, Hongyan Liu 0001, Qingjiang Xiao, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006, Qiang Yang 0004
IEEE/ACM Trans. Netw.8
2024 Toward Full-Coverage and Low-Overhead Profiling of Network-Stack Latency
abstract
In modern data center networks (DCNs), network-stack processing denotes a large portion of the end-to-end latency of TCP flows. So profiling network-stack latency anomalies has been considered as a crucial part in DCN performance diagnosis and troubleshooting. In particular, such profiling requires full coverage (i.e., profiling every TCP packet) and low overhead (i.e., profiling should avoid high CPU consumption in end-hosts). However, existing solutions rely on system calls or tracepoints in end-hosts to implement network-stack latency profiling, leading to either low coverage or high overhead. We propose Torp, a framework that offers full-coverage and low-overhead profiling of network-stack latency. Our key idea is to offload as much of the profiling from costly system calls or tracepoints to the Torp agent built on eBPF modules, and further to include a Torp handler on the ToR switch to accelerate the remaining profiling operations. Torp efficiently coordinates the ToR switch and the Torp agent on end-hosts to jointly execute the entire latency profiling task. We have implemented Torp on$32\times 100$Gbps Tofino switches. Testbed experiments indicate that Torp achieves full coverage and orders of magnitude lower host-side overhead compared to other solutions.
Xiang Chen 0017, Hongyan Liu 0001, Wenbin Zhang 0011, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Xuan Liu 0006, Chunming Wu 0001
IEEE/ACM Trans. Netw.8
2024 Resource-Efficient and Timely Packet Header Vector (PHV) Encoding on Programmable Switches
abstract
The programmable switch offers a limited capacity of packet header vector (PHV) words that store packet header fields and metadata fields defined by network functions. However, existing switch compilers employ inefficient strategies of encoding fields on PHV words. Their encoding wastes scarce PHV words and may result in failures when deploying network functions. In this paper, we propose Melody, a new framework that reuses PHV words for as many fields as possible to achieve resource-efficient PHV encoding. Melody offers a field analyzer and an optimization framework. The analyzer identifies which fields can reuse PHV words while preserving the original packet processing logic. The framework integrates analysis results into its encoding to offer the resource-optimal decisions. Also, to achieve timeliness at runtime, it provides a Greedy-based heuristic, which quickly solves PHV encoding and returns near-optimal results. We evaluate Melody with production-scale network functions. Our results show that Melody reduces the consumption of PHV words by up to 85%.
Xiang Chen 0017, Wenbin Zhang 0011, Hongyan Liu 0001, Jianshan Zhang, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Xuan Liu 0006, Chunming Wu 0001
IEEE/ACM Trans. Netw.10
2024 Toward Resource-Efficient and High- Performance Program Deployment in Programmable Networks
abstract
Programmable switches allow administrators to customize packet processing behaviors in data plane programs. However, existing solutions for program deployment fail to achieve resource efficiency and high packet processing performance. In this paper, we propose SPEED, a system that provides resource-efficient and high-performance deployment for data plane programs. For resource efficiency, SPEED merges input data plane programs by reducing program redundancy. Then it abstracts the substrate network into an one big switch (OBS), and deploys the merged program on the OBS while minimizing resource usage. For high performance, SPEED searches for the performance-optimal mapping between the OBS and the substrate network with respect to network-wide constraints. It also maintains program logic among different switches via inter-device packet scheduling. We have implemented SPEED on a Barefoot Tofino switch. The evaluation indicates that SPEED achieves resource-efficient and high-performance deployment for real data plane programs.
Hongyan Liu 0001, Xiang Chen 0017, Qun Huang 0001, Peiqiao Wang, Dong Zhang 0010, Chunming Wu 0001, Xuan Liu 0006, Qiang Yang 0004
IEEE/ACM Trans. Netw.7
2024 ABOI: AWGR-Based Optical Interconnects for Single-Wavelength and Multi-Wavelength
abstract
Optical interconnect can achieve a substantial increase in the number of nodes and switching capability for data centers, by virtue of their low power consumption and high bandwidth. In this paper, we propose a single-wavelength switch architecture based on two-stage AWGR for data centers. Then the single-wavelength design is extended to support multi-wavelength, which has higher scalability but lower hardware complexity and power consumption. We prove that the single wavelength architecture is internal strictly non-blocking. When the incoming traffic is externally blocked, a buffer allocation algorithm is designed to allocate feedforward and feedback FDL for the blocked packets. Based on the simulation results, the proposed switch can provide higher throughput, lower average latency, and a 0 out-of-order ratio compared to existing solutions.
Xiaoxue Yang, Bing Hu 0002, Chunming Wu 0001
IEEE/ACM Trans. Netw.4
2024 US-Byte: An Efficient Communication Framework for Scheduling Unequal-Sized Tensor Blocks in Distributed Deep Learning
abstract
The communication bottleneck severely constrains the scalability of distributed deep learning, and efficient communication scheduling accelerates distributed DNN training by overlapping computation and communication tasks. However, existing approaches based on tensor partitioning are not efficient and suffer from two challenges: 1) the fixed number of tensor blocks transferred in parallel can not necessarily minimize the communication overheads; 2) although the scheduling order that preferentially transmits tensor blocks close to the input layer can start forward propagation in the next iteration earlier, the shortest per-iteration time is not obtained. In this paper, we propose an efficient communication framework called US-Byte. It can schedule unequal-sized tensor blocks in a near-optimal order to minimize the training time. We build the mathematical model of US-Byte by two phases: 1) the overlap of gradient communication and backward propagation, and 2) the overlap of gradient communication and forward propagation. We theoretically derive the optimal solution for the second phase and efficiently solve the first phase with a low-complexity algorithm. We implement the US-Byte architecture on PyTorch framework. Extensive experiments on two different 8-node GPU clusters demonstrate that US-Byte can achieve up to 1.26x and 1.56x speedup compared to ByteScheduler and WFBP, respectively. We further exploit simulations of 128 GPUs to verify the potential scaling performance of US-Byte. Simulation results show that US-Byte can achieve up to 1.69x speedup compared to the state-of-the-art communication framework.
Yunqi Gao, Bing Hu 0002, Mahdi Boloursaz Mashhadi, A-Long Jin, Pei Xiao 0001, Chunming Wu 0001
IEEE Trans. Parallel Distributed Syst.6
2023 Self-Supervised Interest Transfer Network via Prototypical Contrastive Learning for Recommendation
abstract
Cross-domain recommendation has attracted increasing attention from industry and academia recently. However, most existing methods do not exploit the interest invariance between domains, which would yield sub-optimal solutions. In this paper, we propose a cross-domain recommendation method: Self-supervised Interest Transfer Network (SITN), which can effectively transfer invariant knowledge between domains via prototypical contrastive learning. Specifically, we perform two levels of cross-domain contrastive learning: 1) instance-to-instance contrastive learning, 2) instance-to-cluster contrastive learning. Not only that, we also take into account users' multi-granularity and multi-view interests. With this paradigm, SITN can explicitly learn the invariant knowledge of interest clusters between domains and accurately capture users' intents and preferences. We conducted extensive experiments on a public dataset and a large-scale industrial dataset collected from one of the world's leading e-commerce corporations. The experimental results indicate that SITN achieves significant improvements over state-of-the-art recommendation methods. Additionally, SITN has been deployed on a micro-video recommendation platform, and the online A/B testing results further demonstrate its practical value. Supplement is available at: https://github.com/fanqieCoffee/SITN-Supplement.
Yibin Shen, Sijin Zhou, Xiang Chen 0017, Hongyan Liu 0001, Chunming Wu 0001, Chenyi Lei, Xianhui Wei, Fei Fang 0002
AAAI6
2023 Aigis: Full-Coverage And Low-Overhead Mitigating Against Amplified Reflection DDoS Attacks
abstract
In Internet Service Provider (ISP) networks, Amplified Reflection DDoS (AR-DDoS) attack is one of the main attack categories, which launches gigabytes of traffic with little effort and minimal cost. Thus, the mitigation of AR-DDoS attacks has been considered as a crucial part. In particular, such mitigation requires full coverage (i.e., mitigating AR-DDoS attacks launched from any location) and low overhead (i.e., mitigation should avoid high latency that degrades user experience). However, existing solutions suffer from either limited coverage or high overhead. In this paper, we propose Aigis, a distributed framework that offers full-coverage and low-overhead mitigation of AR-DDoS attacks. Our key idea is to co-design top-of-rack (ToR) switches and end-hosts, which offers line-rate packet processing performance and fine-grained view inherently, to jointly execute endpoint verification. Specifically, Aigis selectively offloads mitigation operations between ToR switches and end-hosts and implements a network-wide epoch synchronization mechanism to guarantee reliable verification. It efficiently coordinates ToR switches and end-hosts to execute the entire mitigation task. We have implemented Aigis on a testbed comprising 32×100 Gbps Tofino switches. Testbed experiments indicate that Aigis achieves complete full coverage and orders of magnitude lower host-side overhead compared to existing solutions.
Tingxin Sun, Jiayi Cai, Kaiwei Guo, Dong Zhang 0010, Xiang Chen 0010, Chunming Wu 0001
GLOBECOM6
2023 MINT: Empowering Multiple Flow Definition Query for Network-Wide Measurement
abstract
Network management tasks rely on precise and fine-grained network information to make correct and appropriate decisions. These tasks (e.g., DDoS detection) require network information with multiple flow definitions to better manage the network. However, the existing works mainly focus on the query of multiple flow definitions on a single switch, without a thoughtful solution for this query in network-wide measurement. In this paper, to address this problem, we overcome several challenges and propose MINT, a system that enables the query for multiple flow definitions in network-wide measurement. The key insights of MINT are: deploying MFSketch to measure multiple flow definitions information on the switch, cutting MFSketch into fixed-size slices, and using in-band telemetry (INT) to carry the slice to the analyzer. Therefore, after the analyzer collects and reorganizes the slices, network operators can query multiple flow definitions information of the whole network for various network management tasks. We implemented a prototype of MINT on a Barefoot Tofino switch. Experimental results show that MINT provides reliable transmission and consistency guarantees while only using switch resources comparable to state-of-the-art works, with less than 1% additional network overhead. Additionally, MFSketch provides accurate measurements for multiple flow definitions query, outperforming other solutions in both accuracy and F1 score.
Jiayi Cai, Zhengyan Zhou, Tingxin Sun, Jiashuo Yu, Longlong Zhu, Chengze Li, Dong Zhang 0010, Chunming Wu 0001
ICC9
2023 Halia: Toward Full-Coverage Network Function Offloading in the Data Plane
abstract
Offloading network functions (NFs) to data plane switches brings remarkable performance benefits. In such offloading, NFs are required to process all the flows of interest (i.e., full coverage) to preserve the quality of services. However, existing solutions fail to guarantee full coverage for NFs. Thus, NFs may miss some essential flows, leading to accuracy drops. In this paper, we propose Halia, a framework that makes NF offloading decisions while ensuring full coverage for NFs. Specifically, Halia formulates the problem of NF offloading as an optimization problem. It encodes the requirement of full coverage as a constraint. Thus, its decisions activate enough NF instances in the substrate network to achieve full coverage.. We have implemented Halia and conducted experiments under multiple realistic network topologies to evaluate Halia. The experimental results indicate that compared to existing solutions, Halia achieves full coverage and high scalability in large-scale networks.
Hongyan Liu 0001, Xiang Chen 0017, Qingjiang Xiao, Kaiwei Guo, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001
ICC8
2023 Traffic-Aware Fast Reroute Mechanism Exploiting Disjoint Subpaths for Named Data Networks
abstract
The popularized named data network (NDN) needs a fast reroute (FRR) to improve resilience in dealing with the norm of link faults. The resilience of the existing disjoint arborescences and paths based approaches is restricted in reroute options, due to the limited number of disjoint end-to-end routes available. Such restrictions can be weakened by exploiting of the disjoint subpaths (DSP) between two intermediate nodes in the primary path (PP), as what DSP can link together is not only the two ends. In this paper, we propose a novel NDN based traffic-aware fast reroute (TA-FRR) mechanism that exploits a new resilient subgraph (RSG) route involving multiple DSPs. TA-FRR proactively constructs an RSG with scalable computational complexity to maximize the resilience and to avoid a too high reroute stretch. In the RSG, TA-FRR efficiently computes the shortest detour path for the disconnected components actively identified according to packet arrival timeouts. We evaluated the performance of TA-FRR in the ndnSIM simulator using synthsized Internet topologies with different degrees of link faults. The results show that TA-FRR can efficiently recover serious link faults in PP with a scalable RSG construction and outperforms the existing approaches with a significantly higher resilience and the maximum stretch similar to the arborescences approach.
Chunming Wu 0001, Qiang Yang 0004, Yinghui Nie
ICC2
2023 MiCuts: Combing Bit-Based Cutting and Splitting for Efficient Packet Classification
abstract
Packet classification is a crucial component in computer networking. To achieve high throughput and low memory consumption, existing solutions apply different heuristics in each construction stage to build efficient decision trees. However, previous studies divide the tree construction process based on the scale of rule subsets which is indirect to the performance goal, leading to massive rule replication and high tree depth. In this paper, we propose MiCuts, a fine-grained framework for packet classification with both high speed and low memory footprint. Its key idea is directly utilizing rule replication and tree depth to divide the tree-building process into three stages, each with suitable optimization goals. First, it partitions rules and builds shallow semi-trees without rule replication via selecting effective bits. Second, it transforms the switching problem of heuristics into an ILP problem and aims to minimize memory consumption while ensuring high lookup speed. Third, it merges some nodes to eliminate memory explosion caused by splitting, where MiCuts combines splitting and linear search. Extensive experimental results on ClassBench show that MiCuts outperforms state-of-the-art approaches, improving lookup speed by 1.71× while reducing memory footprint by 74.4% on average.
Longlong Zhu, Jiashuo Yu, Linying Zheng, Jinfeng Pan, Zhengyan Zhou, Hanze Chen, Dong Zhang 0010, Xiang Chen 0010, Chunming Wu 0001
ICC10
2023 Generating Routes Less Fragile to Improve Disruption-Resilient Transport Protocol for Phasor Data in Optical IoT Networks of Power Transmission Grids
abstract
In the modern power transmission grid, the optical IoT network faces challenges dealing with random fault-prone links during phasor data transport. This can be mitigated by the recently developed routes structured as a multipath sub-graph (MPSG) that could be generated less fragile due to added link redundancy. These MPSGs consist of a primary path (PP) that spans the route, with two hops within the PP connected via redundant subpaths (RSPs). Strengthening MPSGs is crucial for enhancing the disruption-resilient transport protocol (DRTP) and achieving the lowest possible end-to-end packet delivery failure rate (EEFR). DRTP has been successful thanks to its robust and recursive hop-by-hop retransmission process, achieving the lowest EEFR to date. However, fragile MPSGs introduce a risk of retransmission failures, significantly impacting EEFR. Generating MPSGs with minimal fragility is complex, given its NP-complete complexity. To tackle this challenge, we present the HeuMPSG algorithm, designed to minimize failure probability while adhering to constraints on end-to-end packet delivery time (EEDT). During the generation process, HeuMPSG employs heuristic and efficient candidate PP selection, searching for maximally disjoint RSPs, hop by hop. Our evaluation of HeuMPSG alongside DRTP in real grid networks using ndnSIM reveals promising results. Compared to the original algorithm of DRTP, HeuMPSG reduces fragility by an average of 171.51% and operates up to 7.07 times faster. It also significantly reduces the mean EEFR and EEDT by a maximum of 29.11% and 2.41 seconds, respectively. With HeuMPSG, DRTP’s resilience stands out with significantly lower EEFR and EEDT compared to other solutions.
Yinghui Nie, Chunming Wu 0001
ICPADS5
2023 Excalibur: A Scalable and Low-Cost Traffic Testing Framework for Evaluating DDoS Defense Solutions
abstract
To date, security researchers evaluate their solutions of mitigating denial-of-service (DDoS) attacks via kernel-based or kernel-bypassing testing tools. However, kernel-based tools exhibit poor scalability in attack traffic generation while kernel-bypassing tools result in unacceptable monetary cost. We propose Excalibur, a scalable and low-cost testing framework for DDoS defense solutions. The key idea is to leverage the programmable switch to perform testing tasks with Tbps-level scalability and low cost. Specifically, Excalibur coordinates both a server and a programmable switch to jointly perform testing tasks. It realizes flexible attack traffic generation, which requires a large number of resources, in the server while using the switch to increase the sending rate of attack traffic to Tbps-level. Our experiments on a 64×100Gbps Tofino switch show that Excalibur achieves orders-of-magnitude higher scalability and lower cost than existing tools.
Xiang Chen 0017, Hongyan Liu 0001, Tingxin Sun, Qun Huang 0001, Dong Zhang 0010, Xuan Liu 0006, Haifeng Zhou, Chunming Wu 0001
INFOCOM9
2023 Melody: Toward Resource-Efficient Packet Header Vector Encoding on Programmable Switches
abstract
The programmable switch offers a limited capacity of packet header vector (PHV) words that store packet header fields and metadata fields defined by network functions. However, existing switch compilers employ inefficient strategies of encoding fields on PHV words. Their encoding wastes scarce PHV words and may result in failures when deploying network functions. In this paper, we propose Melody, a new framework that reuses PHV words for as many fields as possible to achieve resource-efficient PHV encoding. Melody offers a field analyzer and an optimization framework. The analyzer identifies which fields can reuse PHV words while preserving the original packet processing logic. The framework integrates analysis results into its encoding to offer the resource-optimal decisions. We evaluate Melody with production-scale network functions. Our results show that Melody reduces the consumption of PHV words by up to 85%.
Xiang Chen 0017, Hongyan Liu 0001, Qingjiang Xiao, Jianshan Zhang, Qun Huang 0001, Dong Zhang 0010, Xuan Liu 0006, Chunming Wu 0001
INFOCOM8
2023 P4CTM: Compressed Traffic Pattern Matching Based on Programmable Data Plane
abstract
Pattern matching is an important technology applied to many security applications. Most network service providers choose to compress network traffic for better transmission, which brings the challenges of compressed traffic matching. However, existing works focus on improving the performance of uncompressed traffic matching or only realize the compressed traffic matching on end-host that can not keep pace with the dramatic increase in traffic. In this paper, we present P4CTM, a proof-of-concept method to conduct efficient compressed traffic matching on the programmable data plane. P4CTM uses the two-stage scan scheme to skip some bytes of compressed traffic, the 2-stride DFA combines with the compression algorithm to condense the state space, and the wildcard match to downsize the match action tables in the programmable data plane. The experiment indicates that P4CTM skips 83.10% bytes of compressed traffic, condenses the state space by order of magnitude, and reduces most of the table entries.
Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001
ISCC6
2023 DTRadar: Accelerating Search Process of Decision Trees in Packet Classification
abstract
Packet classification is an essential part of computer networks. Existing algorithms propose a partition process to address the memory explosion problem of the decision tree algorithm caused by the huge number of rules with multiple fields. However, the search process requires traversing multiple trees generated by the partition, which reduces the search efficiency. The existing algorithms take simple approaches to optimize the search process, which is low efficiency or high hardware overhead. In this paper, we propose DTRadar, a framework for expediting the decision tree packet lookup process. Its key idea is building an abstract One-Big-Tree(OBT) for multiple decision trees by establishing the middle data structure. DTRadar considers each decision tree as a splittable tree and organizes these subtrees by intermediate data structures. Extensive experiments show that DTRadar benefits existing decision tree-based solutions in classification time by 61.60%, and the memory footprint only increased by 4.21% on average.
Jiashuo Yu, Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001
ISCC5
2023 In-band Network Telemetry Manipulation Attacks and Countermeasures in Programmable Networks
abstract
In-band Network Telemetry (INT) is a widely used monitoring framework in modern large-scale networks that provides fine-grained visibility into network conditions by inserting telemetry data into packets. However, this mechanism also introduces new vulnerabilities that malicious attackers can exploit. In this paper, we present four In-band Network Telemetry Manipulation Attacks that take advantage of INT's weakness, demonstrating that attackers can cause severe damage with little effort by manipulating INT packets. To address this issue, we design SecureINT, a novel INT prototype that ensures confidentiality and integrity for INT packets. To meet the stringent computational requirements of programmable switches, we comprehensively analyze possible attacks on the deployed encryption/hash algorithms and modify them accordingly without compromising their security. According to the experiments, SecureINT can be deployed on programmable switches using a single pipeline, providing encryption and integrity verification for INT packets with minimal overhead.
Dezhang Kong, Zhengyan Zhou, Yi Shen 0012, Xiang Chen 0017, Qiumei Cheng, Dong Zhang 0010, Chunming Wu 0001
IWQoS7
2023 Vulnerabilities and Attacks of Inter-device Coordination in Programmable Networks
abstract
In programmable networks, some networking systems coordinate data plane switches to realize in-network functions (e.g., in-band network telemetry). However, the vulnerabilities of inter-device coordination are still largely unknown and neglected, which is highly concerning given the increasing popularity of this paradigm. In this paper, we identify three attack scenarios built upon such vulnerabilities, where attackers mislead the behaviors of networking systems that exploit inter-device coordination to execute in-network functions. We implement 20 existing networking systems on Tofino-based switches and a simulator, and attack these systems with the identified attacks. The experimental results indicate that our attacks significantly interfere with the normal operations of the selected networking systems, e.g., the cache hit rate of NetCache drops 38%. Our analysis also demonstrates that none of existing methods can fully mitigate our attacks since they fail to verify the packets for inter-device coordination.
Hongyan Liu 0001, Xiang Chen 0017, Yi Shen 0012, Qun Huang 0001, Zhengyan Zhou, Dong Zhang 0010, Chunming Wu 0001
IWQoS7
2023 AMF: Efficient Browser Interprocess Communication Fuzzing
abstract
With the popularity of computers and mobile devices and the development of the Internet, browsers (applications used to retrieve and display information resources on the World Wide Web) are often included by default and have become an indispensable software. Therefore, research on browser security issues is essential for protecting information assets. Among many browsers in the industry, Chrome, as a cross-platform web browser developed by Google, occupies a large market share in desktop browsers, and its security risks are further amplified as its kernel is used by many other browsers. Therefore, the research on the security issues of Chrome browser is critical for browser security.This paper focuses on the vulnerability detection of the process communication interface in Chrome browser, and designs and implements a fuzzing framework, auto-mojo-fuzz (AMF). The fuzzing process mainly designs a sample optimization technique to ensure the effectiveness of input samples and improve the efficiency of fuzzing. After implementing the AMF solution, we evaluate the generated test samples to demonstrate the effectiveness of the sample optimization technique. We also prove the possibility of discovering more vulnerabilities with AMF, and tests it with the latest version of Chrome browser, finding five unique crashes, four of which are verified as security vulnerabilities, effectively proving the automatic and efficient ability of this framework to discover vulnerabilities in the process communication interfaces in browsers.
Tianxiang Luo, Yiming Tao, Xiao Lei, Shuangxi Chen, Chunming Wu 0001
PST7
2023 Optimizing Program Deployment with libopl in Programmable Networks
abstract
Deploying data plane programs on programmable switches involves complex optimization problems that make the optimal deployment decisions. However, existing deployment frameworks only focus on deploying programs in specific domains (i.e., supporting fixed optimization requirements), resulting in poor scalability. To this end, our goal is to simplify program deployment through general high-level abstractions that capture optimization requirements. In this paper, we present libopl, a generic library that enables administrators to express various optimization requirements when deploying programs and further calculates the optimal deployment plans. Existing frameworks can also use libopl to extend their functionalities to fit more deployment scenarios. To evaluate libopl, we build a Tofino-based testbed and a simulator. Our experimental results show that libopl exhibits comparable or better scalability than stateof-the-art frameworks and only introduces negligible overhead.
Hongyan Liu 0001, Xiang Chen 0017, Yi Shen 0012, Dong Zhang 0010, Chunming Wu 0001
SECON5
2023 RFT: Toward Highly Reliable Flow Data Transmission in Network Measurement
abstract
How to satisfy the latency and reliability requirements of flow data transfer is an essential problem. To address this problem, we propose RFT, a framework that aims to satisfy the user-specified latency and reliability requirements of flow data transfer, especially in the situation where the network resources are insufficient. Firstly, we formulate the problem of satisfying the user-specified latency and reliability requirements of data transfer via mixed integer linear programming (MILP), and a heuristic algorithm is then designed to solve it in a polynomialtime. Secondly, to satisfy these requirements under insufficient network resources, we proposed a greedy-based algorithm used to select the minimum number of links added to the network, which can be deployed with low cost, especially in production networks such as data centers. Finally, we have implemented RFT on a 64$\times$100 Gbps Intel Barefoot Tofino switch. Our experimental results indicate that RFT satisfies the user-specified latency and reliability requirements in all test cases at acceptable costs, even when the network resources are insufficient.
Xiang Chen 0017, Di Wang 0003, Zhengyan Zhou, Wenhai Wang, Chunming Wu 0001, Haifeng Zhou
SECON7
2023 An Attention-Based Deep Generative Model for Anomaly Detection in Industrial Control Systems
abstract
Anomaly detection is critical for the secure and reliable operation of industrial control systems. As our reliance on such complex cyber-physical systems grows, it becomes paramount to have automated methods for detecting anomalies, preventing attacks, and responding intelligently. {This paper presents a novel deep generative model to meet this need. The proposed model follows a variational autoencoder architecture with a convolutional encoder and decoder to extract features from both spatial and temporal dimensions. Additionally, we incorporate an attention mechanism that directs focus towards specific regions, enhancing the representation of relevant features and improving anomaly detection accuracy. We also employ a dynamic threshold approach leveraging the reconstruction probability and make our source code publicly available to promote reproducibility and facilitate further research. Comprehensive experimental analysis is conducted on data from all six stages of the Secure Water Treatment (SWaT) testbed, and the experimental results demonstrate the superior performance of our approach compared to several state-of-the-art baseline techniques.
Mayra Alexandra Macas Carrasco, Chunming Wu 0001, Walter Fuertes
WEBIST2
2023 Adversarial attacks against Windows PE malware detection: A survey of the state-of-the-art
Xiang Ling 0001, Lingfei Wu 0001, Jiangyu Zhang, Zhenqing Qu, Xiang Chen 0017, Yaguan Qian, Chunming Wu 0001, Shouling Ji, Tianyue Luo, JingZheng Wu
Comput. Secur.8
2023 Code classification with graph neural networks: Have you ever struggled to make it work?
Xin Liu 0050, Qingguo Zhou, Jianwei Zhuge, Chunming Wu 0001
Expert Syst. Appl.5
2023 OF-WFBP: A near-optimal communication mechanism for tensor fusion in distributed deep learning
Yunqi Gao, Zechao Zhang, Bing Hu 0002, A-Long Jin, Chunming Wu 0001
Parallel Comput.5
2023 Towards desirable decision boundary by Moderate-Margin Adversarial Training
Xiaoyu Liang 0003, Yaguan Qian, Jianchang Huang, Xiang Ling 0001, Bin Wang 0062, Chunming Wu 0001, Wassim Swaileh
Pattern Recognit. Lett.6
2023 Automatic Performance-Optimal Offloading of Network Functions on Programmable Switches
abstract
In network function virtualization (NFV), network functions (NFs) are chained as a service function chain (SFC) to enhance NF management with low cost and high flexibility. Recent NFV solutions indicate that the packet processing performance of SFCs can be significantly improved by offloading NFs to programmable switches. However, such offloading requires a deep understanding of heterogeneous NF properties (e.g., NF resource consumption and NF performance behaviors) to achieve the maximum SFC performance. Unfortunately, none of existing solutions provide automatic analysis of these NF properties. Thus, network administrators have to manually examine the source codes of NFs and profile various NF properties by hand, which is extremely time-consuming and laborious. In this article, we propose LightNF, a novel system that simplifies NF offloading in programmable networks. LightNF automatically dissects comprehensive NF properties by means of code analysis and performance profiling while eliminating manual efforts. It then leverages its analysis results of NF properties in its SFC placement so as to make the performance-optimal offloading decisions. We have implemented LightNF on Tofino-based hardware programmable switches. We perform extensive experiments to evaluate LightNF with a real-world testbed and large-scale simulation. Our experiments show that LightNF outperforms existing solutions with an orders-of-magnitude reduction in per-packet processing latency and 9.5× improvement in SFC throughput.
Xiang Chen 0017, Hongyan Liu 0001, Dong Zhang 0010, Zili Meng, Qun Huang 0001, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006, Qiang Yang 0004
IEEE Trans. Cloud Comput.7
2023 Mangling Rules Generation With Density-Based Clustering for Password Guessing
abstract
Rule-based password generation is one of the most effective and often employed techniques in the highly compute-intensive password recovery process. However, it is challenging to design and maintain a practical password mangling ruleset, which is a time-consuming task requiring specialized expertise. This paper therefore introduced MDBSCAN (Modified Density-Based Spatial Clustering of Applications with Noise), a novel density-based cluster approach in machine learning, to build an automatic password mangling rule generator. To evaluate the proposed method, cross-checks across 4 different real-world password datasets leaked from popular Internet services and applications are adopted. The results indicate that the proposed generator could produce high-quality mangling rules with a better hit rate and enhance current mangling rules by identifying hidden or omitted rules. The proposed approach also shows strong interpretability and computational efficiency. When examining the RockYou password dataset with the top 77 rules, the hit rate may rise by 11% to 104% proportionally to other well-known solutions. Furthermore, by combining the top 77 rules generated by MDBSCAN with those from other rulesets, 3–12.67% more real-world passwords can be retrieved.
Shunbin Li, Ruyun Zhang 0001, Chunming Wu 0001, Hanguang Luo
IEEE Trans. Dependable Secur. Comput.4
2023 Stalker Attacks: Imperceptibly Dropping Sketch Measurement Accuracy on Programmable Switches
abstract
Due to limited memory usage and provably high accuracy, sketches running on programmable switches have been commonly used by the literature for network measurement. However, their vulnerabilities are still largely unknown and neglected, which is highly concerning given the increasing popularity of network measurement. In this paper, we identify the Stalker attacks, where attackers aim to degrade the accuracy of sketches running on programmable switches. More precisely, attackers tamper with some sketch operations during sketch deployment atop programmable switches. At runtime, the tampered sketch will record highly inaccurate flow data, which degrades measurement accuracy. We implement Stalker attacks on Tofino switches. The results indicate that Stalker attacks significantly drop the accuracy of network management applications, e.g., reducing the F1 score of heavy hitter detection to zero. However, our analysis indicates that none of existing methods can detect Stalker attacks since they can hardly verify the correctness of sketch operations. Finally, we analyze potential defense mechanisms and identify challenges to enable further research in this context.
Xiang Chen 0017, Hongyan Liu 0001, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006, Muhammad Khurram Khan
IEEE Trans. Inf. Forensics Secur.6
2023 Multilevel Graph Matching Networks for Deep Graph Similarity Learning
abstract
While the celebrated graph neural networks (GNNs) yield effective representations for individual nodes of a graph, there has been relatively less success in extending to the task of graph similarity learning. Recent work on graph similarity learning has considered either global-level graph-graph interactions or low-level node-node interactions, however, ignoring the rich cross-level interactions (e.g., between each node of one graph and the other whole graph). In this article, we propose a multilevel graph matching network (MGMN) framework for computing the graph similarity between any pair of graph-structured objects in an end-to-end fashion. In particular, the proposed MGMN consists of a node-graph matching network (NGMN) for effectively learning cross-level interactions between each node of one graph and the other whole graph, and a siamese GNN to learn global-level interactions between two input graphs. Furthermore, to compensate for the lack of standard benchmark datasets, we have created and collected a set of datasets for both the graph-graph classification and graph-graph regression tasks with different sizes in order to evaluate the effectiveness and robustness of our models. Comprehensive experiments demonstrate that MGMN consistently outperforms state-of-the-art baseline models on both the graph-graph classification and graph-graph regression tasks. Compared with previous work, multilevel graph matching network (MGMN) also exhibits stronger robustness as the sizes of the two input graphs increase.
Xiang Ling 0001, Lingfei Wu 0001, Saizhuo Wang, Tengfei Ma 0001, Fangli Xu, Alex X. Liu, Chunming Wu 0001, Shouling Ji
IEEE Trans. Neural Networks Learn. Syst.7
2023 Toward Low-Latency and Accurate State Synchronization for Programmable Networks
abstract
Programmable switches empower stateful packet processing, in which incoming packets continuously update states in the data plane, while applications in the control plane read and write states. However, since the data plane and control plane are separated, a consistent view of states in both planes is required for stateful packet processing. Existing approaches suffer from either high latency or low accuracy. In this paper, we propose ApproSync, a framework that offers approximate state synchronization with low latency and high accuracy. To achieve low latency, ApproSync directly transfers states between switch ASICs and the control plane by bypassing switch operating systems. To achieve high accuracy, ApproSync utilizes the resources in the switch ASIC to realize rate control in state synchronization, such that it avoids potential state loss. It also bounds the divergence between the states in the data plane and that in the control plane under limited link capacity. We prototype ApproSync on Barefoot Tofino switches. The experimental results indicate that compared to existing approaches, ApproSync achieves order-of-magnitude latency reduction while maintaining high accuracy of state synchronization. Also, our experiments demonstrate that ApproSync provides significant latency benefits to existing network management applications and well preserves high application-level accuracy.
Xiang Chen 0017, Hongyan Liu 0001, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006, Qiang Yang 0004
IEEE/ACM Trans. Netw.6
2023 Eliminating Control Plane Overload via Measurement Task Placement
abstract
Recent efforts in network measurement place measurement tasks on programmable switches to measure high-speed traffic. These tasks extract flow data, i.e., events, from packets and send events to the control plane. However, the tasks may generate massive events in a short time. In this context, the links transferring events to the control plane and the control plane servers that handle events may be overloaded, i.e., control plane overload. None of existing solutions can eliminate control plane overload. In this paper, we propose MTP, a framework that eliminates control plane overload via careful measurement task placement. Our key idea is to allocate enough resources for each task during task placement to avoid control plane overload at runtime. For each task, MTP estimates its maximum possible rate of sending events to the control plane. Then its optimization framework addresses the resource restrictions of both switches and the control plane. The experiments on Tofino switches indicate that MTP outperforms existing solutions with higher accuracy in several use cases.
Xiang Chen 0017, Hongyan Liu 0001, Dong Zhang 0010, Qun Huang 0001, Haifeng Zhou, Chunming Wu 0001, Qiang Yang 0004
IEEE/ACM Trans. Netw.6
2023 Combination Attacks and Defenses on SDN Topology Discovery
abstract
The topology discovery service in Software-Defined Networking (SDN) provides the controller with a global view of the substrate network topology, allowing for central management of the entire network. Unfortunately, emerging topology attacks can poison the network topology and result in unforeseeable disasters. Although researchers have made great efforts to mitigate this problem, security hazards still exist. In this paper, we propose Invisible Assailant Attack (IAA), the first combination topology attack capable of injecting and maintaining fake links even when 12 existing defense strategies are deployed simultaneously. IAA consists of 14 attack phases that apply multiple attack strategies. Attackers skillfully disguise the attack traffic in each phase so that it looks like normal network traffic, and perform these phases in a well-planned sequence, thereby bypassing existing defenses step by step. To mitigate this attack, we propose a Route Path Verification (RPV) mechanism that orchestrates multiple defense strategies to identify fake links. According to the experiments, RPV can successfully detect IAA with low overhead: its detection completes within 1 ms while its per-flow storage consumption is only a few KB.
Dezhang Kong, Yi Shen 0012, Xiang Chen 0017, Qiumei Cheng, Hongyan Liu 0001, Dong Zhang 0010, Xuan Liu 0006, Shuangxi Chen, Chunming Wu 0001
IEEE/ACM Trans. Netw.9
2023 Performance Tuning via Lean Measurements for Acceleration of Network Functions Virtualization
abstract
Network Functions Virtualization (NFV) replaces the specialized hardware with the software-based forwarding to promise the flexibility, scalability and automation benefits. With an increasing range of applications, NFV must ultimately forward packets at rates that are comparable to the native and specialized hardware-based approaches. However, the transition packet forwarding from specialized hardware to software-based has turned out to be more challenging than expected. Thus, NFV acceleration is desperately needed to play a crucial role in the development of NFV. It is an interesting issue how to address the persistent performance tuning in a way that provides far greater flexibility to meet the demands of power. The existing developments are very inefficient, since that the uncontrollable and unanticipated performance regressions frequently occur. Besides, the environments for full system simulations are traditionally expensive and time consuming to evaluate the system performance. In this paper, we propose the methodology named as “NFV Acceleration via Lean Measurements (NALM)” to tune the performance for the NFV acceleration. NALM provides a holistic measurement approach through combining individual measures to quickly identify the bottlenecks, which can help developers with a better understanding of the design tradeoffs. Moreover, the environments for large scale performance simulation are replaced by a debugger. Thus, the waste is eliminated in terms of time consumption and infrastructure costs of the full system simulation. The systematic analysis of the multi-cores speedup ratio highlights the potential optimization space and rules. We further propose the improvement recommendations on efficient practices. The experiments evaluate the specific effects, and the relationship between the metrics and forwarding performance.
Qiang Wu 0018, Xiangping Bryce Zhai, Chunming Wu 0001, Fangliang Lou, Hongke Zhang
IEEE/ACM Trans. Netw.4
2022 TableGuard: A Novel Security Mechanism Against Flow Table Overflow Attacks in SDN
abstract
One of the most important components of Software-Defined Networking (SDN) is the flow table. It receives flow rules from the controller and uses them to handle network traffic. However, a flow table can only store a few thousand flow rules, which makes it an attractive target for table overflow attacks. These attacks force the controller to populate the flow table with a large number of meaningless flow rules, which prevents normal flows from finding matching rules and therefore having to be reported to the controller. It results in a significant latency overhead, degrading the performance of the whole network. In this paper, we present a key characteristic of table overflow attacks: even though attackers can change some critical attack parameters (e.g., attack speed) to avoid detection, proactive flows from the attacked port always occupy a stable proportion in the flow table regardless of the attack form. In light of this finding, we propose TableGuard, a novel security mechanism that uses the proactive flow rule number as the detection metric and applies a statistical approach to help filter malicious flows. The experiments demonstrate that TableGuard can mitigate both high-rate and low-rate table overflow attacks. Compared with existing defenses, TableGuard has the best mitigation performance and the minimal overhead on normal flows.
Dezhang Kong, Chunming Wu 0001, Yi Shen 0012, Xiang Chen 0017, Hongyan Liu 0001, Dong Zhang 0010
GLOBECOM2
2022 KVLB: An In-network Key-Value Load Balancer using Multi-Valued Hash
abstract
Today's Internet service architectures rely extensively on distributed key-value stores (KV-stores) to meet their performance requirements. One of the bottlenecks lies in the un-balanced load among key-value store nodes caused by the skewed workloads. With the flexibility and power of programmable switch ASICs, in-network computing becomes a propeller of application performance. This paper introduces KVLB, a new system that uses the programmable switch to achieve load balancing between key-value store nodes. KVLB uses selective replication of hot items and allocates replica node locations to the hot items through multi-value hash. This allows the switch to reroute the hot item to the replica node through a multi-valued hash calculation and requires fewer hardware resources for programmable switch ASICs. Our experimental results on an initial prototype show that KVLB improves the throughput of KV-stores at various degrees of skew and rely only on a small amount of switch hardware resources.
Xikun Zheng, Dong Zhang 0010, Zhengyan Zhou, Jingwen Lv, Chunming Wu 0001
GLOBECOM5
2022 Toward Low-Overhead Inter-Switch Coordination in Network-Wide Data Plane Program Deployment
abstract
In modern networks, administrators realize their desired functions such as network measurement in several data plane programs. They often employ the network-wide program deployment paradigm that decomposes input programs into match-action tables (MATs) while deploying each MAT on a specific programmable switch. Since MATs may be deployed on different switches, existing solutions propose the inter-switch coordination that uses the per-packet header space to deliver crucial packet processing information among switches. However, such coordination introduces non-trivial per-packet byte overhead, leading to significant end-to-end network performance degradation. In this paper, we propose Hermes, a program deployment framework that aims to minimize the per-packet byte overhead. The key idea of Hermes is to formulate the network-wide program deployment as a mixed-integer linear programming (MILP) problem with the objective of minimizing the per-packet byte overhead. In view of the NP hardness of the MILP problem, Hermes further offers a greedy-based heuristic that solves the problem in a near-optimal and timely manner. We have implemented Hermes on Tofino-based switches. Our experiments show that compared to existing frameworks, Hermes decreases the per-packet byte overhead by 156 bytes while preserving end-to-end performance in terms of flow completion time and goodput.
Xiang Chen 0017, Hongyan Liu 0001, Qingjiang Xiao, Kaiwei Guo, Tingxin Sun, Xiang Ling 0001, Xuan Liu 0006, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001
ICDCS12
2022 SketchGuide: Reconfiguring Sketch-based Measurement on Programmable Switches
abstract
Sketches enable efficient and fine-grained network measurement results with configurable resource-performance trade-offs. While sketch configurations are guided by theories, the current theoretical guidelines are either impractical or deficient for sketch configurations on emerging programmable switches. To better configure sketches on programmable switches, we (1) systematically analyze the limitations of sketch configuration guidelines on programmable hardware switches (i.e., unguided parameters, accuracy profiles, and resource budgets); (2) propose a generic and practical framework called SketchGuide to automate efficient sketch configurations on programmable switches; (3) implement SketchGuide on a Barefoot Tofino switch and compare SketchGuide to the state-of-the-art sketches by conducting extensive experiments. Our evaluations demonstrate that SketchGuide can automatically configure unguided parameters given resource budgets. SketchGuide reduces the hardware resource footprint by 52.92%-99.28% compared with current guidelines without impacting fidelity.
Zhengyan Zhou, Jingwen Lv, Lingfei Cheng, Xiang Chen 0017, Tianzhu Zhang 0002, Qun Huang 0001, Jiayu Luo, Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001
ICNP10
2022 Torp: Full-Coverage and Low-Overhead Profiling of Host-Side Latency
abstract
In data center networks (DCNs), host-side packet processing accounts for a large portion of the end-to-end latency of TCP flows. Thus, the profiling of host-side latency anomalies has been considered as a crucial part in DCN performance diagnosis and troubleshooting. In particular, such profiling requires full coverage (i.e., profiling every TCP packet handled by end-hosts) and low overhead (i.e., profiling should avoid high CPU consumption in end-hosts). However, existing solutions fully rely on end-hosts to implement host-side latency profiling, leading to low coverage or high overhead. In this paper, we propose Torp, a framework that offers full-coverage and low-overhead profiling of host-side latency. Our key idea is to offload profiling operations to top-of-rack (ToR) switches, which inherently offer full coverage and line-rate packet processing performance. Specifically, Torp selectively offloads profiling operations to the ToR switch based on switch limitations. It efficiently coordinates the ToR switch and end-hosts to execute the entire latency profiling task. We have implemented Torp on 32×100Gbps Tofino switches. Testbed experiments indicate that Torp achieves full coverage and orders of magnitude lower host-side overhead compared to other solutions.
Xiang Chen 0017, Hongyan Liu 0001, Junyi Guo, Qun Huang 0001, Dong Zhang 0010, Chunming Wu 0001, Haifeng Zhou
INFOCOM7
2022 MalGraph: Hierarchical Graph Neural Networks for Robust Windows Malware Detection
abstract
With the ever-increasing malware threats, malware detection plays an indispensable role in protecting information systems. Although tremendous research efforts have been made, there are still two key challenges hindering them from being applied to accurately and robustly detect malwares. Firstly, most of them represent executables with shallow features, but ignore their semantic and structural information. Secondly, they are primarily based on representations that can be easily modified by attackers and thus cannot provide robustness against adversarial attacks. To tackle the challenges, we present MalGraph, which first represents executables with hierarchical graphs and then uses an end-to-end learning framework based on graph neural networks for malware detection. In particular, a hierarchical graph consists of a function call graph that captures the interaction semantics among different functions at the inter-function level and corresponding control-flow graphs for learning the structural semantics of each function at the intra-function level. We argue the abstraction and hierarchy nature of hierarchical graphs makes them not only easy to capture rich structural information of executables, but also be immune to adversarial attacks. Evaluations show that MalGraph not only outperforms state-of-the-art malware detection, but also exhibits stronger robustness against adversarial attacks by a large margin.
Xiang Ling 0001, Lingfei Wu 0001, Zhenqing Qu, Jiangyu Zhang, Tengfei Ma 0001, Bin Wang 0062, Chunming Wu 0001, Shouling Ji
INFOCOM9
2022 Escala: Timely Elastic Scaling of Control Channels in Network Measurement
abstract
In network measurement, data plane switches measure traffic and report events (e.g., heavy hitters) to the control plane via control channels. The control plane makes decisions to process events. However, current network measurement suffers from two problems. First, when traffic bursts occur, massive events are reported in a short time so that the control channels may be overloaded due to limited bandwidth capacity. Second, only a few events are reported in normal cases, making control channels underloaded and wasting network resources. In this paper, we propose Escala to provide the elastic scaling of control channels at runtime. The key idea is to dynamically migrate event streams among control channels to regulate the loads of these channels. Escala offers two components, including an Escala monitor that detects scaling situations based on realtime network statistics, and an optimization framework that makes scaling decisions to eliminate overload and underload situations. We have implemented a prototype of Escala on Tofino-based switches. Extensive experiments show that Escala achieves timely elastic scaling while preserving high application-level accuracy.
Hongyan Liu 0001, Xiang Chen 0017, Qun Huang 0001, Dezhang Kong, Jinbo Sun, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001
INFOCOM8
2022 DeepThrottle: Deep Reinforcement Learning for Router Throttling to Defend Against DDoS Attack in SDN
abstract
The router throttling mechanism provides us a chance to prevent DDoS attack proactively through rate-limiting suspicious traffic before effective detection mechanism. The existing search-based and learning-based studies are highly customized to server load and can hardly cope with the constantly changing server load and unseen scenarios. To address the problem above, we design a self-evolutionary DDoS defense system, DeepThrottle, based on deep reinforcement learning (DRL) and router throttling mechanism in software defined network (SDN). The experimental results demonstrate that the DeepThrottle improves the passing ratio of normal traffic to the victim server, and reduces the server load under unseen attack scenarios compared with the state-of-the-art RL-based method.
Shuhan Chen, Congqi Shen, Chunming Wu 0001, Yi Shen 0012
IPCCC3
2022 Libra: A Stateful Layer-4 Load Balancer with Fair Load Distribution
abstract
Layer-4 (L4) load balancers (LBs) are essential for data centers, dispatching incoming connections among thousands of servers. There are two critical requirements for L4 LBs: i) load balancing fairness, i.e., the ability to assign the load to servers in proportion to their capacity; ii) per-connection consistency (PCC), i.e., all the packets belonging to the same connection should be forwarded to the same server. Howbeit, existing LBs at large sacrifice load balancing fairness to mitigate PCC violations, which cannot satisfy both requirements in the meantime. In this paper, we present Libra, a stateful L4 LB that supports fair load distribution, PCC, memory efficiency, and resilience to resource depletion attacks. Libra makes load balancing decisions resorting to the proposed Weighted M-Least-Connection First (WMLCF) mechanism considering the real-time load and available processing capacity of servers, hence enabling load balancing fairness. We prototype Libra in a programmable software switch—BMv2 using P4 language and conduct extensive flow-level simulation to evaluate the performance. The evaluation indicates that Libra significantly improves load balancing fairness (over 95%), fully ensures PCC, and reduces the average flow completion time by 17.27–42.55% compared to existing mechanisms.
Xingong Guo, Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001
IPCCC4
2022 FROD: An Efficient Framework for Optimizing Decision Trees in Packet Classification
abstract
To perform efficient packet classification, decision tree-based methods conduct decision trees via hand-tuned heuristics. Then the performance testing and optimization are executed to ensure an excellent searching speed and space overhead. Specifically, when the performance is below expectation, existing solutions attempt to optimize the algorithms, such as conducting more sophisticated heuristics. However, reconstruction or adjustment for algorithms produces an intolerable time overhead due to the long optimization period, caused by uncertain performance benefits and high pre-processing time. In this paper, we propose FROD, an efficient framework for optimizing the decision trees directly in packet classification. FROD raises a meticulous evaluation to accurately appraise decision trees constructed by different heuristics. It then seeks out the bottleneck components via a lightweight heuristic. After that, FROD searches the optimal division for inferior components considering structural constraints and characteristics of traffic distribution. Evaluation on ClassBench shows that FROD benefits existing decision tree-based solutions in classification time by 41% and memory footprint by 19% on average, and reduces classification time by up to 64%.
Longlong Zhu, Jiashuo Yu, Jiayi Cai, Jinfeng Pan, Zhigao Li, Zhengyan Zhou, Dong Zhang 0010, Chunming Wu 0001
IWQoS8
2022 Efficient middlebox scaling for virtualized intrusion prevention systems in software-defined networks
Junchi Xing, Chunming Wu 0001, Haifeng Zhou, Qiumei Cheng, Danrui Yu, Mayra Alexandra Macas Carrasco
Sci. China Inf. Sci.2
2022 A survey on deep learning for cybersecurity: Progress, challenges, and opportunities
Mayra Alexandra Macas Carrasco, Chunming Wu 0001, Walter Fuertes
Comput. Networks2
2022 GAAT: Group Adaptive Adversarial Training to Improve the Trade-Off Between Robustness and Accuracy
abstract
Adversarial training is by far one of the most effective methods to improve the robustness of deep neural networks against adversarial examples. However, the trade-off between robustness and accuracy is still a challenge in adversarial training. Previous methods used adversarial examples with a fixed perturbation budget or specific perturbation budgets for each example, which is inefficient in improving the trade-off and lacks the ability to control the trade-off flexibly. In this paper, we show that the largest element of logit, [Formula: see text], can roughly represent the minimum distance between an example and its neighboring decision boundary. Thus, we propose group adaptive adversarial training (GAAT) that divides the training dataset into several groups based on [Formula: see text] and develops a binary search algorithm to determine the group perturbation budgets for each group. Using the group perturbation budgets to perform adversarial training can fine-tune the trade-off between robustness and accuracy. Extensive experiments conducted on CIFAR-10 and ImageNet-30 show that our GAAT can achieve a more perfect trade-off than TRADES, MMA, and MART.
Yaguan Qian, Xiaoyu Liang 0003, Ming Kang 0006, Bin Wang 0062, Zhaoquan Gu, Chunming Wu 0001
Int. J. Pattern Recognit. Artif. Intell.7
2022 V-Fuzz: Vulnerability Prediction-Assisted Evolutionary Fuzzing for Binary Programs
abstract
Fuzzing is a technique of finding bugs by executing a target program recurrently with a large number of abnormal inputs. Most of the coverage-based fuzzers consider all parts of a program equally and pay too much attention to how to improve the code coverage. It is inefficient as the vulnerable code only takes a tiny fraction of the entire code. In this article, we design and implement an evolutionary fuzzing framework called V-Fuzz, which aims to find bugs efficiently and quickly in limited time for binary programs. V-Fuzz consists of two main components: 1) a vulnerability prediction model and 2) a vulnerability-oriented evolutionary fuzzer. Given a binary program to V-Fuzz, the vulnerability prediction model will give a prior estimation on which parts of a program are more likely to be vulnerable. Then, the fuzzer leverages an evolutionary algorithm to generate inputs which are more likely to arrive at the vulnerable locations, guided by the vulnerability prediction result. The experimental results demonstrate that V-Fuzz can find bugs efficiently with the assistance of vulnerability prediction. Moreover, V-Fuzz has discovered ten common vulnerabilities and exposures (CVEs), and three of them are newly discovered.
Yuwei Li 0002, Shouling Ji, Chenyang Lyu, Jianhai Chen, Qinchen Gu, Chunming Wu 0001, Raheem A. Beyah
IEEE Trans. Cybern.7
2022 EI-MTD: Moving Target Defense for Edge Intelligence against Adversarial Attacks
abstract
Edge intelligence has played an important role in constructing smart cities, but the vulnerability of edge nodes to adversarial attacks becomes an urgent problem. A so-called adversarial example can fool a deep learning model on an edge node for misclassification. Due to the transferability property of adversarial examples, an adversary can easily fool a black-box model by a local substitute model. Edge nodes in general have limited resources, which cannot afford a complicated defense mechanism like that on a cloud data center. To address the challenge, we propose a dynamic defense mechanism, namely EI-MTD. The mechanism first obtains robust member models of small size through differential knowledge distillation from a complicated teacher model on a cloud data center. Then, a dynamic scheduling policy, which builds on a Bayesian Stackelberg game, is applied to the choice of a target model for service. This dynamic defense mechanism can prohibit the adversary from selecting an optimal substitute model for black-box attacks. We also conduct extensive experiments to evaluate the proposed mechanism, and results show that EI-MTD could protect edge intelligence effectively against adversarial attacks in black-box settings.
Yaguan Qian, Yankai Guo, Qiqi Shao, Jiamin Wang 0003, Bin Wang 0062, Zhaoquan Gu, Xiang Ling 0001, Chunming Wu 0001
ACM Trans. Priv. Secur.8
2021 V-Shuttle: Scalable and Semantics-Aware Hypervisor Virtual Device Fuzzing
abstract
With the wide application and deployment of cloud computing in enterprises, virtualization developers and security researchers are paying more attention to cloud computing security. The core component of cloud computing products is the hypervisor, which is also known as the virtual machine monitor (VMM) that can isolate multiple virtual machines in one host machine. However, compromising the hypervisor can lead to virtual machine escape and the elevation of privilege, allowing attackers to gain the permission of code execution in the host. Therefore, the security analysis and vulnerability detection of the hypervisor are critical for cloud computing enterprises. Importantly, virtual devices expose many interfaces to a guest user for communication, making virtual devices the most vulnerable part of a hypervisor. However, applying fuzzing to the virtual devices of a hypervisor is challenging because the data structures transferred by DMA are constructed in a nested form according to protocol specifications. Failure to understand the protocol of the virtual devices will make the fuzzing process stuck in the initial fuzzing stage, resulting in inefficient fuzzing.
Xingwei Lin, Xuhong Zhang 0002, Yongkang Jia, Shouling Ji, Chunming Wu 0001, Xinlei Ying, Jiashui Wang
CCS6
2021 MTP: Avoiding Control Plane Overload with Measurement Task Placement
abstract
In programmable networks, measurement tasks are placed on programmable switches to keep pace with high-speed traffic. At runtime, programmable switches send events to the control plane for further processing. However, existing solutions for task placement overlook the limitations of control plane resources. Thus, excessive events may overload the control plane. In this paper, we propose MTP, a system that eliminates control plane overload via careful task placement. For each task, MTP analyzes its structure to estimate its maximum possible rate of sending events to the control plane. Then it builds an optimization framework that addresses the resource restrictions of both switches and the control plane. We have implemented MTP on Barefoot Tofino switches. The experimental results indicate that MTP outperforms existing solutions with higher accuracy across four real use cases.
Xiang Chen 0017, Qun Huang 0001, Peiqiao Wang, Hongyan Liu 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001
INFOCOM8
2021 Intelligent DDoS Detection in Botnet Combined with Packet-Level Features under SDN
abstract
Botnets provide a fundamental infrastructure for various kinds of network attacks, such as DDoS attack, etc. Detecting DDoS attack in botnet is a long-standing challenge. In this paper, we propose a novel method to detect DDoS attack in botnet through the analysis of packet forwarding and network traffic. The primary idea is to collect features from both overall network traffic and packet to describe the attack pattern and conduct a detection model with high accuracy. Firstly, apart from overall network traffic, we calculate several statistics related to looking up functions of flow tables during packet forwarding on switches to describe attack pattern. Secondly, these features are put into a deep learning model to detect DDoS attack. We perform evaluations under Software Defined Networking (SDN) paradigm. In particular, we compare the proposed method with traditional methods. The experimental results demonstrate that the proposed method is meaningful in improving the accuracy of DDoS attack detection by adding packet-level features extracted from packet forwarding.
Shuhan Chen, Congqi Shen, Danrui Yu, Yuqin Wu, Chunming Wu 0001
ISNCC5
2021 LightNF: Simplifying Network Function Offloading in Programmable Networks
abstract
In network function virtualization (NFV), network functions (NFs) are chained as a service function chain (SFC) to enhance NF management with high flexibility. Recent solutions indicate that the processing performance of SFCs can be significantly improved by offloading NFs to programmable switches. However, such offloading requires a deep understanding of NF properties to achieve the maximum SFC performance, which brings non-trivial burdens to network administrators. In this paper, we propose LightNF, a novel system that simplifies NF offloading in programmable networks. LightNF automatically dissects comprehensive NF properties (e.g., NF performance behaviors) via code analysis and performance profiling while eliminating manual efforts. It then leverages the analyzed NF properties in its SFC placement so as to produce the performance-optimal offloading. We have implemented a LightNF prototype. Our experiments show that LightNF outperforms state-of-the-art solutions with an orders-of-magnitude reduction in per-packet processing latency and 9.5× improvement in SFC throughput.
Xiang Chen 0017, Qun Huang 0001, Peiqiao Wang, Zili Meng, Hongyan Liu 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001
IWQoS10
2021 UNIFUZZ: A Holistic and Pragmatic Metrics-Driven Platform for Evaluating Fuzzers
Yuwei Li 0002, Shouling Ji, Sizhuang Liang, Wei-Han Lee, Yueyao Chen, Chenyang Lyu, Chunming Wu 0001, Raheem A. Beyah, Peng Cheng 0001, Kangjie Lu, Ting Wang 0006
USENIX Security Symposium8
2021 Machine learning based malicious payload identification in software-defined networking
abstract
Deep packet inspection (DPI) has been extensively investigated in software-defined networking (SDN) as complicated attacks may intractably inject malicious payloads in the packets. Existing proprietary pattern-based or port-based third-party DPI tools can suffer from limitations in efficiently processing a large volume of data traffic. In this paper, a novel OpenFlow-enabled deep packet inspection (OFDPI) approach is proposed based on the SDN paradigm to provide adaptive and efficient packet inspection. First, OFDPI prescribes an early detection at the flow-level granularity by checking the IP addresses of each new flow via OpenFlow protocols. Then, OFDPI allows for deep packet inspection at the packet-level granularity: (i) for unencrypted packets, OFDPI extracts the features of accessible payloads, including tri-gram frequency based on Term Frequency and Inverted Document Frequency (TF–IDF) and linguistic features. These features are concatenated into a sparse matrix representation and are then applied to train a binary classifier with logistic regression rather than matching with specific pattern combinations. In order to balance the detection accuracy and performance bottleneck of the SDN controller, OFDPI introduces an adaptive packet sampling window based on the linear prediction; and (ii) for encrypted packets, OFDPI extracts notable features of packets and then trains a binary classifier with a decision tree, instead of decrypting the encrypted traffic to weaken user privacy. A prototype of OFDPI is implemented on the Ryu SDN controller and the Mininet platform. The performance and the overhead of the proposed solution are assessed using the real-world datasets through experiments. The numerical results indicate that OFDPI can provide a significant improvement in detection accuracy with acceptable overheads.
Qiumei Cheng, Chunming Wu 0001, Haifeng Zhou, Dezhang Kong, Dong Zhang 0010, Junchi Xing
J. Netw. Comput. Appl.2
2021 Adversarial Examples versus Cloud-Based Detectors: A Black-Box Empirical Study
abstract
Deep learning has been broadly leveraged by major cloud providers, such as Google, AWS and Baidu, to offer various computer vision related services including image classification, object identification, illegal image detection, etc. While recent works extensively demonstrated that deep learning classification models are vulnerable to adversarial examples, cloud-based image detection models, which are more complicated than classifiers, may also have similar security concern but not get enough attention yet. In this paper, we mainly focus on the security issues of real-world cloud-based image detectors. Specifically, (1) based on effective semantic segmentation, we propose four attacks to generate semantics-aware adversarial examples via only interacting with black-box APIs; and (2) we make the first attempt to conduct an extensive empirical study of black-box attacks against real-world cloud-based image detectors. Through the comprehensive evaluations on five major cloud platforms: AWS, Azure, Google Cloud, Baidu Cloud, and Alibaba Cloud, we demonstrate that our image processing based attacks can reach a success rate of approximately 100 percent, and the semantic segmentation based attacks have a success rate over 90 percent among different detection services, such as violence, politician, and pornography detection. We also proposed several possible defense strategies for these security challenges in the real-life situation.
Xurong Li, Shouling Ji, Juntao Ji, Zhenyu Ren, Yushan Liu 0004, Chunming Wu 0001
IEEE Trans. Dependable Secur. Comput.7
2021 A Practical Black-Box Attack on Source Code Authorship Identification Classifiers
abstract
Existing researches have recently shown that adversarial stylometry of source code can confuse source code authorship identification (SCAI) models, which may threaten the security of related applications such as programmer attribution, software forensics, etc. In this work, we propose source code authorship disguise (SCAD) to automatically hide programmers' identities from authorship identification, which is more practical than the previous work that requires to known the output probabilities or internal details of the target SCAI model. Specifically, SCAD trains a substitute model and develops a set of semantically equivalent transformations, based on which the original code is modified towards a disguised style with small manipulations in lexical features and syntactic features. When evaluated under totally black-box settings, on a real-world dataset consisting of 1,600 programmers, SCAD induces state-of-the-art SCAI models to cause above 30% misclassification rates. The efficiency and utility-preserving properties of SCAD are also demonstrated with multiple metrics. Furthermore, our work can serve as a guideline for developing more robust identification methods in the future.
Qianjun Liu, Shouling Ji, Changchang Liu, Chunming Wu 0001
IEEE Trans. Inf. Forensics Secur.4
2021 Deep Graph Matching and Searching for Semantic Code Retrieval
abstract
Code retrieval is to find the code snippet from a large corpus of source code repositories that highly matches the query of natural language description. Recent work mainly uses natural language processing techniques to process both query texts (i.e., human natural language) and code snippets (i.e., machine programming language), however, neglecting the deep structured features of query texts and source codes, both of which contain rich semantic information. In this article, we propose an end-to-end deep graph matching and searching (DGMS) model based on graph neural networks for the task of semantic code retrieval. To this end, we first represent both natural language query texts and programming language code snippets with the unified graph-structured data, and then use the proposed graph matching and searching model to retrieve the best matching code snippet. In particular, DGMS not only captures more structural information for individual query texts or code snippets, but also learns the fine-grained similarity between them by cross-attention based semantic matching operations. We evaluate the proposed DGMS model on two public code retrieval datasets with two representative programming languages (i.e., Java and Python). Experiment results demonstrate that DGMS significantly outperforms state-of-the-art baseline models by a large margin on both datasets. Moreover, our extensive ablation studies systematically investigate and illustrate the impact of each part of DGMS.
Xiang Ling 0001, Lingfei Wu 0001, Saizhuo Wang, Tengfei Ma 0001, Fangli Xu, Alex X. Liu, Chunming Wu 0001, Shouling Ji
ACM Trans. Knowl. Discov. Data8
2021 Intrinsic Security and Self-Adaptive Cooperative Protection Enabling Cloud Native Network Slicing
abstract
With the emergence of cloud native technology, the network slicing enables automatic service orchestration, flexible network scheduling and scalable network resource allocation, which profoundly affects the traditional security solution. Security is regarded as a technology independent of the cloud native architecture in the initial design, traditional passive defense such as “reinforced” and “stacked” is relied on to achieve system security protection. The lack of intrinsic security mechanisms makes the system capability insufficient when faces the uncertain threat brought by vulnerabilities and backdoors under the ecosystem of opening-up and sharing. The static nature of existing networks and computing systems makes them easy to be compromised and hard to defend, and thus it is urgent to provide intrinsic security and proactive protection against the unpredictable attacks. To this end, this paper proposes a novel paradigm named intrinsic cloud security (iCS) from the perspective of dynamic defense. The dynamic defense provides component-level security, and has complementary and consistency with the cloud native environment. In particular, iCS introduces mimic defense and moving target defense (MTD), and makes full use of the new features introduced by cloud native to implement an intrinsic and proactive defense mechanism with acceptable costs and efficiency. The iCS paradigm achieves seamless integration and symbiosis evolution between security and cloud native. We implement a trial of iCS based on 5GC commercial system and evaluate its performance on costs, efficiency and attack success. The result shows that the iCS enhanced mode always can provide a better and more stable defense effects.
Qiang Wu 0018, Chunming Wu 0001, Xincheng Yan, Qiumei Cheng
IEEE Trans. Netw. Serv. Manag.2
2020 SRA: Switch Resource Aggregation for Application Offloading in Programmable Networks
abstract
Programmable switches empower network applications with line-rate packet processing performance by allowing the offloading of applications. However, the resource of a programmable switch is extremely limited, which significantly limits the application offloading. Existing solutions to the problem either provide poor efficiency or suffer from accuracy drop. In this paper, we propose SRA, a system that loosens switch resource constraints for application offloading via resource aggregation. SRA provides administrators with an intuitive compiler directive to customize application offloading by resource aggregation. According to compiler directives, it automatically places the program on the substrate network, while maintaining original packet processing logics. We implement a prototype of SRA in P4, and establish an experimental testbed consisting of three 32$\times$100 Gbps Barefoot switches. The experimental results indicate that SRA enhances two real-world applications with sufficient resources while maintaining high performance.
Hongyan Liu 0001, Xiang Chen 0017, Qun Huang 0001, Haifeng Zhou, Dong Zhang 0010, Chunming Wu 0001
GLOBECOM6
2020 TPDD: A Two-Phase DDoS Detection System in Software-Defined Networking
abstract
Distributed Denial of Service (DDoS) attack is one of the most severe threats to the current network security. As a new network architecture, Software-Defined Networking (SDN) draws notable attention from both industry and academia. The characteristics of SDN such as centralized management and flow-based traffic monitoring make it an ideal platform to defend against DDoS attacks. When designing a network intrusion detection system (NIDS) in SDN, how to obtain fine-grained flow information with minimal overhead to the SDN architecture is a problem to be solved. In this paper, we propose TPDD, a two-phase DDoS detection system to detect DDoS attacks in SDN. In the first phase, we utilize the characteristics of SDN to collect coarse-grained flow information from the core switches and locate the potential victim. Then we monitor the edge switches located close to the potential victim to obtain finer-grained traffic information in the second phase. The collection method of each phase fully considers the impact on the bandwidth between the controller and switches. Without modifying the existing flow rules, the collection module can obtain sufficient information about traffic. By using entropy-based and machine learning-based methods, the detection module can effectively detect anomalies and identify whether the potential victim marked in the first phase is the target of attacks. Experimental results show that TPDD can effectively detect DDoS attacks with little overhead.
Yi Shen 0012, Chunming Wu 0001, Dezhang Kong
ICC2
2020 ApproSync: Approximate State Synchronization for Programmable Networks
abstract
Programmable switches empower stateful packet processing, in which incoming packets continuously update states in the data plane, while applications in the control plane read and write states. However, as the data plane and control plane are separated, a consistent view of states in both planes is required for stateful packet processing. Existing approaches suffer from either high latency or low accuracy. In this paper, we propose ApproSync, a framework that offers approximate state synchronization with low latency and high accuracy. To achieve low latency, ApproSync directly transfers states between switch ASICs and the control plane by bypassing switch operating systems. To achieve high accuracy, ApproSync utilizes the resources in the switch ASIC to realize rate control in state synchronization, such that it avoids potential state loss. It also bounds the divergence between the states in the data plane and that in the control plane under limited link capacity. We prototype ApproSync on Barefoot Tofino switches. The experimental results indicate that compared to existing approaches, ApproSync achieves order-of-magnitude latency reduction while maintaining high accuracy.
Xiang Chen 0017, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001
ICNP5
2020 SPEED: Resource-Efficient and High-Performance Deployment for Data Plane Programs
abstract
Programmable switches allow network administrators to customize packet processing behaviors in data plane programs. However, existing solutions for program deployment fail to achieve resource efficiency and high packet processing performance. In this paper, we propose SPEED, a system that provides resource-efficient and high-performance deployment for data plane programs. For resource efficiency, SPEED merges input data plane programs by reducing program redundancy. Then it abstracts the substrate network into an one big switch (OBS), and deploys the merged program on the OBS while minimizing resource usage. For high performance, SPEED searches for the performance-optimal mapping between the OBS and the substrate network with respect to network-wide constraints. It also maintains program logics among different switches via inter-device packet scheduling. We have implemented SPEED on a Barefoot Tofino switch. The evaluation indicates that SPEED achieves resource-efficient and high-performance deployment for real data plane programs.
Xiang Chen 0017, Hongyan Liu 0001, Qun Huang 0001, Peiqiao Wang, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001
ICNP7
2020 Multi-variant network address hopping to defend stealthy crossfire attack
Chunming Wu 0001
Sci. China Inf. Sci.3
2020 Adversarial examples detection through the sensitivity in space mappings
abstract
Adversarial examples (AEs) against deep neural networks (DNNs) raise wide concerns about the robustness of DNNs. Existing detection mechanisms are often limited to a given attack algorithm. Therefore, it is highly desirable to develop a robust detection approach that remains effective for a large group of attack algorithms. In addition, most of the existing defences only perform well for small images (e.g. MNIST and Canadian institute for advanced research (CIFAR)) rather than large images (e.g. ImageNet). In this paper, the authors propose a robust and effective defence method for analysing the sensitivity of various AEs, especially in a much harder case (large images). Their method first creates a feature map from the input space to the new feature space, by utilising 19 different feature mapping methods. Then, a detector is learned with the machine‐learning algorithm to recognise the unique distribution of AEs. Their extensive evaluations on their proposed detector show that their detector can achieve: (i) low false‐positive rate (<1%), (ii) high true‐positive rate (higher than 98%), (iii) low overhead (<0.1 s per input), and (iv) good robustness (work well across different learning models, attack algorithms, and parameters), which demonstrate the efficacy of the proposed detector in practise.
Xurong Li, Shouling Ji, Juntao Ji, Zhenyu Ren, Chunming Wu 0001, Bo Li 0026, Ting Wang 0006
IET Comput. Vis.5
2019 Think That Attackers Think: Using First-Order Theory of Mind in Intrusion Response System
abstract
The intrusion response system is dedicated to automatically respond to sophisticated network intrusions, which is a sequential decision-making problem for autonomous agents. The current Markov decision process (MDP) or stochastic games based solutions suffer from several weaknesses: (i) The MDP- based approach is unable to explicitly model the opponents; (ii) The Nash equilibrium approach of stochastic games cannot handle the condition with multi equilibria. Existing studies have not considered the cognitive ability of the agent and lack of explicit opponent modeling. Inspired by recursive reasoning, this paper introduces a theory of mind (ToM)-based stochastic game-theoretic approach to reason about the beliefs and behaviors of the attackers. Each agent maintains different order ToM beliefs concerning his opponent with explicit opponent modeling. In order to accurately predict the attacker's action with nested beliefs, we utilize the Bayesian attack graph (BAG) to model multi-step attacks scenarios. In addition, the agent is allowed to learn from new information to adjust his beliefs and learning speed. Simulation results validate that ToM modeling performs well in the intrusion response system than random defense actions. Besides, a defender with first-order ToM beliefs always wins an attacker with zero-order ToM beliefs.
Qiumei Cheng, Chunming Wu 0001, Dezhang Kong
GLOBECOM2
2019 An Unsupervised Framework for Anomaly Detection in a Water Treatment System
abstract
Current Cyber-Physical Systems (CPSs) are sophisticated, complex, and equipped with networked sensors and actuators. As such, they have become further exposed to cyber-attacks. Recent catastrophic events have demonstrated that standard, human-based management of anomaly detection in complex systems is not efficient enough and have underlined the significance of automated detection, intelligent and rapid response. Nevertheless, existing anomaly detection frameworks usually are not capable of dealing with the dynamic and complicated nature of the CPSs. In this study, we introduce an unsupervised framework for anomaly detection based on an Attention-based Spatio-Temporal Autoencoder. In particular, we first construct statistical correlation matrices to characterize the system status across different time steps. Next, a 2D convolutional encoder is employed to encode the patterns of the correlation matrices, whereas an Attention-based Convolutional LSTM Encoder-Decoder (ConvLSTM-ED) is used to capture the temporal dependencies. More precisely, we introduce an input attention mechanism to adaptively select the most significant input features at each time step. Finally, the 2D convolutional decoder reconstructs the correlation matrices. The differences between the reconstructed correlation matrices and the original ones are used as indicators of anomalies. Extensive experimental analysis on data collected from all six stages of Secure Water Treatment (SWaT) testbed, a scaled-down version of a real-world industrial water treatment plant, demonstrates that the proposed model outperforms the state-of-the-art baseline techniques.
Mayra Alexandra Macas Carrasco, Chunming Wu 0001
ICMLA2
2019 Hiding and Trapping: A Deceptive Approach for Defending against Network Reconnaissance with Software-Defined Network
abstract
Network reconnaissance aims at gathering as much information as possible before an attack is launched. Meanwhile, static host address configuration facilitates network reconnaissance. Currently, more sophisticated network reconnaissance has been emerged with the adaptive and cooperative features. To address this, in this paper, we present Hiding and Trapping (HaT), which is a deceptive approach to disrupt adversarial network reconnaissance with the help of the software-defined networking (SDN) paradigm. HaT is able to hide valuable hosts from attackers and to trap them into decoy nodes through strategic and holistic host address mutation according to characteristic of adversaries. We implement a prototype of HaT, and evaluate its performance by experiments. The experimental results show that HaT is capable to effectively disrupt adversarial network reconnaissance with better deceptive performance than the existing address randomization approach.
Junchi Xing, Haifeng Zhou, Chunming Wu 0001
IPCCC4
2019 A Deep ConvNet-Based Countermeasure to Mitigate Link Flooding Attacks Using Software-Defined Networks
abstract
Recently, link flooding attacks (LFA) have been observed as a serious threat for cutting off the Internet connectivity through congesting critical links. A LFA typically utilizes legitimate and low-rate flows, which makes it extremely hard to be detected and, subsequently, to be mitigated. In this paper, we present LF-Shield, that is a deep convolutional neural network (ConvNet) based countermeasure to accurately detect and efficiently mitigate LFAs using software-defined network (SDN) paradigm. LF-Shield can identify malicious bots that launch LFA flows by extracting end-hosts' traffic features and afterwards, classifying the type of end-hosts based on deep ConvNet. Then, LF-Shield mitigates LFAs without affecting legitimate end-hosts through blocking the classified malicious bots and limiting the bandwidths of inactive or newly-accessed end-hosts. A LF-Shield prototype is implemented for evaluating its performance by several experiments. The experimental results demonstrate that LF-Shield can identify malicious bots with an accuracy of 96.4% and mitigate LFAs with the 93.1% reduction in link degradation ratio, with negligible impact on legitimate end-hosts.
Junchi Xing, Jingjing Cai, Chunming Wu 0001
ISCC4
2019 Research on Executive Control Strategy of Mimic Web Defense Gateway
abstract
Resource Scheduling Strategy on heterogeneous executives, which is the “real” service provider, is studied in this paper based on Mimic Defense Theory. Since the redundancy and heterogeneity of executives, their response times differ. In order to mitigate the Bucket Effect on the response time of heterogeneous executives in the existing mimic defense methods, this paper proposes a method based on feedback control model that executive source allocation is adjusted dynamically to improve the response time of the executive set. The comparative experiment shows that this scheduling strategy has a great improvement in response process time, comparing with the existing scheduling strategy.
Shuang-Xi Chen, Xin-Yue Jiang, Chunming Wu 0001
ISNCC4
2019 A Decentralized Multi-ruling Arbiter for Cyberspace Mimicry Defense
abstract
Cyberspace Mimicry Defense (CMD) has been widely used to achieve intrusion prevention against unknown system vulnerabilities or backdoors. The multi-ruling arbiter is a key part in CMD. This paper focuses on the problem of multi-ruling arbiter under data injection attack from the perspective of attacker and defender. We build a decentralized multi-ruling arbiter model for arbitration and introduced a standard iteration process to achieve consensus without attackers. We describe two data injection attack models for decentralized multi-ruling arbiter, namely random data injection attack and stealthy data injection attack. Further, we characterize the negative effect of the data injection attack on the performance of multi-ruling correctness. In order to mitigate the negative effect of random data injection attack, we propose a reliable multi-ruling arbitration approach based on adaptive threshold. By cutting future communication with the malicious neighbor, the decentralized multi-ruling arbiter is robust against random data injection attacks. Simulation results show that the proposed arbitration approach can effectively defend against random data injection attacks.
Congqi Shen, Shuang-Xi Chen, Chunming Wu 0001
ISNCC3
2019 DEEPSEC: A Uniform Platform for Security Analysis of Deep Learning Model
abstract
Deep learning (DL) models are inherently vulnerable to adversarial examples - maliciously crafted inputs to trigger target DL models to misbehave - which significantly hinders the application of DL in security-sensitive domains. Intensive research on adversarial learning has led to an arms race between adversaries and defenders. Such plethora of emerging attacks and defenses raise many questions: Which attacks are more evasive, preprocessing-proof, or transferable? Which defenses are more effective, utility-preserving, or general? Are ensembles of multiple defenses more robust than individuals? Yet, due to the lack of platforms for comprehensive evaluation on adversarial attacks and defenses, these critical questions remain largely unsolved. In this paper, we present the design, implementation, and evaluation of DEEPSEC, a uniform platform that aims to bridge this gap. In its current implementation, DEEPSEC incorporates 16 state-of-the-art attacks with 10 attack utility metrics, and 13 state-of-the-art defenses with 5 defensive utility metrics. To our best knowledge, DEEPSEC is the first platform that enables researchers and practitioners to (i) measure the vulnerability of DL models, (ii) evaluate the effectiveness of various attacks/defenses, and (iii) conduct comparative studies on attacks/defenses in a comprehensive and informative manner. Leveraging DEEPSEC, we systematically evaluate the existing adversarial attack and defense methods, and draw a set of key findings, which demonstrate DEEPSEC's rich functionality, such as (1) the trade-off between misclassification and imperceptibility is empirically confirmed; (2) most defenses that claim to be universally applicable can only defend against limited types of attacks under restricted settings; (3) it is not necessary that adversarial examples with higher perturbation magnitude are easier to be detected; (4) the ensemble of multiple defenses cannot improve the overall defense capability, but can improve the lower bound of the defense effectiveness of individuals. Extensive analysis on DEEPSEC demonstrates its capabilities and advantages as a benchmark platform which can benefit future adversarial learning research.
Xiang Ling 0001, Shouling Ji, Jiaxu Zou, Jiannan Wang 0002, Chunming Wu 0001, Bo Li 0026, Ting Wang 0006
IEEE Symposium on Security and Privacy5
2019 Open ICT-PaaS platform enabling 5G network slicing
abstract
With a traditional static model of telecommunications service, it is difficult to deal with the uncertain factors and differentiation of massive mobile Internet services. Cloud computing, software‐defined networking (SDN), and network function virtualisation (NFV) drive telecom networks to a new round of network reconstruction. The IaaS platform can partially solve the ‘production tool’ problem of operators; however, the ‘production relationship’ problem persists. After infrastructure reconstruction is completed, due to the 5G service vision of networks on demand and sliced networks, a PaaS environment is introduced to implement a new‐generation of virtual network functions (VNFs). Based on the analysis of the necessity and feasibility of an Information and Communications Technology PaaS (ICT‐PaaS) Platform, following the technical trend of 5G, the ICT‐PaaS platform is proposed to construct a future network featuring elasticity, automation, flexibility, and openness. The adoption of lightweight VNF design ideas based on componentisation, containerisation, and microservice can better satisfy the requirements of flexible ‘network slicing’ than the traditional heavy monolith VNF‐based on VMs. The measured results show that the optimisation implementations can significantly improve network forwarding performance. From the evaluation result of the cost benefits of it is possible to highlight potentials of the platform introduction.
Qiang Wu 0018, Chunming Wu 0001
IET Commun.2
2019 Reliable Communication in Transmission Grids based on Nondisjoint Path Aggregation Using Software-Defined Networking
abstract
In electrical transmission grids, the redundant communication paths nondisjointly overlapping at links can be established between certain substations and the control center to guarantee reliable packet delivery under link failures. However, the generation of nondisjoint paths with multiplicative and concave constraints is with the NP-complete complexity and the failovers can lead to out-of-order packets. This paper presents an OpenFlow-based nondisjoint path aggregation mechanism to heuristically compute the constrained nondisjoint paths in a centralized fashion and reorganize the out-of-order packets at edge switches. The solution is evaluated through simulations of the IEEE 30-bus network scenario under 2% link failure rate and the result confirms its effectiveness: the packet delivery success rate is significantly improved in comparison with the current nondisjoint and disjoint path algorithms. TCP throughput is improved by 121.62% with the packet reordering. Also, the memory usage for the packet reordering buffer is reduced by 76.563% compared with the distributed cognitive packet network.
Qiang Yang 0004, Chunming Wu 0001
IEEE Trans. Ind. Informatics4
2018 Generating stable biometric keys for flexible cloud computing authentication using finger vein
Zhendong Wu, Longwei Tian, Ping Li 0018, Ting Wu 0001, Ming Jiang 0009, Chunming Wu 0001
Inf. Sci.6
2018 PAME: Evolutionary membrane computing for virtual network embedding
Chunyan Yu, Qi Lian, Dong Zhang 0010, Chunming Wu 0001
J. Parallel Distributed Comput.4
2018 Easy Path Programming: Elevate Abstraction Level for Network Functions
abstract
As datacenter networks become increasingly programmable with proliferating network functions, network programming languages have emerged to simplify the program development of the network functions. While network functions exhibit high level abstraction over operations on the traffic flow and the interconnections among the operations, the existing languages usually require programming with detailed knowledge about the packet processing patterns at the switches. Such a mismatch between the program abstraction and development details makes developing network functions a nontrivial task. To solve the problem, this paper introduces the easy path programming (EP2) framework. EP2 offers a high-level abstraction to simplify the program design process of the network functions. EP2 also provides a language that captures the common properties of network functions and uses predicates and primitives as basic language components. Specifically, predicates describe when to handle a flow with a global view of the flow dynamics; and primitives describe how to choose a path for a specific flow. Furthermore, EP2 has its own runtime system to support the language and the abstraction model, especially to hide the low level packet-processing behavior at the data plane from the programmers. Throughout this paper, cases are given to illustrate the EP2 abstraction model, language details and benefits. The expressiveness of EP2, the potential overhead of the runtime system and the efficiency of the network functions generated by EP2 are evaluated. The results show that EP2 can achieve comparable performance while reducing programming efforts.
Fei Chen 0009, Chunming Wu 0001, Xiaoyan Hong, Bin Wang 0062
IEEE/ACM Trans. Netw.2
2018 SDN-RDCD: A Real-Time and Reliable Method for Detecting Compromised SDN Devices
Haifeng Zhou, Chunming Wu 0001, Zhouhao Lu, Qiumei Cheng
IEEE/ACM Trans. Netw.2
2017 PBUF: Sharing Buffer to Mitigate Flooding Attacks
abstract
Software defined networking (SDN) is a promising network architecture, which decouples the control plane and data plane of a network. However, SDN opens some security challenges, such as man-in-the-middle attacks, spoofing attacks, flooding attacks and so on. In this paper, we focus on flooding attacks which consume the switch buffer and controller resource resulting in SDN framework resource overloaded. To prevent SDN framework from flooding attack, we present a defense approach called PBUF (Packet forwarding based on BUFfer sharing), which pools the idle switches to mitigate threat issues. This approach consists of buffer management and packet forwarding modules. The buffer management module gleans the statistics of incoming packets and then analyzes these statistics to estimate the buffer size by network calculus. Considering that a lot of table-miss packets will be generated and stored in buffer when the flooding attack is happening, the packet forwarding module is designed to forward these table-miss packets to idle switches to prevent the switch or controller to be overloaded. These table-miss packets will be buffered in idle switches and then sent to controller in a limited rate by generating packet_in messages. The simulation results show that PBUF is effective and only introduces a little overhead in SDN framework.
Chang-Ting Lin, Chunming Wu 0001, Yifei Tian, Zhenyu Wen, Shouling Ji
ICPADS2
2017 HSTS Measurement and an Enhanced Stripping Attack Against HTTPS
Xurong Li, Chunming Wu 0001, Shouling Ji, Qinchen Gu, Raheem A. Beyah
SecureComm2
2017 H _2 DoS: An Application-Layer DoS Attack Towards HTTP/2 Protocol
Xiang Ling 0001, Chunming Wu 0001, Shouling Ji
SecureComm2
2017 SDN-LIRU: A Lossless and Seamless Method for SDN Inter-Domain Route Updates
abstract
Maintaining service availability during an inter-domain route update is a challenge in both conventional networks and software-defined networks (SDNs). In the update process, asynchronous reconfigurations to border forwarding devices in different domains will incur transient anomalies with numerous packet losses and service disruptions. Based on current SDN inter-domain routing mechanisms, we in this paper propose a lossless and seamless method for SDN inter-domain route updates. This method is lightweight, and it has no requirement to add extra switch functionality or to extend SDN southbound protocols. The primary idea of this method is to achieve a lossless inter-domain route update by communications and collaborations among relevant domains. Motivated by this idea, we first identify three different domain categories for the update, i.e., domains only on the new inter-domain route, domains on both the old and new inter-domain routes, and domains only on the old inter-domain route. We further find that the transient anomalies are able to be avoided by reconfiguring the related border switches of the three categories of domains in order. Four update steps are then designed to keep the orderly update. Furthermore, we present the theoretical proof of the effectiveness of this method. Finally, based on our prototype implementation, the proposed method is also validated by simulation studies, and the simulation results indicate that this method succeeds in avoiding packet loss and maintaining service availability during the update.
Haifeng Zhou, Chunming Wu 0001, Qiumei Cheng, Qianjun Liu
IEEE/ACM Trans. Netw.2
2016 Engineering traffic uncertainty in the OpenFlow data plane
abstract
This paper is driven by a simple question of whether traffic engineering in Software Defined Networking (SDN) can react quickly to bursty and unpredictable changes in traffic demand. The key challenge is to strike a careful balance between the overhead (frequently involving the SDN controller) and performance (the degree of congestion measured as the maximum load and the balance between the minimum and the maximum loads). Exploiting OpenFlow (OF) features, quick shift of routing paths for unpredictable traffic bursty is the focal point of this work. It is achieved by using a dual routing scheme and letting the data plane to select the appropriate path in reacting to uncertainty in traffic load. The proposed work is called DUCE (Demand Uncertainty Configuration sElection). Further, we describe a traffic distribution model, an optimization solution that calculates congestion-free traffic distribution plan which guarantees that each switch can select one of the paths in a distributed way, and moreover, OF details about detaching the functionality of responding to the demand uncertainty from the control plane and delegating it to the data plane. Simulations are performed validating the efficiency of DUCE under various network scenarios.
Fei Chen 0009, Chunming Wu 0001, Xiaoyan Hong, Zhouhao Lu, Zhouhao Wang, Chang-Ting Lin
INFOCOM2
2016 A multipath resource updating approach for distributed controllers in software-defined network
Xiaochun Wu, Chunming Wu 0001, Chang-Ting Lin, Qiang Wu 0018, Bin Wang 0062
Sci. China Inf. Sci.2
2016 Traffic matrix estimation: A neural network approach with extended input and expectation maximization iteration
Haifeng Zhou, Liansheng Tan, Chunming Wu 0001
J. Netw. Comput. Appl.4
2015 Embedding Algorithm for Virtualizing Content-Centric Networks in a Shared Substrate
abstract
Network virtualization enables diversified network architectures to coexist in a substrate network. Content-centric network (CCN), as one of the major proposals for the future network, can be deployed in the virtualized environment. However, the recent work that has attempted in this direction has not addressed issues concerning the unique resource usage of CCN in resource allocations during the embedding procedure when instantiating a virtual CCN (VCCN). The unique problem is the cache component, i.e., the content store, in CCN. This paper will develop a VCCN Mapping Algorithm (VCCNMA) that optimizes the performances of VCCN considering the content store request in addition to CPU and bandwidth. Since the resource allocation problem is NP-hard, this paper thus proposes a heuristic algorithm. Several sets of simulation experiments are performed to evaluate the acceptance ratio, the network cost, the storage utilization ratio and load balance. Results show that the algorithm can achieve better performances. Moreover, the paper also studies the tradeoffs between the network cost and the load balance.
Shengquan Liao, Xiaoyan Hong, Chunming Wu 0001, Ming Jiang 0009
GLOBECOM3
2015 Improving QoS in SDN with lossless multi-domain reconfigurations
abstract
In this poster, we propose a novel approach of multidomain reconfigurations in SDN, termed as Lossless Reconfiguration (LR), to avoid packet loss and maintain the availability of services during the reconfiguration process. By leveraging the advantages of SDN, e.g., centralized control in one domain and feasible cooperation between the controllers of different domains, LR offers a better solution to the transient problem. First, we identify three categories of domains in the reconfiguration process. We then develop four steps to reconfigure the different categories of domains in order. By the synchronization of the involved controllers in each reconfiguration step, the transient problem can be resolved without any device modification requirement.
Haifeng Zhou, Chunming Wu 0001, Wen Gao 0001, Ming Jiang 0009, Tingting Pan
IWQoS2
2014 Programming network via Distributed Control in Software-Defined Networks
abstract
Programming a network for innovative services or for function improvements has never been easier using Software-Defined Networking (SDN). However, the programming tasks can also be significantly complicated by the asynchrony of data plane states and complexities of service control states. In order to reduce the complexity for programming a network service in Distributed Control Plane of SDN, we propose a proGRAmming Control (GRACE) layer as a generic solution, which provides two key features, namely, reconfigurability and reusability. Their implementations deal with aforementioned challenges, thus to achieve the consistency of the data plane states and the reusability of the service control states at the distributed controllers. This paper introduces the reconfigurability and reusability with their design goals and their impact on the programmability of DCP. We further use two popular network services, ICN (Information-Centric Networking) and CDN (Content Distribution Networks) to illustrate these concepts. NS-3 simulations and PlanetLab emulations are conducted to show the advantage of using the GRACE layer for ICN and CDN. Results show that the ICN Interest delay is reduced by 19.6% and CDN request delay is reduced by 81% in extremely harsh network conditions.
Chunming Wu 0001, Xiaoyan Hong, Ming Jiang 0009
ICC2
2014 Dynamic load distribution with hop-by-hop forwarding based on max-min one-way delay
Fei Chen 0009, Chunming Wu 0001, Bin Wang 0062, Yaguan Qian, Xiaochun Wu
Sci. China Inf. Sci.2
2014 A secure routing model based on distance vector routing algorithm
Bin Wang 0062, Chunming Wu 0001, Qiang Yang 0004, Pan Lai, Julong Lan
Sci. China Inf. Sci.2
2013 Multicast virtual network mapping for supporting multiple description coding-based video applications
Yuting Miao, Qiang Yang 0004, Chunming Wu 0001, Ming Jiang 0009, Jinzhou Chen
Comput. Networks3
2012 Robust dynamic bandwidth allocation method for virtual networks
abstract
Multiple virtual networks sharing an underlying substrate network is considered a promising tool to diversify and reshape the future inter-networking paradigm. As a simple and straightforward approach, the static bandwidth allocation in virtual networks (VNs) can often be inefficient in practice, and the adaptive allocation scheme with a small time-scale may lead to the transient network and service instability. Due to the fact that the traffic patterns vary over time, the challenge still remains to meet the expected resources allocation whilst promote the network scalability and robustness. In this paper, based on the robust optimization theory we present a robust dynamic approach which periodically identifies bandwidth allocation to VNs to work reasonable well for a range of traffic patterns over a period of time, rather than certain traffic pattern instance. This problem is formulated as a robust optimization problem using path-flow model aiming to compute the minimum-cost bandwidth allocation. Through the primal decomposition, we present a distributed algorithm which consists of two components running in individual VNs and the substrate network respectively. The numerical result obtained from simulation experiments demonstrates the strength and the effectiveness of the proposed algorithm in terms of convergence and acceptance ratio.
Min Zhang 0029, Chunming Wu 0001, Qiang Yang 0004, Ming Jiang 0009
ICC2
2010 Mapping Multicast Service-Oriented Virtual Networks with Delay and Delay Variation Constraints
abstract
As a key issue of building a virtual network (VN), the VN mapping problem can be addressed by various state-of-the-art algorithms. While these algorithms are efficient for the construction of unicast service-oriented VNs, they are generally not suitable for multicast cases. In this paper, we investigate the mapping problem in the context of virtual multicast service-oriented network subject to delay and delay variation constraints (VMNDDVC). We present a novel and efficient heuristic algorithm to tackle this problem based on a sliding window approach. The primary objective of this algorithm is in two-fold: to minimize the cost of VMNDDVC request mapping, and to achieve load balancing so as to increase the acceptance ratio of virtual multicast network (VMN) requests. The numerical results obtained from extensive simulation experiments demonstrate the effectiveness of the proposed approach and superiority than existing solutions in terms of VN mapping acceptance ratio, total revenue and cost in the long term.
Min Zhang 0029, Chunming Wu 0001, Ming Jiang 0009, Qiang Yang 0004
GLOBECOM2