EDBT 2026 Demo / reviewers in the wild / expert
Haifeng Zhou
dblp:81/6572
· DBLP profile ↗
44ranked-venue papers
6as first author
34since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 36 · 6 first-author · 26 since 2021Systems, architecture and hardware · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MonPlan: Taming Network Measurement with Accurate and Resource-Efficient Sketch-INT Co-Design
Xiang Chen 0017, Linying Zheng, Longlong Zhu, Zedi Chen, Qing Shu, Jialu Tian, Siqi Dong, Qun Huang 0001, Jianshan Zhang, Xuan Liu 0006, Haifeng Zhou, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001 |
INFOCOM | 11 |
| 2026 | MPulse: A Programmable and Autonomic Fault Detection System via Hierarchical Liveness Exchange
Di Wang 0003, Haifeng Zhou, Zhengyan Zhou, Jiayu Luo |
INFOCOM | 2 |
| 2026 | Holm: A DPU-Based Robust Host-Side Latency Monitoring with Low-Overhead and Selective Full-Coverage
Haifeng Zhou, Di Wang 0003, Wenbin Zhang 0011, Dianxing Tang, Zhengyan Zhou, Chunming Wu 0001 |
INFOCOM | 1 |
| 2026 | Achieving Precise Host Congestion Mitigation by DPU Offloading
Haifeng Zhou, Di Wang 0003, Dianxing Tang, Zhengyan Zhou, Chunming Wu 0001 |
INFOCOM | 1 |
| 2026 | SketchPipe: Toward Accurate Sketch-based Network Measurement on Multi-Pipeline Switches with Splitless Sketch Placement
Xiang Chen 0017, Longlong Zhu, Linying Zheng, Hongyang Du 0001, Dong Zhang 0010, Jianshan Zhang, Xuan Liu 0006, Qun Huang 0001, Dusit Niyato, Haifeng Zhou, Chunming Wu 0001, Hongyan Liu 0001, Kui Ren 0001 |
NSDI | 10 |
| 2025 | Phantom: Virtualizing Switch Register Resources for Accurate Sketch-based Network MeasurementabstractSketches have proven to be useful for measuring traffic. They store measurement results in the registers of data plane switches. However, they suffer from the short of switch register resources, limiting their measurement accuracy. Xiang Chen 0017, Hongyan Liu 0001, Zhengyan Zhou, Wenbin Zhang 0011, Hongyang Du 0001, Dong Zhang 0010, Xuan Liu 0006, Haifeng Zhou, Dusit Niyato, Qun Huang 0001, Chunming Wu 0001, Kui Ren 0001 |
EuroSys | 9 |
| 2025 | Carrera: Enabling High-Performance eBPF-based Sketches in Network MeasurementabstractTo achieve dynamic network measurement, trends build sketches on eBPF to avoid service interruptions. However, existing eBPF-based sketches suffer from high CPU consumption, leading to poor throughput and high latency and making them hard to measure high-speed traffic. Optimizing their performance requires users to refactor codes based on each sketch’s characteristics on eBPF, which is highly complex and time-consuming.In this paper, we argue that users should write sketches without concerning low-level eBPF performance optimizations, with the deployment automatically activating cross-sketch performance optimizations. We present Carrera, a library that offers domain-specific optimizations for eBPF-based sketches. Our contributions are (1) systematically analyzing the performance bottlenecks of eBPF-based sketches through microbenchmarks, (2) identifying practical optimizations, including hardware offloading, SIMD-accelerated hashing, traffic-aware flow index caching, prefetched randomization, and active data collection, to address the identified bottlenecks in eBPF-based sketches, (3) evaluating these optimizations with state-of-the-art sketches and demonstrating that Carrera improves throughput by up to 65% and reduces latency by up to 93% via testbed experiments. Xiang Chen 0017, Xin Yao 0008, Longlong Zhu, Linying Zheng, Hongyan Liu 0001, Jianshan Zhang, Dong Zhang 0010, Xuan Liu 0006, Qun Huang 0001, Haifeng Zhou, Chunming Wu 0001 |
ICNP | 11 |
| 2025 | SkewTide: Bridging Efficiency and Tail Latency in Key-Value Stores via Kernel Re-ArchitectureabstractKey-value stores are the key building block of online services such as e-commerce. However, highly skewed workloads (i.e., skewed access frequency and request size) may cause severe load imbalance and head-of-line blocking, resulting in significant performance penalty (e.g., low throughput and high latency). Existing works mitigate skewed workloads, but often struggle to balance CPU efficiency with low tail latency or require specialized hardware. In this paper, we present SkewTide, an in-kernel architecture that breaks this trade-off through workload-aware request pre-processing and bypassing unnecessary network stack operations. Moreover, SkewTide carefully orchestrates size-aware parsing, sharding, caching, and queueing in the kernel. Both designs enable efficient CPU multiplexing and preserve low tail latency without specialized hardware. We implement SkewTide as an out-of-the-box framework using eBPF, making it readily deployable in existing key-value store infrastructure. Evaluation with YCSB traces shows that SkewTide achieves up to 8.1× higher throughput, 37% lower 99th-percentile latency, and 32% lower CPU usage compared to existing systems. Jinghan Zu, Zhengyan Zhou, Lingfei Cheng, Zhongfeng Jin, Haifeng Zhou, Chunming Wu 0001 |
ICNP | 5 |
| 2025 | P4-IDet: A Programmable Switch-Based Framework for Real-Time and High-Accuracy Traffic Anomaly Detection in ICPSsabstractThe rise of Industry 4.0 exposes traditionally isolated Industrial Cyber-Physical Systems (ICPSs) to increasing network attacks, posing serious security threats and potential damage. Traffic anomaly detection is essential for identifying such attacks. Nevertheless, existing work faces a dilemma between high accuracy and real-time performance. In this paper, we resolve this dilemma through P4-IDet, a novel traffic anomaly detection framework based on programmable switches, achieving both high accuracy and real-time performance. P4-IDet first deploys a low-complexity detector in the data plane to stamp timestamps, extract traffic features, and perform line-rate preliminary detection. Only suspicious packets and their features are uploaded to a server for fine-grained analysis by a high-accuracy machine learning model. To further reduce the upload and accelerate detection, a Bayesian optimizer adaptively tunes detection rules based on differences between detection results of the switch and the server. Moreover, P4-IDet can be integrated with existing detection models to enhance accuracy and real-time performance. Finally, we implement the prototype on a Barefoot Tofino 2.0 switch using the P4 language and an x86 server, and validate it on a large-scale ICPS platform with real-world industrial systems. Experiments show 5.6–41.1% accuracy gains, a 28.93% reduction in machine learning model workload, and 8.31–25.90% improvements in real-time performance. Jiayu Luo, Zhengyan Zhou, Qiaoxiong Tang, Ruohan Chen, Xiang Chen 0017, Chao Pei, Qiang Yang 0004, Wenhai Wang, Haifeng Zhou |
IECON | 13 |
| 2025 | TurboCache: Empowering Switch-Accelerated Key-Value Caches with Accurate and Fast Cache UpdatesabstractRecent key-value (KV) caches are offloaded to programmable switches to offer high query processing performance. However, they suffer from both low accuracy in hot key detection and high latency in cache updates due to the strict limitations on switch registers. We propose TurboCache, a switch-accelerated KV cache with accurate hot key detection and fast cache updates. Our key idea is to leverage the switch recirculation capability to build a novel data structure that caches hot KV pairs. With this hardware-compatible cache data structure, TurboCache designs efficient data plane algorithms that accurately detects new hot keys and quickly updates its cache entirely within switch ASIC pipelines. We have implemented TurboCache on a${64}\times {100}$Gbps Tofino switch. Testbed results indicate that TurboCache improves the hot key detection accuracy and decreases the cache update latency of existing KV caches by several orders of magnitude. Xiang Chen 0017, Longlong Zhu, Linying Zheng, Lingfei Cheng, Jianshan Zhang, Xu Yang 0002, Dong Zhang 0010, Xuan Liu 0006, Xiaoming Lu, Xun Yi, Ibrahim Khalil 0001, Albert Y. Zomaya, Haifeng Zhou, Chunming Wu 0001 |
INFOCOM | 13 |
| 2025 | Efficient active flow control strategy for confined square cylinder wake using deep learning-based surrogate model and reinforcement learning
Mustafa Z. Yousif, Minze Xu, Haifeng Zhou, Linqi Yu, HeeChang Lim |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | SpotMon: Enabling General Hotspot Monitoring in Key-Value StoresabstractKey-value stores are essential to online services such as e-commerce. In key-value stores, a hotspot (i.e., frequently accessed items) may cause severe load imbalances, high response latency, and Service Level Agreement (SLA) violations. However, existing works only focus on specific types of hotspots, thus overlooking other types of hotspots and leading to blind spots. In this paper, we propose SpotMon, a system that enables general hotspot monitoring in key-value stores. Specifically, we (1) systematically identify the generality requirements of hotspot monitoring from existing works, (2) formulate general hotspot monitoring as an arbitrary partial spot query problem, (3) measure the hotness of hotspot candidates with a new vector expression, (4) propose hotspot encoding, filtering, decoding, and querying to support general queries without focusing on specific hotspots, (5) leverage the in-network visibility of programmable switches to identify system-wide hotspots. Our extensive experiments indicate that SpotMon provides high accuracy (e.g., F1 score from 0.88 to 1) and enables efficient hotspot mitigations (e.g., up to$4.03 \times$MQPS). Zhengyan Zhou, Jinhan Zu, Enhao Huang, Haifeng Zhou, Dong Zhang 0010, Xiang Chen 0017, Chunming Wu 0001 |
ICNP | 5 |
| 2024 | Accelerating Sketch-based End-Host Traffic Measurement with Automatic DPU OffloadingabstractSketch-based traffic measurement is a crucial building block for monitoring traffic statistics and ensuring the quality of services of end-host applications. However, existing approaches for building sketches in end-hosts exhibit poor packet processing performance or high CPU consumption. In this paper, we propose MPU, which automatically offloads sketch-based measurement to the emerging hardware, DPU. MPU consists of a sketch analyzer that profiles sketch resource consumption and an optimization framework that formulates the offloading problem and maximizes sketch performance on DPU. We implement MPU on the NVIDIA BlueField DPU. Our testbed results indicate that MPU achieves 85% lower per-packet processing latency and 47% higher traffic measurement accuracy when compared to existing approaches. Xiang Chen 0017, Wenbin Zhang 0011, Xin Yao 0008, Zizheng Wang, Hongyan Liu 0001, Qun Huang 0001, Xuan Liu 0006, Haifeng Zhou, Chunming Wu 0001 |
INFOCOM | 10 |
| 2024 | Eagle: Toward Scalable and Near-Optimal Network-Wide Sketch Deployment in Network MeasurementabstractSketches are useful for network measurement thanks to their low resource overheads and theoretically bounded accuracy. However, their network-wide deployment suffers from the trade-off between optimality and scalability: (1) Most solutions rely on mixed integer linear programming (MILP) solvers to provide the optimal decisions. But they are time-consuming and can hardly scale to large-scale deployment scenarios. (2) While heuristics achieve scalability, they deteriorate resource and performance overheads. We propose Eagle, a framework that achieves scalable and near-optimal network-wide sketch deployment. Our key idea is to decompose network-wide sketch deployment into sub-problems. Such decomposition allows Eagle to (1) simultaneously optimize switch resource consumption and end-to-end performance (retaining optimality), and (2) incorporate time-saving techniques into sub-problem solving (achieving scalability). Compared to existing solutions, Eagle improves scalability by up to 255× with negligible loss of optimality. It has also saved administrators in a production network days of efforts and reduced the operation time from O(hour) to O(second). Xiang Chen 0017, Qingjiang Xiao, Hongyan Liu 0001, Qun Huang 0001, Dong Zhang 0010, Xuan Liu 0006, Longbing Hu, Haifeng Zhou, Chunming Wu 0001, Kui Ren 0001 |
SIGCOMM | 8 |
| 2024 | Terra: Low-latency and reliable event collection in network measurement
Hongyan Liu 0001, Xiang Chen 0017, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006, Muhammad Khurram Khan |
J. Netw. Comput. Appl. | 6 |
| 2024 | Toward Scalable and Low-Cost Traffic Testing for Evaluating DDoS Defense SolutionsabstractTo date, security researchers evaluate their solutions of mitigating distributed denial-of-service (DDoS) attacks via kernel-based or kernel-bypassing testing tools. However, kernel-based tools exhibit poor scalability in attack traffic generation while kernel-bypassing tools incur unacceptable monetary cost. We propose Excalibur, a scalable and low-cost testing framework for evaluating DDoS defense solutions. The key idea is to leverage the emerging programmable switch to empower testing tasks with Tbps-level scalability and low cost. Specifically, Excalibur offers intent-based primitives to enable academic researchers to customize testing tasks on demand. Moreover, in view of switch resource limitations, Excalibur coordinates both a server and a programmable switch to jointly perform testing tasks. It realizes flexible attack traffic generation, which requires a large number of resources, in the server while using the switch to increase the sending rate of attack traffic to Tbps-level. We have implemented Excalibur on a$64\times 100$Gbps Tofino switch. Our experiments on a$64\times 100$Gbps Tofino switch show that Excalibur achieves orders-of-magnitude higher scalability and lower cost than existing tools. Xiang Chen 0017, Hongyan Liu 0001, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006 |
IEEE/ACM Trans. Netw. | 5 |
| 2024 | Hermes: Low-Overhead Inter-Switch Coordination in Network-Wide Data Plane Program DeploymentabstractNetwork administrators usually realize network functions in data plane programs. They employ the network-wide program deployment that decomposes input programs into match-action tables (MATs) while deploying each MAT on a specific switch. Since MATs may be deployed on different switches, existing solutions propose the inter-switch coordination that uses the per-packet header space to deliver crucial packet processing information among switches. However, such coordination incurs non-trivial per-packet byte overhead, leading to end-to-end performance degradation. We propose, a framework that aims to minimize the per-packet byte overhead. The key idea is to formulate network-wide program deployment as a mixed-integer programming (MIP) problem with the objective of minimizing the per-packet byte overhead. Also, offers a greedy-based heuristic that solves the problem in a near-optimal and timely manner. We have implemented on Tofino switches. Compared to existing frameworks, decreases the per-packet byte overhead by 156 bytes while preserving end-to-end performance in terms of flow completion time and goodput. Xiang Chen 0017, Hongyan Liu 0001, Qingjiang Xiao, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006, Qiang Yang 0004 |
IEEE/ACM Trans. Netw. | 6 |
| 2024 | Toward Full-Coverage and Low-Overhead Profiling of Network-Stack LatencyabstractIn modern data center networks (DCNs), network-stack processing denotes a large portion of the end-to-end latency of TCP flows. So profiling network-stack latency anomalies has been considered as a crucial part in DCN performance diagnosis and troubleshooting. In particular, such profiling requires full coverage (i.e., profiling every TCP packet) and low overhead (i.e., profiling should avoid high CPU consumption in end-hosts). However, existing solutions rely on system calls or tracepoints in end-hosts to implement network-stack latency profiling, leading to either low coverage or high overhead. We propose Torp, a framework that offers full-coverage and low-overhead profiling of network-stack latency. Our key idea is to offload as much of the profiling from costly system calls or tracepoints to the Torp agent built on eBPF modules, and further to include a Torp handler on the ToR switch to accelerate the remaining profiling operations. Torp efficiently coordinates the ToR switch and the Torp agent on end-hosts to jointly execute the entire latency profiling task. We have implemented Torp on$32\times 100$Gbps Tofino switches. Testbed experiments indicate that Torp achieves full coverage and orders of magnitude lower host-side overhead compared to other solutions. Xiang Chen 0017, Hongyan Liu 0001, Wenbin Zhang 0011, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Xuan Liu 0006, Chunming Wu 0001 |
IEEE/ACM Trans. Netw. | 6 |
| 2024 | Resource-Efficient and Timely Packet Header Vector (PHV) Encoding on Programmable SwitchesabstractThe programmable switch offers a limited capacity of packet header vector (PHV) words that store packet header fields and metadata fields defined by network functions. However, existing switch compilers employ inefficient strategies of encoding fields on PHV words. Their encoding wastes scarce PHV words and may result in failures when deploying network functions. In this paper, we propose Melody, a new framework that reuses PHV words for as many fields as possible to achieve resource-efficient PHV encoding. Melody offers a field analyzer and an optimization framework. The analyzer identifies which fields can reuse PHV words while preserving the original packet processing logic. The framework integrates analysis results into its encoding to offer the resource-optimal decisions. Also, to achieve timeliness at runtime, it provides a Greedy-based heuristic, which quickly solves PHV encoding and returns near-optimal results. We evaluate Melody with production-scale network functions. Our results show that Melody reduces the consumption of PHV words by up to 85%. Xiang Chen 0017, Wenbin Zhang 0011, Hongyan Liu 0001, Jianshan Zhang, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Xuan Liu 0006, Chunming Wu 0001 |
IEEE/ACM Trans. Netw. | 8 |
| 2023 | Halia: Toward Full-Coverage Network Function Offloading in the Data PlaneabstractOffloading network functions (NFs) to data plane switches brings remarkable performance benefits. In such offloading, NFs are required to process all the flows of interest (i.e., full coverage) to preserve the quality of services. However, existing solutions fail to guarantee full coverage for NFs. Thus, NFs may miss some essential flows, leading to accuracy drops. In this paper, we propose Halia, a framework that makes NF offloading decisions while ensuring full coverage for NFs. Specifically, Halia formulates the problem of NF offloading as an optimization problem. It encodes the requirement of full coverage as a constraint. Thus, its decisions activate enough NF instances in the substrate network to achieve full coverage.. We have implemented Halia and conducted experiments under multiple realistic network topologies to evaluate Halia. The experimental results indicate that compared to existing solutions, Halia achieves full coverage and high scalability in large-scale networks. Hongyan Liu 0001, Xiang Chen 0017, Qingjiang Xiao, Kaiwei Guo, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
ICC | 7 |
| 2023 | Excalibur: A Scalable and Low-Cost Traffic Testing Framework for Evaluating DDoS Defense SolutionsabstractTo date, security researchers evaluate their solutions of mitigating denial-of-service (DDoS) attacks via kernel-based or kernel-bypassing testing tools. However, kernel-based tools exhibit poor scalability in attack traffic generation while kernel-bypassing tools result in unacceptable monetary cost. We propose Excalibur, a scalable and low-cost testing framework for DDoS defense solutions. The key idea is to leverage the programmable switch to perform testing tasks with Tbps-level scalability and low cost. Specifically, Excalibur coordinates both a server and a programmable switch to jointly perform testing tasks. It realizes flexible attack traffic generation, which requires a large number of resources, in the server while using the switch to increase the sending rate of attack traffic to Tbps-level. Our experiments on a 64×100Gbps Tofino switch show that Excalibur achieves orders-of-magnitude higher scalability and lower cost than existing tools. Xiang Chen 0017, Hongyan Liu 0001, Tingxin Sun, Qun Huang 0001, Dong Zhang 0010, Xuan Liu 0006, Haifeng Zhou, Chunming Wu 0001 |
INFOCOM | 8 |
| 2023 | RFT: Toward Highly Reliable Flow Data Transmission in Network MeasurementabstractHow to satisfy the latency and reliability requirements of flow data transfer is an essential problem. To address this problem, we propose RFT, a framework that aims to satisfy the user-specified latency and reliability requirements of flow data transfer, especially in the situation where the network resources are insufficient. Firstly, we formulate the problem of satisfying the user-specified latency and reliability requirements of data transfer via mixed integer linear programming (MILP), and a heuristic algorithm is then designed to solve it in a polynomialtime. Secondly, to satisfy these requirements under insufficient network resources, we proposed a greedy-based algorithm used to select the minimum number of links added to the network, which can be deployed with low cost, especially in production networks such as data centers. Finally, we have implemented RFT on a 64$\times$100 Gbps Intel Barefoot Tofino switch. Our experimental results indicate that RFT satisfies the user-specified latency and reliability requirements in all test cases at acceptable costs, even when the network resources are insufficient. Xiang Chen 0017, Di Wang 0003, Zhengyan Zhou, Wenhai Wang, Chunming Wu 0001, Haifeng Zhou |
SECON | 8 |
| 2023 | Fuzzy encoding and decoding approaches for 2-TCLE and their applications in multi-criteria decision making
Yaya Liu, Haifeng Zhou, Rosa M. Rodríguez 0001, Luis Martínez-López 0001 |
Inf. Sci. | 2 |
| 2023 | Automatic Performance-Optimal Offloading of Network Functions on Programmable SwitchesabstractIn network function virtualization (NFV), network functions (NFs) are chained as a service function chain (SFC) to enhance NF management with low cost and high flexibility. Recent NFV solutions indicate that the packet processing performance of SFCs can be significantly improved by offloading NFs to programmable switches. However, such offloading requires a deep understanding of heterogeneous NF properties (e.g., NF resource consumption and NF performance behaviors) to achieve the maximum SFC performance. Unfortunately, none of existing solutions provide automatic analysis of these NF properties. Thus, network administrators have to manually examine the source codes of NFs and profile various NF properties by hand, which is extremely time-consuming and laborious. In this article, we propose LightNF, a novel system that simplifies NF offloading in programmable networks. LightNF automatically dissects comprehensive NF properties by means of code analysis and performance profiling while eliminating manual efforts. It then leverages its analysis results of NF properties in its SFC placement so as to make the performance-optimal offloading decisions. We have implemented LightNF on Tofino-based hardware programmable switches. We perform extensive experiments to evaluate LightNF with a real-world testbed and large-scale simulation. Our experiments show that LightNF outperforms existing solutions with an orders-of-magnitude reduction in per-packet processing latency and 9.5× improvement in SFC throughput. Xiang Chen 0017, Hongyan Liu 0001, Dong Zhang 0010, Zili Meng, Qun Huang 0001, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006, Qiang Yang 0004 |
IEEE Trans. Cloud Comput. | 6 |
| 2023 | Stalker Attacks: Imperceptibly Dropping Sketch Measurement Accuracy on Programmable SwitchesabstractDue to limited memory usage and provably high accuracy, sketches running on programmable switches have been commonly used by the literature for network measurement. However, their vulnerabilities are still largely unknown and neglected, which is highly concerning given the increasing popularity of network measurement. In this paper, we identify the Stalker attacks, where attackers aim to degrade the accuracy of sketches running on programmable switches. More precisely, attackers tamper with some sketch operations during sketch deployment atop programmable switches. At runtime, the tampered sketch will record highly inaccurate flow data, which degrades measurement accuracy. We implement Stalker attacks on Tofino switches. The results indicate that Stalker attacks significantly drop the accuracy of network management applications, e.g., reducing the F1 score of heavy hitter detection to zero. However, our analysis indicates that none of existing methods can detect Stalker attacks since they can hardly verify the correctness of sketch operations. Finally, we analyze potential defense mechanisms and identify challenges to enable further research in this context. Xiang Chen 0017, Hongyan Liu 0001, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006, Muhammad Khurram Khan |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Toward Low-Latency and Accurate State Synchronization for Programmable NetworksabstractProgrammable switches empower stateful packet processing, in which incoming packets continuously update states in the data plane, while applications in the control plane read and write states. However, since the data plane and control plane are separated, a consistent view of states in both planes is required for stateful packet processing. Existing approaches suffer from either high latency or low accuracy. In this paper, we propose ApproSync, a framework that offers approximate state synchronization with low latency and high accuracy. To achieve low latency, ApproSync directly transfers states between switch ASICs and the control plane by bypassing switch operating systems. To achieve high accuracy, ApproSync utilizes the resources in the switch ASIC to realize rate control in state synchronization, such that it avoids potential state loss. It also bounds the divergence between the states in the data plane and that in the control plane under limited link capacity. We prototype ApproSync on Barefoot Tofino switches. The experimental results indicate that compared to existing approaches, ApproSync achieves order-of-magnitude latency reduction while maintaining high accuracy of state synchronization. Also, our experiments demonstrate that ApproSync provides significant latency benefits to existing network management applications and well preserves high application-level accuracy. Xiang Chen 0017, Hongyan Liu 0001, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006, Qiang Yang 0004 |
IEEE/ACM Trans. Netw. | 5 |
| 2023 | Eliminating Control Plane Overload via Measurement Task PlacementabstractRecent efforts in network measurement place measurement tasks on programmable switches to measure high-speed traffic. These tasks extract flow data, i.e., events, from packets and send events to the control plane. However, the tasks may generate massive events in a short time. In this context, the links transferring events to the control plane and the control plane servers that handle events may be overloaded, i.e., control plane overload. None of existing solutions can eliminate control plane overload. In this paper, we propose MTP, a framework that eliminates control plane overload via careful measurement task placement. Our key idea is to allocate enough resources for each task during task placement to avoid control plane overload at runtime. For each task, MTP estimates its maximum possible rate of sending events to the control plane. Then its optimization framework addresses the resource restrictions of both switches and the control plane. The experiments on Tofino switches indicate that MTP outperforms existing solutions with higher accuracy in several use cases. Xiang Chen 0017, Hongyan Liu 0001, Dong Zhang 0010, Qun Huang 0001, Haifeng Zhou, Chunming Wu 0001, Qiang Yang 0004 |
IEEE/ACM Trans. Netw. | 5 |
| 2022 | Toward Low-Overhead Inter-Switch Coordination in Network-Wide Data Plane Program DeploymentabstractIn modern networks, administrators realize their desired functions such as network measurement in several data plane programs. They often employ the network-wide program deployment paradigm that decomposes input programs into match-action tables (MATs) while deploying each MAT on a specific programmable switch. Since MATs may be deployed on different switches, existing solutions propose the inter-switch coordination that uses the per-packet header space to deliver crucial packet processing information among switches. However, such coordination introduces non-trivial per-packet byte overhead, leading to significant end-to-end network performance degradation. In this paper, we propose Hermes, a program deployment framework that aims to minimize the per-packet byte overhead. The key idea of Hermes is to formulate the network-wide program deployment as a mixed-integer linear programming (MILP) problem with the objective of minimizing the per-packet byte overhead. In view of the NP hardness of the MILP problem, Hermes further offers a greedy-based heuristic that solves the problem in a near-optimal and timely manner. We have implemented Hermes on Tofino-based switches. Our experiments show that compared to existing frameworks, Hermes decreases the per-packet byte overhead by 156 bytes while preserving end-to-end performance in terms of flow completion time and goodput. Xiang Chen 0017, Hongyan Liu 0001, Qingjiang Xiao, Kaiwei Guo, Tingxin Sun, Xiang Ling 0001, Xuan Liu 0006, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
ICDCS | 10 |
| 2022 | Torp: Full-Coverage and Low-Overhead Profiling of Host-Side LatencyabstractIn data center networks (DCNs), host-side packet processing accounts for a large portion of the end-to-end latency of TCP flows. Thus, the profiling of host-side latency anomalies has been considered as a crucial part in DCN performance diagnosis and troubleshooting. In particular, such profiling requires full coverage (i.e., profiling every TCP packet handled by end-hosts) and low overhead (i.e., profiling should avoid high CPU consumption in end-hosts). However, existing solutions fully rely on end-hosts to implement host-side latency profiling, leading to low coverage or high overhead. In this paper, we propose Torp, a framework that offers full-coverage and low-overhead profiling of host-side latency. Our key idea is to offload profiling operations to top-of-rack (ToR) switches, which inherently offer full coverage and line-rate packet processing performance. Specifically, Torp selectively offloads profiling operations to the ToR switch based on switch limitations. It efficiently coordinates the ToR switch and end-hosts to execute the entire latency profiling task. We have implemented Torp on 32×100Gbps Tofino switches. Testbed experiments indicate that Torp achieves full coverage and orders of magnitude lower host-side overhead compared to other solutions. Xiang Chen 0017, Hongyan Liu 0001, Junyi Guo, Qun Huang 0001, Dong Zhang 0010, Chunming Wu 0001, Haifeng Zhou |
INFOCOM | 8 |
| 2022 | Escala: Timely Elastic Scaling of Control Channels in Network MeasurementabstractIn network measurement, data plane switches measure traffic and report events (e.g., heavy hitters) to the control plane via control channels. The control plane makes decisions to process events. However, current network measurement suffers from two problems. First, when traffic bursts occur, massive events are reported in a short time so that the control channels may be overloaded due to limited bandwidth capacity. Second, only a few events are reported in normal cases, making control channels underloaded and wasting network resources. In this paper, we propose Escala to provide the elastic scaling of control channels at runtime. The key idea is to dynamically migrate event streams among control channels to regulate the loads of these channels. Escala offers two components, including an Escala monitor that detects scaling situations based on realtime network statistics, and an optimization framework that makes scaling decisions to eliminate overload and underload situations. We have implemented a prototype of Escala on Tofino-based switches. Extensive experiments show that Escala achieves timely elastic scaling while preserving high application-level accuracy. Hongyan Liu 0001, Xiang Chen 0017, Qun Huang 0001, Dezhang Kong, Jinbo Sun, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
INFOCOM | 7 |
| 2022 | Efficient middlebox scaling for virtualized intrusion prevention systems in software-defined networks
Junchi Xing, Chunming Wu 0001, Haifeng Zhou, Qiumei Cheng, Danrui Yu, Mayra Alexandra Macas Carrasco |
Sci. China Inf. Sci. | 3 |
| 2021 | MTP: Avoiding Control Plane Overload with Measurement Task PlacementabstractIn programmable networks, measurement tasks are placed on programmable switches to keep pace with high-speed traffic. At runtime, programmable switches send events to the control plane for further processing. However, existing solutions for task placement overlook the limitations of control plane resources. Thus, excessive events may overload the control plane. In this paper, we propose MTP, a system that eliminates control plane overload via careful task placement. For each task, MTP analyzes its structure to estimate its maximum possible rate of sending events to the control plane. Then it builds an optimization framework that addresses the resource restrictions of both switches and the control plane. We have implemented MTP on Barefoot Tofino switches. The experimental results indicate that MTP outperforms existing solutions with higher accuracy across four real use cases. Xiang Chen 0017, Qun Huang 0001, Peiqiao Wang, Hongyan Liu 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
INFOCOM | 7 |
| 2021 | LightNF: Simplifying Network Function Offloading in Programmable NetworksabstractIn network function virtualization (NFV), network functions (NFs) are chained as a service function chain (SFC) to enhance NF management with high flexibility. Recent solutions indicate that the processing performance of SFCs can be significantly improved by offloading NFs to programmable switches. However, such offloading requires a deep understanding of NF properties to achieve the maximum SFC performance, which brings non-trivial burdens to network administrators. In this paper, we propose LightNF, a novel system that simplifies NF offloading in programmable networks. LightNF automatically dissects comprehensive NF properties (e.g., NF performance behaviors) via code analysis and performance profiling while eliminating manual efforts. It then leverages the analyzed NF properties in its SFC placement so as to produce the performance-optimal offloading. We have implemented a LightNF prototype. Our experiments show that LightNF outperforms state-of-the-art solutions with an orders-of-magnitude reduction in per-packet processing latency and 9.5× improvement in SFC throughput. Xiang Chen 0017, Qun Huang 0001, Peiqiao Wang, Zili Meng, Hongyan Liu 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
IWQoS | 8 |
| 2021 | Machine learning based malicious payload identification in software-defined networkingabstractDeep packet inspection (DPI) has been extensively investigated in software-defined networking (SDN) as complicated attacks may intractably inject malicious payloads in the packets. Existing proprietary pattern-based or port-based third-party DPI tools can suffer from limitations in efficiently processing a large volume of data traffic. In this paper, a novel OpenFlow-enabled deep packet inspection (OFDPI) approach is proposed based on the SDN paradigm to provide adaptive and efficient packet inspection. First, OFDPI prescribes an early detection at the flow-level granularity by checking the IP addresses of each new flow via OpenFlow protocols. Then, OFDPI allows for deep packet inspection at the packet-level granularity: (i) for unencrypted packets, OFDPI extracts the features of accessible payloads, including tri-gram frequency based on Term Frequency and Inverted Document Frequency (TF–IDF) and linguistic features. These features are concatenated into a sparse matrix representation and are then applied to train a binary classifier with logistic regression rather than matching with specific pattern combinations. In order to balance the detection accuracy and performance bottleneck of the SDN controller, OFDPI introduces an adaptive packet sampling window based on the linear prediction; and (ii) for encrypted packets, OFDPI extracts notable features of packets and then trains a binary classifier with a decision tree, instead of decrypting the encrypted traffic to weaken user privacy. A prototype of OFDPI is implemented on the Ryu SDN controller and the Mininet platform. The performance and the overhead of the proposed solution are assessed using the real-world datasets through experiments. The numerical results indicate that OFDPI can provide a significant improvement in detection accuracy with acceptable overheads. Qiumei Cheng, Chunming Wu 0001, Haifeng Zhou, Dezhang Kong, Dong Zhang 0010, Junchi Xing |
J. Netw. Comput. Appl. | 3 |
| 2020 | SRA: Switch Resource Aggregation for Application Offloading in Programmable NetworksabstractProgrammable switches empower network applications with line-rate packet processing performance by allowing the offloading of applications. However, the resource of a programmable switch is extremely limited, which significantly limits the application offloading. Existing solutions to the problem either provide poor efficiency or suffer from accuracy drop. In this paper, we propose SRA, a system that loosens switch resource constraints for application offloading via resource aggregation. SRA provides administrators with an intuitive compiler directive to customize application offloading by resource aggregation. According to compiler directives, it automatically places the program on the substrate network, while maintaining original packet processing logics. We implement a prototype of SRA in P4, and establish an experimental testbed consisting of three 32$\times$100 Gbps Barefoot switches. The experimental results indicate that SRA enhances two real-world applications with sufficient resources while maintaining high performance. Hongyan Liu 0001, Xiang Chen 0017, Qun Huang 0001, Haifeng Zhou, Dong Zhang 0010, Chunming Wu 0001 |
GLOBECOM | 4 |
| 2020 | ApproSync: Approximate State Synchronization for Programmable NetworksabstractProgrammable switches empower stateful packet processing, in which incoming packets continuously update states in the data plane, while applications in the control plane read and write states. However, as the data plane and control plane are separated, a consistent view of states in both planes is required for stateful packet processing. Existing approaches suffer from either high latency or low accuracy. In this paper, we propose ApproSync, a framework that offers approximate state synchronization with low latency and high accuracy. To achieve low latency, ApproSync directly transfers states between switch ASICs and the control plane by bypassing switch operating systems. To achieve high accuracy, ApproSync utilizes the resources in the switch ASIC to realize rate control in state synchronization, such that it avoids potential state loss. It also bounds the divergence between the states in the data plane and that in the control plane under limited link capacity. We prototype ApproSync on Barefoot Tofino switches. The experimental results indicate that compared to existing approaches, ApproSync achieves order-of-magnitude latency reduction while maintaining high accuracy. Xiang Chen 0017, Qun Huang 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
ICNP | 4 |
| 2020 | SPEED: Resource-Efficient and High-Performance Deployment for Data Plane ProgramsabstractProgrammable switches allow network administrators to customize packet processing behaviors in data plane programs. However, existing solutions for program deployment fail to achieve resource efficiency and high packet processing performance. In this paper, we propose SPEED, a system that provides resource-efficient and high-performance deployment for data plane programs. For resource efficiency, SPEED merges input data plane programs by reducing program redundancy. Then it abstracts the substrate network into an one big switch (OBS), and deploys the merged program on the OBS while minimizing resource usage. For high performance, SPEED searches for the performance-optimal mapping between the OBS and the substrate network with respect to network-wide constraints. It also maintains program logics among different switches via inter-device packet scheduling. We have implemented SPEED on a Barefoot Tofino switch. The evaluation indicates that SPEED achieves resource-efficient and high-performance deployment for real data plane programs. Xiang Chen 0017, Hongyan Liu 0001, Qun Huang 0001, Peiqiao Wang, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
ICNP | 6 |
| 2019 | P4SC: Towards High-Performance Service Function Chain Implementation on the P4-Capable Device
Xiang Chen 0017, Dong Zhang 0010, Haifeng Zhou |
IM | 5 |
| 2019 | Hiding and Trapping: A Deceptive Approach for Defending against Network Reconnaissance with Software-Defined NetworkabstractNetwork reconnaissance aims at gathering as much information as possible before an attack is launched. Meanwhile, static host address configuration facilitates network reconnaissance. Currently, more sophisticated network reconnaissance has been emerged with the adaptive and cooperative features. To address this, in this paper, we present Hiding and Trapping (HaT), which is a deceptive approach to disrupt adversarial network reconnaissance with the help of the software-defined networking (SDN) paradigm. HaT is able to hide valuable hosts from attackers and to trap them into decoy nodes through strategic and holistic host address mutation according to characteristic of adversaries. We implement a prototype of HaT, and evaluate its performance by experiments. The experimental results show that HaT is capable to effectively disrupt adversarial network reconnaissance with better deceptive performance than the existing address randomization approach. Junchi Xing, Haifeng Zhou, Chunming Wu 0001 |
IPCCC | 3 |
| 2018 | MATReduce: Towards High-Performance P4 Pipeline by Reducing Duplicate Match OperationsabstractP4 provides operators with the ability to program the packet processing pipeline of the data plane device. The match-action table (MAT) is a basic component of the P4 pipeline that matches the packet and performs an action on the matched packet. However, different MATs may execute duplicate match operations that decreases the performance of the P4 pipeline. To this end, we present MATReduce, a framework that optimizes the P4 pipeline by reducing duplicate match operations between MATs. MATReduce is composed of two key components, the preprocessor and the runtime manager. By introducing the compound MAT and rewriting the P4 control flow, the preprocessor merges duplicate match operations of the P4 pipeline while maintaining the program semantics. At runtime, the runtime manager converts user rules to actual rules for maintaining the policy consistency. Our preliminary experimental results show that MATReduce provides significant performance improvement, including a 23.90% throughput increase and a 34.37% delay decrease on the software target, and a 45.19% delay decrease on the hardware target. Xiang Chen 0017, Dong Zhang 0010, Haifeng Zhou |
GLOBECOM | 3 |
| 2018 | SDN-RDCD: A Real-Time and Reliable Method for Detecting Compromised SDN Devices
Haifeng Zhou, Chunming Wu 0001, Zhouhao Lu, Qiumei Cheng |
IEEE/ACM Trans. Netw. | 1 |
| 2017 | SDN-LIRU: A Lossless and Seamless Method for SDN Inter-Domain Route UpdatesabstractMaintaining service availability during an inter-domain route update is a challenge in both conventional networks and software-defined networks (SDNs). In the update process, asynchronous reconfigurations to border forwarding devices in different domains will incur transient anomalies with numerous packet losses and service disruptions. Based on current SDN inter-domain routing mechanisms, we in this paper propose a lossless and seamless method for SDN inter-domain route updates. This method is lightweight, and it has no requirement to add extra switch functionality or to extend SDN southbound protocols. The primary idea of this method is to achieve a lossless inter-domain route update by communications and collaborations among relevant domains. Motivated by this idea, we first identify three different domain categories for the update, i.e., domains only on the new inter-domain route, domains on both the old and new inter-domain routes, and domains only on the old inter-domain route. We further find that the transient anomalies are able to be avoided by reconfiguring the related border switches of the three categories of domains in order. Four update steps are then designed to keep the orderly update. Furthermore, we present the theoretical proof of the effectiveness of this method. Finally, based on our prototype implementation, the proposed method is also validated by simulation studies, and the simulation results indicate that this method succeeds in avoiding packet loss and maintaining service availability during the update. Haifeng Zhou, Chunming Wu 0001, Qiumei Cheng, Qianjun Liu |
IEEE/ACM Trans. Netw. | 1 |
| 2016 | Traffic matrix estimation: A neural network approach with extended input and expectation maximization iteration
Haifeng Zhou, Liansheng Tan, Chunming Wu 0001 |
J. Netw. Comput. Appl. | 1 |
| 2015 | Improving QoS in SDN with lossless multi-domain reconfigurationsabstractIn this poster, we propose a novel approach of multidomain reconfigurations in SDN, termed as Lossless Reconfiguration (LR), to avoid packet loss and maintain the availability of services during the reconfiguration process. By leveraging the advantages of SDN, e.g., centralized control in one domain and feasible cooperation between the controllers of different domains, LR offers a better solution to the transient problem. First, we identify three categories of domains in the reconfiguration process. We then develop four steps to reconfigure the different categories of domains in order. By the synchronization of the involved controllers in each reconfiguration step, the transient problem can be resolved without any device modification requirement. Haifeng Zhou, Chunming Wu 0001, Wen Gao 0001, Ming Jiang 0009, Tingting Pan |
IWQoS | 1 |