EDBT 2026 Demo / reviewers in the wild / expert
Menghao Zhang 0001
dblp:188/5710-1
· DBLP profile ↗
36ranked-venue papers
8as first author
31since 2021 · last 2026
0000-0001-5274-5512ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 22 · 3 first-author · 19 since 2021Security and privacy · 8 · 4 first-author · 6 since 2021Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HyLink: Harnessing PCIe and Dedicated Interconnects for Efficient Collective Communication
Yuezheng Liu, Menghao Zhang 0001, Xuebin Song, Juner Shen, Chunming Hu, Xudong Liu 0001 |
APNet | 2 |
| 2026 | Cross-Data Center Training with Heterogeneous Accelerators: Protocols and Evaluation
Bohua Xu, Xiongyan Tang, Dongyue Zhang, Xiaohe Hu, Lexi Xu, Xiaoxiang Wang, Menghao Zhang 0001 |
ICC | 9 |
| 2026 | MEGATRACE: Troubleshooting Hang and Slowdown in Large-Scale LLM Training Clusters
Fangzheng Jiao, Menghao Zhang 0001, Jiaxun Huang, Yanmin Jia, Xiaohe Hu, Bohua Xu, Chunming Hu |
ICDCS | 2 |
| 2026 | Vedrfolnir: RDMA Network Performance Anomalies Diagnosis in Collective CommunicationsabstractCollective communication becomes increasingly crucial as large language models rapidly evolve, but the RDMA it uses inevitably faces network performance anomalies (NPAs). Vedrfolnir is an accurate and efficient diagnosis system for RDMA NPAs in collective communication, which (1) constructs waiting graphs through algorithm decomposition, (2) adaptively detects anomalies while efficiently collecting diagnostic data, and (3) precisely analyzes performance bottlenecks and root causes. Evaluation shows that Vedrfolnir can achieve accurate diagnosis results with low overhead. Menghao Zhang 0001, Xiheng Li, Fangzheng Jiao, Xiao Li 0044, Jiaxun Huang, Chunming Hu |
INFOCOM | 2 |
| 2026 | Janus: Enabling Expressive and Efficient ACLs in High-speed RDMA Clouds
Ziteng Chen, Menghao Zhang 0001, Jiahao Cao 0001, Xuzheng Chen, Qiyang Peng |
NDSS | 2 |
| 2026 | Weaver: Diagnosing Extra Kernel and Synchronization Interference in GPU WorkloadsabstractModern GPU workloads use batching, asynchronous execution, kernel overlap, and GPU sharing to improve utilization, but these optimizations make kernel-level slowdown hard to diagnose. A target kernel may be blocked by extra synchronization or helper kernels, or slowed by concurrent kernels after it starts. Existing tools provide GPU visibility, but they either remain too heavy for continuous online use or focus on coarse-grained symptoms, leaving fine-grained kernel-level root causes to manual analysis. This poster presents Weaver, a low-overhead cross-layer diagnosis framework for GPU kernels. Weaver builds a semantic execution graph from operator, kernel timeline, and warp/block evidence, distinguishes blocked and slowed kernels, and reports an interpretable root-cause chain. Our prototype evaluation shows that Weaver continuously collects runtime evidence with low overhead and accurately localizes anomalies. Menghao Zhang 0001 |
SIGCOMM | 2 |
| 2025 | Mnemosyne: Lightweight and Fast Error Recovery for LLM Training in a Just-In-Time Manner
Jinyi Xia, Menghao Zhang 0001, Jiaxun Huang, Yuezheng Liu, Xiaohe Hu, Xudong Liu 0001, Chunming Hu |
APNet | 2 |
| 2025 | Cauchy: A Cost-Efficient LLM Serving System through Adaptive Heterogeneous DeploymentabstractRecent advances in large language models (LLMs) have intensified the need for serving LLMs that are cost-efficient and QoS-guaranteed. Existing frameworks often co-locate computationally distinct prefill and decode instances on homogeneous GPUs, overlooking their unique resource demands and under-utilizing heterogeneous GPUs. This leads to suboptimal resource utilization and increased capital expenditure. We present Cauchy, a LLM serving framework that adaptively deploys prefill and decode computation to the most suitable heterogeneous GPUs and dynamically schedules user requests. At the core of Cauchy is choosing proper GPU Combo, a conceptual GPU combination encompassing diverse GPU configurations, for their cost efficiency in running prefill-decode pairs. Cauchy deploys a set of combos to satisfy QoS requirements (e.g., goodput) of LLM inference. Cauchy further employs hierarchical scheduling to handle user requests, using opportunistic scheduling within the allocated GPU Combos and a goodput-weighted round-robin policy across GPU Combos. Dynamic autoscaling is used to stabilize the cost-efficiency in the face of surging requests. Experiments show that Cauchy achieves up to a 38.3% improvement in Tokens/USD efficiency over the state-of-the-art baselines, while maintaining strict Service Level Objectives (SLOs). Our work highlights the importance of leveraging workload and GPU heterogeneity to achieve superior cost-efficient LLM serving. Renyu Yang, Yuxi Luo, Menghao Zhang 0001, Li Li 0029, Chunming Hu, Tianyu Wo, Chengru Song, Jin Ouyang |
SoCC | 6 |
| 2025 | SuperFE: A Scalable and Flexible Feature Extractor for ML-based Traffic Analysis ApplicationsabstractThe feature extractor component in today's ML-based traffic analysis applications is becoming a key bottleneck. While mainstream software-based approaches can support flexible feature extraction, they fail to scale to multi-100Gbps network speed easily. Meanwhile, hardware-accelerated solutions can scale to high throughput, but cannot flexibly support generic traffic analysis applications. In this paper, we propose SuperFE, a feature extraction framework that allows users to extract traffic features efficiently and flexibly. SuperFE leverages the capabilities of both new-generation programmable switches and SmartNICs, with three key designs. First, SuperFE presents a user-friendly and extensible interface to support customized feature extraction policies, shielding underlying hardware implementation details and complexities. Second, SuperFE introduces a high-performance multi-granularity key-vector cache system in the programmable switches to batch necessary feature metadata for massive amounts of packets. Third, SuperFE exploits the multi-core parallel and hierarchical memory of SoC-based SmartNICs to achieve efficient feature computation with diverse streaming algorithms. Evaluations using our prototype demonstrate that SuperFE enables various state-of-the-art traffic analysis applications to efficiently extract features from multi-100Gbps raw traffic without compromising detection accuracy, and achieves nearly two orders of magnitude higher throughput than the software-based counterparts. Menghao Zhang 0001, Cheng Guo 0007, Renyu Yang, Han Bao 0011, Xiao Li 0044, Mingwei Xu 0001, Tianyu Wo, Chunming Hu |
EuroSys | 1 |
| 2025 | JEEVES: The Valet Who Masters the Art of Cross-DC Training SchedulingabstractAs model sizes continue to grow and the capacity of a single data center becomes insufficient, training models across multiple data centers efficiently is becoming increasingly important. In this paper, we first show that existing parallelism strategies perform poorly under limited bandwidth and high latency of cross-DC links. To address this, we propose JEEVES, a framework that extends the pipeline parallelism across DCs to minimize iteration time under memory constraints. We identify that the key lies in a good schedule of computation and communication, and propose communication-aware schedule, memory-aware stage division and inter-replica coordinated schedule. Simulations show that JEEVES improves iteration time by up to 43% when training a 175B-parameter model. Xuebin Song, Menghao Zhang 0001, Yuan Yang 0001, Mingwei Xu 0001 |
HotNets | 3 |
| 2025 | THEMIS: Addressing Congestion-Induced Unfairness in Long-Haul RDMA NetworksabstractRDMA is promising for enhancing the performance of cross-datacenter (DC) services. However, deploying RDMA over wide-area networks introduces severe congestion control unfairness, primarily due to asymmetric congestion feedback delays between inter-DC flows and intra-DC flows. As a result, intra-DC flows often bear the full burden of congestion response, leading to drastically increased flow completion times (FCT). In this work, we identify two key forms of unfairness — near-source and near-destination — depending on whether congestion occurs near the sender or receiver of inter-DC flows. Based on this, we propose THEMIS, a fairness maintenance patch for long-haul RDMA networks. To mitigate near-source unfairness, THEMIS devises a Proactive Notification Point to shorten the congestion feedback loop within a single DC. To alleviate near-destination unfairness, THEMIS introduces a Temporary Reaction Point to temporarily slow down the target inter-DC flow until the sender receives the corresponding congestion feedback. We implement an open-source prototype of THEMIS, and evaluate it on both real-world testbed and large-scale simulations. Compared to DCQCN, Annulus and BiCC, THEMIS reduces the intra-DC FCT by up to 79.2%, 63.6% and 55.6%, and decreases overall FCT by up to 61.2%, 31.9% and 59.5% respectively. Zihan Niu, Menghao Zhang 0001, Renjie Xie, Yuan Yang 0001, Xiaohe Hu |
ICNP | 2 |
| 2025 | Hawkeye: Diagnosing RDMA Network Performance Anomalies with PFC ProvenanceabstractRDMA is becoming increasingly prevalent from private data centers to public multi-tenant clouds, due to its remarkable performance improvement. However, its lossless traffic control, i.e., PFC, introduces new complexities in network performance anomalies (NPAs) due to its cascading congestion spreading property, which usually incurs complaints from customers/applications about certain flows' performance degradation. Existing studies fall short in fine-grained visibility of PFC impact and traceability of PFC causality, and are thus ineffective in diagnosing the root causes for RDMA NPAs. In this paper, we propose Hawkeye, an accurate and efficient RDMA NPA diagnosis system based on PFC provenance. Hawkeye comprises 1) a fine-grained PFC-aware telemetry mechanism to record the PFC impact on flows; 2) an in-network PFC causality analysis and tracing mechanism to quickly and efficiently collect causal telemetry for diagnosis; and 3) a provenance-based diagnosis algorithm to comprehensively present the anomaly breakdown, identifying the anomaly type and root causes accurately. Through extensive evaluations on both NS-3 simulations and a Tofino testbed, Hawkeye can quickly and accurately diagnose multiple RDMA NPAs with over 90% precision and 1–4 orders of magnitude lower overhead than baselines. Menghao Zhang 0001, Xiao Li 0044, Qiyang Peng, Mingwei Xu 0001, Xiaohe Hu, Jiahai Yang 0001, Xingang Shi |
SIGCOMM | 2 |
| 2024 | Near-Lossless Gradient Compression for Data-Parallel Distributed DNN TrainingabstractData parallelism has become a cornerstone in scaling up the training of deep neural networks (DNNs). However, the communication overhead associated with synchronizing gradients across multiple nodes has emerged as a significant bottleneck, adversely affecting training efficiency and leading to a surge in large-scale distributed model training costs. By leveraging insights into the statistical characteristics of gradients, we present GComp, a near-lossless gradient compression scheme designed to reduce the communication burden during data-parallel training significantly. GComp develops an optimized Huffman encoding/decoding strategy to compress gradient exponents effectively. Additionally, it introduces an innovative multi-level quantization method for mantissa, complemented by a pruning strategy that eliminates zero-valued gradients. These integrated approaches significantly reduce the volume of data for synchronization, while virtually not affecting the DNN model's training accuracy. We conduct comprehensive evaluations of GComp, demonstrating that our method can decrease the communication volume by as much as 67.1%, and enhance training speed by up to 1.9×. Xue Li 0024, Cheng Guo 0007, Kun Qian 0021, Menghao Zhang 0001, Mengyu Yang, Mingwei Xu 0001 |
SoCC | 4 |
| 2024 | Kale: Elastic GPU Scheduling for Online DL Model TrainingabstractLarge-scale GPU clusters have been widely used for effectively training both online and offline deep learning (DL) jobs. However, elastic scheduling in most cases of resource schedulers is dedicated for offline model training where resource adjustment is planned ahead of time. The native autoscaling policy is on the basis of pre-defined threshold and, if applied directly in online model training, often suffers from belated resource adjustment, leading to diminished model accuracy. In this paper, we present Kale, a novel elastic GPU scheduling system to improve the performance of online DL model training. Through traffic forecasting and resource-throughput modeling, Kale automatically pinpoints the number of required GPUs that best accommodate the on-the-fly data samples before performing stabilized autoscaling. An advanced data shuffling strategy is further employed for balancing uneven samples among different training workers, thereby improving the runtime efficacy. Experiments show that Kale substantially outperforms the state-of-the-art solutions. Compared with the default HPA autoscaling strategy, Kale reduces the accumulated lag and downtime by 69.2% and 33.1%, respectively, whilst lowering the SLO violation rate from 19.57% to just 2.6%. Kale has been deployed at Kuaishou's production-level GPU clusters and successfully underpins real-time video recommendation and advertisement at scale. Renyu Yang, Jin Ouyang, Weihan Jiang, Tianyu Ye, Menghao Zhang 0001, Sui Huang, Chengru Song, Di Zhang 0026, Tianyu Wo, Chunming Hu |
SoCC | 6 |
| 2024 | Paraleon: Automatic and Adaptive Tuning for DCQCN Parameters in RDMA NetworksabstractRDMA is a kernel-bypass and transport-offload technology that provides high throughput and low delay for datacenter networks, and DCQCN is the default and most widely used congestion control algorithm in large-scale RDMA networks. DCQCN involves over 10 parameters at RNICs and switches, and their settings significantly affect network performance, currently relying heavily on exhaustive manual tuning. Although some automatic methods are proposed to tune a subset of DCQCN parameters, none of them comprehensively address all parameters at both RNICs and switches, resulting in compromised network performance. In this paper, we propose Paraleon, an automatic and adaptive system to tune DCQCN parameters comprehensively. We design a millisecond-level sketch-based monitoring mechanism for accurate network-wide measurement, which collects runtime metrics as feedback to guide the tuning process. We also analyze the complicated parameter impacts on network performance, and leverage an improved heuristic searching algorithm for timely performance optimization with better efficiency and convergence. We implement Paraleon and conduct extensive experiments in both NS3 simulations and a real-world testbed. The results show that Paraleon achieves$3.8 \% \sim 61.4 \%$higher performance than existing tuning schemes. Ziteng Chen, Menghao Zhang 0001, Jiahao Cao 0001, Yang Jing, Mingwei Xu 0001, Renjie Xie, Fangzheng Jiao, Xiaohe Hu |
ICNP | 2 |
| 2024 | LoRDMA: A New Low-Rate DoS Attack in RDMA Networks
Menghao Zhang 0001, Yuying Du, Ziteng Chen, Mingwei Xu 0001, Renjie Xie, Jiahai Yang 0001 |
NDSS | 2 |
| 2024 | Cactus: Obfuscating Bidirectional Encrypted TCP Traffic at Client SideabstractAs the mainstream encrypted protocols adopt TCP protocol to ensure lossless data transmissions, the privacy of encrypted TCP traffic becomes a significant focus for adversaries. They can leverage Deep Learning (DL) models to infer the sensitive information from encrypted TCP traffic by analyzing its packet size, direction, and timing information. To defend against such DL-based traffic analysis attacks, recent advances reshape the encrypted traffic and achieve desired results. However, they typically require deploying cooperative modules on both communication endpoints and only support specific applications, such as browsers. In this paper, we propose Cactus, a client-side plug-in to obfuscate bidirectional encrypted TCP traffic for a wide range of applications transparently using the inherent TCP semantics and the emerging eBPF technique. In particular, Cactus provides four effective operations to enable bidirectional traffic obfuscation while preserving communication semantics of applications. Besides, Cactus empowers users to specify which applications to conduct traffic obfuscation and what obfuscation level for each application. We conduct comprehensive experiments to demonstrate that Cactus can effectively obfuscate encrypted TCP traffic with low overhead to hinder the traffic analysis efforts in website fingerprinting and application identification. Renjie Xie, Jiahao Cao 0001, Yuxi Zhu, Yi He 0020, Hanyi Peng, Mingwei Xu 0001, Kun Sun 0001, Enhuan Dong, Qi Li 0002, Menghao Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 12 |
| 2024 | IMap: Toward a Fast, Scalable and Reconfigurable In-Network Scanner With Programmable SwitchesabstractNetwork scanning has been a standard measurement technique to understand a network’s security situations, e.g., revealing security vulnerabilities, monitoring service deployments. However, probing a large-scale scanning space with existing network scanners is both difficult and slow, since they are all implemented on commodity servers and deployed at the network edge. To address this, we introduce IMap, a fast, scalable and reconfigurable in-network scanner based on programmable switches. In designing IMap, we overcome key restrictions posed by computation models and memory resources of programmable switches, and devise numerous techniques and optimizations, including an address-random and rate-adaptive probe packet generation mechanism, and a correct and efficient response packet processing scheme, to turn a switch into a practical runtime-reconfigurable high-speed network scanner. We implement an open-source prototype of IMap, and evaluate it with extensive testbed experiments and real-world deployments in our campus network. Evaluation results show that even with one switch port enabled, IMap can survey all ports of our campus network (i.e., a total of up to 25 billion scanning space) in 8 minutes. This demonstrates a nearly 4 times faster scanning speed and 1.5 times higher scanning accuracy than the state of the art, which shows that IMap has great potentials to be the next-generation terabit network scanner with all switch ports enabled. Besides, our experiments also show that IMap supports the reconfiguration of scanning tasks at runtime, without incurring switch downtime. Leveraging IMap, we also discover several potential security threats in our campus network, and report them to our network administrators responsibly. Menghao Zhang 0001, Cheng Guo 0007, Han Bao 0011, Mingwei Xu 0001, Hongxin Hu |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Exploring Dynamic Rule Caching Under Dependency Constraints for Programmable Switches: Theory, Algorithm, and ImplementationabstractTernary Content Addressable Memory (TCAM) enables fast lookup and is widely used by routers and switches to support policy-based forwarding. Due to high cost and small capacity, only a small subset of important rules can be cached in TCAM, so determining it is critical to increasing the hit ratio. This is more challenging than traditional caching problems because of complicated rule dependency relationships. Existing works are based on heuristics and they don’t work well under all practical scenarios. Worse still, the lack of fundamental understanding of the design space, complexity, and optimality makes all explorations in mystery. In this paper, we use a modeling-based method to formulate the problem, prove its complexity, and propose DROPS, a dynamic rule caching framework with a much higher hit ratio. In particular, we deduce the rule selection problem into a multi-dimensional rule space transformation problem. Thus, we are no longer limited by using the intrinsic rules; rather, we can transform original rules into “new rules” equivalently without rule dependency. We design non-trivial rule placement and update algorithms and implement them in programmable switches. In the experimental evaluation, we show that our method outperforms all existing methods. Xinhao Deng 0001, Mingwei Xu 0001, Qi Li 0002, Weijie Wu, Yuan Yang 0001, Menghao Zhang 0001, Yu Zhou 0008 |
IEEE Trans. Netw. Serv. Manag. | 6 |
| 2023 | Multa: Enabling Cost-efficient Multi-task Network Traffic AnalysisabstractMachine learning based network traffic analysis has been widely used to combat malicious behaviors in various network scenarios. However, most of the applications are specialized for a single task and cannot cover a wide spectrum of security threats. Building a machine learning pipeline from scratch for each traffic analysis task is painstaking and time-consuming. And it is not practical to deploy several traffic analysis tasks together due to the high runtime overhead. In this paper, we propose Multa to enhance the deployability of traffic analysis applications. Multa is a framework that leverages Multi-Task Learning (MTL) - an emerging paradigm in machine learning - to enable the efficient parallel deployment of multiple traffic analysis tasks. Specifically, we devise a sequence-based traffic feature representation that is informative and general enough to serve a diverse range of tasks. Then we use the LSTM-based Multi-gate Mixture-of-Experts (LMMoE) model to learn task characteristics and propose a size tuning algorithm to optimize its performance. Considering different data curating methods, we also present two different training paradigms to construct Multa. Our experiments show that Multa can reduce the space occupancy of the model by up to 50% and accelerate the inference process by more than 100% in multi-task scenarios, which indicates better deployability in real-world networks. Cheng Guo 0007, Menghao Zhang 0001, Mingwei Xu 0001 |
GLOBECOM | 2 |
| 2023 | HyperClassifier: Accurate, Extensible and Scalable Traffic Classification with Programmable SwitchesabstractTraffic classification provides substantial benefits for service differentiation, security policy enforcement, and traffic engineering. However, accurately classifying large volumes of network traffic using existing solutions is pretty challenging, as they are typically implemented on commodity servers with slow CPUs for packet processing. To address this, we leverage the opportunity provided by emerging programmable switches and propose HyperClassifier as a solution to achieve accurate, extensible, and scalable traffic classification. HyperClassifier designs an efficient classifying table with an effective flow expiration mechanism that enables lightweight packet inspection on resource-limited switches. We implement an open-source prototype of HyperClassifier on a hardware Tofino switch and conduct extensive evaluations. The results of our evaluation demonstrate that, compared to existing solutions, HyperClassifier can provide orders of magnitude higher classification throughput with comparable classification accuracy. Yichi Xu, Jiamin Cao, Menghao Zhang 0001, Ying Liu 0024, Mingwei Xu 0001 |
ICC | 4 |
| 2023 | Poster: Chameleon: Automatic and Adaptive Tuning for DCQCN Parameters in RDMA NetworksabstractDatacenter Quantized Congestion Notification (DCQCN) [12] is the default congestion control algorithm for Mellanox RDMA (Remote Direct Memory Access) NICs [2] in RoCEv2 (RDMA over Converged Ethernet v2) networks, one of the most widely used NICs in leading industry companies [4, 5, 7, 9]. In DCQCN, firstly switches mark packets with ECN (Explicit Congestion Notification) when the queue length exceeds ECN thresholds, then receivers respond to ECN-marked packets with CNPs (Congestion Notification Packets), and finally senders reduce transmission rate when receiving CNPs. DCQCN has 10+ parameters at both NICs and switches, including Alpha Update, Rate Increase & Decrease, Notification Point and ECN thresholds [3], and these parameters have a non-negligible impact on the network performance. Our experiments also verify the network performance of common AI (Artificial Intelligence) training workloads in RoCEv2 networks (e.g., all-to-all collective communication) is greatly influenced by different DCQCN parameter settings (§3). Therefore, when deploying applications in practice, the DCQCN parameters need to be carefully tested and tuned to improve the network performance. Ziteng Chen, Menghao Zhang 0001, Mingwei Xu 0001 |
SIGCOMM | 2 |
| 2023 | Rosetta: Enabling Robust TLS Encrypted Traffic Classification in Diverse Network Environments with TCP-Aware Traffic Augmentation
Renjie Xie, Jiahao Cao 0001, Enhuan Dong, Kun Sun 0001, Qi Li 0002, Licheng Shen, Menghao Zhang 0001 |
USENIX Security Symposium | 8 |
| 2023 | NetHCF: Filtering Spoofed IP Traffic With Programmable SwitchesabstractIn this paper, we identify the opportunity of using programmable switches to improve the state of the art in spoofed IP traffic filtering, and proposeNetHCF, a line-rate in-network system to filter spoofed traffic. One key challenge in the design ofNetHCFis to handle the restrictions stemmed from the limited computational model and memory resources of programmable switches. We address this by decomposing the HCF scheme into two complementary parts, by aggregating the IP-to-Hop-Count (IP2HC) mapping table for efficient memory usage, and by designing adaptive mechanisms to handle routing changes, IP popularity changes, and network activity dynamics. We implement an open-source prototype ofNetHCF, and conduct extensive evaluations. The evaluation results demonstrate thatNetHCFis able to process most legitimate traffic in 1$\mu$s, filter spoofed IP traffic effectively under network dynamics, with less than 30% of switch resource occupation. Menghao Zhang 0001, Chang Liu 0021, Mingwei Xu 0001, Guofei Gu |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2023 | Bolt: Scalable and Cost-Efficient Multistring Pattern Matching With Programmable SwitchesabstractMulti-string pattern matching is a crucial building block for many network security applications and thus of great importance. Since every byte of a packet has to be inspected by a large set of patterns, it often becomes a bottleneck of these applications and dominates the performance of an entire system. Many existing studies have been devoted to alleviating this performance bottleneck either by algorithm optimization or hardware acceleration. However, neither one provides the desired scalability and costs that keep pace with the drastic increase in network bandwidth and traffic today. To address these issues, in this paper, we present BOLT, a scalable and cost-efficient multi-string pattern matching system leveraging the capability of emerging programmable switches. BOLT combines the following techniques: (1) an efficient state encoding scheme to fit a large number of strings into the limited memory on a programmable switch; (2) a variable$k$-stride transition mechanism to increase the throughput significantly with the same level of memory cost; and(3)a compactpattern2rulemapping method to accommodate multiple co-existing strings in one rule. We implement a prototype of BOLT and make its source code publicly available. Extensive evaluations demonstrate that BOLT can provide multi-hundred Gbps throughput and scales well with various pattern sets and workloads. Menghao Zhang 0001, Chang Liu 0021, Ying Liu 0024, Mingwei Xu 0001 |
IEEE/ACM Trans. Netw. | 2 |
| 2022 | IMap: Fast and Scalable In-Network Scanning with Programmable Switches
Menghao Zhang 0001, Cheng Guo 0007, Han Bao 0011, Mingwei Xu 0001, Hongxin Hu |
NSDI | 2 |
| 2022 | NetEC: Accelerating Erasure Coding Reconstruction With In-Network AggregationabstractIn distributed storage systems, Erasure Coding (EC) is a crucial technology to enable high data availability. By downloading parity data from survived machines, EC can reconstruct lost data with much lower storage overheads than data replication. However, this reduction in storage cost comes at the expense of extra performance problems:low reconstruction rate,high degraded read latency, andhigh host CPU utilization. Our analysis shows that these performance problems are deeply rooted in thehost-basedEC processing. To resolve these problems, we present NetEC, an in-network accelerating framework that fully offloads EC to the new generation programmable switching ASICs. We propose Explicit Buffer Size Notification (EBSN) to constrain decoding buffer usage, and design an on-switch one-to-many TCP proxy to integrate EBSN with TCP. We also design two parallel Galois Field (GF) offloading methods—table lookup and bitmatrix methods—to maximize parsable bytes. We implement NetEC on programmable switches and integrate it with HDFS. Extensive evaluations show that NetEC improves the reconstruction rate by 2.7x-6.8x, reduces the degraded read latency significantly, and removes the host CPU overhead completely. We also emulate multi-rack scenarios and show that NetEC is able to support$\sim$∼GB/s reconstruction rate and tens of concurrent tasks. Yi Qiao, Menghao Zhang 0001, Yu Zhou 0008, Han Zhang 0009, Mingwei Xu 0001, Jun Bi, Jilong Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | Switches are Scanners Too!: A Fast and Scalable In-Network Scanner with Programmable SwitchesabstractNetwork scanning has been a standard measurement technique to understand the network's security situations, however, probing a large-scale scanning space with existing network scanners is both difficult and slow. To address this issue, we introduce IMap, a fast and scalable in-network scanner based on programmable switches. In designing IMap, we overcome key restrictions posed by computation models and memory resources of programmable switches, and devise numerous techniques and optimizations to turn a switch into a practical high-speed network scanner. We conduct preliminary experiments on the open-source prototype of IMap and evaluation results show that IMap can survey all addresses (i.e., 6 Class B Addresses) and all ports of our campus network in 8 minutes, nearly 4 times faster than state-of-the-art network scanners. As an ongoing work, we plan to continuously improve the design and implementation of IMap, and hope IMap can serve as a foundation for designing next-generation terabit network scanners. Menghao Zhang 0001, Cheng Guo 0007, Han Bao 0011, Mingwe Xu, Hongxin Hu |
HotNets | 2 |
| 2021 | Making Multi-String Pattern Matching Scalable and Cost-Efficient with Programmable Switching ASICsabstractMulti-string pattern matching is a crucial building block for many network security applications, and thus of great importance. Since every byte of a packet has to be inspected by a large set of patterns, it often becomes a bottleneck of these applications and dominates the performance of an entire system. Many existing works have been devoted to alleviate this performance bottleneck either by algorithm optimization or hardware acceleration. However, neither one provides the desired scalability and costs that keep pace with the dramatic increase of the network bandwidth and network traffic today. In this paper, we present BOLT, a scalable and cost-efficient multi-string pattern matching system leveraging the capability of emerging programmable switches. BOLT combines the following two techniques, a smart state encoding scheme to fit a large number of strings into the limited memory on the programmable switch, and a variable k-stride transition mechanism to increase the throughput significantly with the same level of memory costs. We implement a prototype of BOLT and make its source code publicly available. Extensive evaluations demonstrate that BOLT could provide orders of magnitude improvement in throughput which is scalable with pattern sets and workloads, and could also significantly decrease the number of entries and memory requirement. Menghao Zhang 0001, Chang Liu 0021, Ying Liu 0024, Xuya Jia, Mingwei Xu 0001 |
INFOCOM | 2 |
| 2021 | Enabling Performant, Flexible and Cost-Efficient DDoS Defense With Programmable SwitchesabstractDistributed Denial-of-Service (DDoS) attacks have become a critical threat to the Internet. Due to the increasing number of vulnerable Internet of Things (IoT) devices, attackers can easily compromise a large set of nodes and launch high-volume DDoS attacks from the botnets. State-of-the-art DDoS defenses, however, have not caught up with the fast development of the attacks. Middlebox-based defenses can achieve high performance with specialized hardware; however, these defenses incur a high cost, and deploying new defenses typically requires a device upgrade. On the other hand, software-based defenses are highly flexible, but software-based packet processing leads to high performance overheads. In this article, we propose Poseidon, a system that addresses these limitations in today's DDoS defenses. It leverages emerging programmable switches, which can be reconfigured in the field without additional hardware upgrades. Users of Poseidon can specify their defense strategies in a modular fashion in the form of a set of defense primitives; this can be further customized easily for each network and extended to include new defenses. Poseidon then maps the defense primitives to run on programmable switches-and when necessary, on server software-for effective defense. When attacks change, Poseidon can reconfigure the underlying defense primitives to respond to the new attack patterns. Evaluations using our prototype demonstrate that Poseidon can effectively defend against high-volume attacks, easily support customization of defense strategies, and adapt to dynamic attacks with low overheads. Menghao Zhang 0001, Chang Liu 0021, Mingwei Xu 0001, Ang Chen 0001, Hongxin Hu, Guofei Gu, Qi Li 0002 |
IEEE/ACM Trans. Netw. | 2 |
| 2021 | Control Plane Reflection Attacks and Defenses in Software-Defined NetworksabstractSoftware-Defined Networking (SDN) continues to be deployed spanning from enterprise data centers to cloud computing with the proliferation of various SDN-enabled hardware switches and dynamic control plane applications. However, state-of-the-art SDN-enabled hardware switches have rather limited downlink message processing capability, especially for Flow-Mod and Statistic Query, which may not suffice the huge need of dynamic control plane applications. In this paper, we systematically study the interactions between the control plane applications and the data plane switches, and present two new attacks, namely Control Plane Reflection Attacks, to exploit the limited processing capability of SDN-enabled hardware switches. The reflection attacks adopt direct and indirect data plane events to force the control plane to issue massive expensive downlink messages towards SDN switches. Moreover, we propose a two-phase probing-triggering attack strategy, which makes the reflection attacks much more efficient and powerful. Experiments on a testbed with 3 different physical OpenFlow switches demonstrate that the attacks can lead to catastrophic results such as hurting the establishment of new flows and even disruption of connection between SDN controller and switches. To mitigate such attacks, we present several countermeasures from different perspectives. In particular, we propose a novel, systematical defense framework, SwitchGuard, to detect anomalies of downlink messages and prioritize these messages based on a novel monitoring granularity, i.e., host-application pair (HAP). Implementations and evaluations demonstrate that SwitchGuard can effectively reduce the latency for legitimate hosts and applications under the control plane reflection attacks with only minor overheads. Menghao Zhang 0001, Lei Xu 0024, Jiasong Bai, Mingwei Xu 0001, Guofei Gu |
IEEE/ACM Trans. Netw. | 1 |
| 2020 | Poseidon: Mitigating Volumetric DDoS Attacks with Programmable Switches
Menghao Zhang 0001, Chang Liu 0021, Ang Chen 0001, Hongxin Hu, Guofei Gu, Qi Li 0002, Mingwei Xu 0001 |
NDSS | 1 |
| 2019 | NETHCF: Enabling Line-rate and Adaptive Spoofed IP Traffic FilteringabstractIn this paper, we design NETHCF, a line-rate in-network system for filtering spoofed traffic. NETHCF leverages the opportunity provided by programmable switches to design a novel defense against spoofed IP traffic, and it is highly efficient and adaptive. One key challenge stems from the restrictions of the computational model and memory resources of programmable switches. We address this by decomposing the HCF system into two complementary components-one component for the data plane and another for the control plane. We also aggregate the IP-to-Hop-Count (IP2HC) mapping table for efficient memory usage, and design adaptive mechanisms to handle end-to-end routing changes, IP popularity changes, and network activity dynamics. We have built a prototype on a hardware Tofino switch, and our evaluation demonstrates that NETHCF can achieve line-rate and adaptive traffic filtering with low overheads. Menghao Zhang 0001, Chang Liu 0021, Ang Chen 0001, Guofei Gu, Hai-Xin Duan |
ICNP | 2 |
| 2019 | When NFV Meets ANN: Rethinking Elastic Scaling for ANN-based NFsabstractNetwork Function Virtualization (NFV) provides middleboxes with substantial elasticity from a system level, and Artificial Neural Network (ANN) empowers middleboxes with great intelligence from an algorithm-level perspective. However, when ANN-based Network Functions (NFs) want to take advantage of the elasticity of NFV, our study finds that huge gaps exist between the existing approaches and the ideal goals for the elasticity control of ANN-based NFs. By revealing the key differences between ANN-based NFs and traditional NFs, we propose LEGO, an innovative framework that provides systematic mechanisms for traffic splitting, instance partition and runtime management to enable correct and efficient scaling of ANN-based NFs. Preliminary implementation and evaluation demonstrate the feasibility and effectiveness of the LEGO system. The major purpose of this paper is to highlight these challenges and sketch out a new roadmap towards ANN-based NFV paradigm. Menghao Zhang 0001, Jiasong Bai, Zili Meng, Hongda Li 0002, Hongxin Hu, Mingwei Xu 0001 |
ICNP | 1 |
| 2019 | Tripod: Towards a Scalable, Efficient and Resilient Cloud GatewayabstractCloud gateways are fundamental components of a cloud platform, where various network functions (e.g., L4/L7 load balancing, network address translation, stateful firewall, and SYN proxy) are deployed to process millions of connections and billions of packets. Providing high-performance and failure-resilient packet processing with a scalable traffic management mechanism is crucial to ensuring the quality of service of a cloud provider, and hence is of great importance. Many network functions nowadays are implemented in software with commodity servers for low cost and high flexibility. However, existing software-based network function frameworks oftentimes provide part of these features, while cannot satisfy all three requirements above simultaneously. To address these issues, in this paper, we introduce TRIPOD, a novel network function framework specialized for cloud gateways. Having identified the fundamental limitations of loosely coupling traffic, processing logic and state, TRIPOD jointly manages these three elements with the unique characteristics of cloud gateways, which is enabled by a simple, efficient traffic processing mechanism, and a high performance state management service. Adopting several effective techniques and optimizations, TRIPOD is able to achieve scalable traffic management (<;100 flow rules for even ~Tbps traffic), high performance (reducing 40% of latency compared with state of the art) and failure resilience (similar packet/connection loss rate compared to state of the art), with reasonable overheads (less than 10% of the workload traffic) even under an extremely heavy traffic, making it a good fit for cloud gateways. Menghao Zhang 0001, Jun Bi, Kai Gao 0001, Yi Qiao, Zhaogeng Li, Hongxin Hu |
IEEE J. Sel. Areas Commun. | 1 |
| 2018 | Control Plane Reflection Attacks in SDNs: New Attacks and Countermeasures
Menghao Zhang 0001, Lei Xu 0024, Jun Bi, Guofei Gu, Jiasong Bai |
RAID | 1 |