EDBT 2026 Demo / reviewers in the wild / expert
Longlong Zhu
dblp:257/2179
· DBLP profile ↗
42ranked-venue papers
7as first author
41since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 34 · 4 first-author · 34 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Turbolearn: Harnessing Accurate and Line-Rate Deep Learning on Programmable Switches
Zhifan Jiang, Longlong Zhu, Jiashuo Yu, Linying Zheng, Chunming Wu 0001, Xiang Chen 0017 |
APNet | 2 |
| 2026 | TurboLearn: Harnessing Accurate and Line-Rate Deep Learning on Programmable SwitchesabstractThe intelligent data plane (IDP) embeds deep learning (DL) models on switches for line-rate traffic analysis, but hardware constraints often force simplified models, reducing accuracy, while complex models like Transformers remain undeployable. We present TurboLearn, which achieves high accuracy and line-rate performance by co-designing inference across the switch ASIC and switch OS on the same switch. TurboLearn uses three key techniques: (1) a hardware fast path for lightweight classification and a software normal path for complex models, (2) confidence-based selective inference that escalates only low-confidence packets, and (3) confidence-calibrated knowledge distillation, where the normal path teaches the fast path. On Intel Tofino2 switches across three real-world traffic tasks, TurboLearn supports models that existing IDPs cannot deploy, improves macro-F1 by up to 31.31%, and keeps over 90% of traffic on the fast path with zero throughput loss. Zhifan Jiang, Longlong Zhu, Jiashuo Yu, Linying Zheng, Chunming Wu 0001, Xiang Chen 0017 |
APNet | 2 |
| 2026 | APTMatch: Empowering Dynamic Rule Updating for Learning-based Packet Classification
Jiashuo Yu, Longlong Zhu, Zongye Lin, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001 |
ICC | 3 |
| 2026 | HYDRA: A Hybrid Synthesizer for Asymmetric Mixture-of-Experts Communication Scheduling
Lida Liao, Qianxun Xu, Xuanwei Si, Hongyan Liu 0001, Jiashuo Yu, Zongye Lin, Qiaoling Hu, Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001 |
ICC | 11 |
| 2026 | DCAD: Dual-Condition Adaptive Consistency Model for IoT Privacy-Preserving Video Anomaly Detection
Shaopeng Zhou, Chaohao Li, Haonan Yan, Longlong Zhu, Chunming Wu 0001, Bin Wang 0062 |
ICIC (2) | 5 |
| 2026 | MonPlan: Taming Network Measurement with Accurate and Resource-Efficient Sketch-INT Co-Design
Xiang Chen 0017, Linying Zheng, Longlong Zhu, Zedi Chen, Qing Shu, Jialu Tian, Siqi Dong, Qun Huang 0001, Jianshan Zhang, Xuan Liu 0006, Haifeng Zhou, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001 |
INFOCOM | 3 |
| 2026 | LTD: Low-Overhead Topology Discovery using Programmable Data Planes
Dezhang Kong, Minghao Li 0012, Shi Lin, Zhenhua Xu 0004, Longlong Zhu, Linying Zheng, Xiang Chen 0017, Changting Lin, Xuan Liu 0006, Dong Zhang 0010, Chunming Wu 0001 |
INFOCOM | 5 |
| 2026 | SketchPipe: Toward Accurate Sketch-based Network Measurement on Multi-Pipeline Switches with Splitless Sketch Placement
Xiang Chen 0017, Longlong Zhu, Linying Zheng, Hongyang Du 0001, Dong Zhang 0010, Jianshan Zhang, Xuan Liu 0006, Qun Huang 0001, Dusit Niyato, Haifeng Zhou, Chunming Wu 0001, Hongyan Liu 0001, Kui Ren 0001 |
NSDI | 2 |
| 2026 | Proteus: Towards Accurate and Low-overhead In-Network Malicious Traffic DetectionabstractNetwork intrusion detection systems (NIDS) are essential for web security by identifying and dropping malicious traffic. Existing in-network NIDS leverage the Tbps-level packet processing capability of programmable switches to achieve high-speed flow classification. They translate complex trained machine learning models to decision trees (DTs), where DTs are deployed on programmable switches via single-DT or multiple-DT deployment. However, they face a fundamental trade-off: single-DT deployment suffers from low classification accuracy due to over-pruning of trees, while multiple-DT deployment suffers from high overhead due to deploying multiple tree replicas. In this paper, we propose Proteus, an in-network malicious traffic detection system that achieves both high classification accuracy and low overhead. Its key idea is to split the original DT into critical and normal sub-trees, where these sub-trees have different impacts on overall accuracy. More precisely, Proteus first splits a DT into one critical and several normal sub-trees for adapting to the accuracy requirement and switch resource budgets. Second, it minimizes coordination overhead between sub-trees while ensuring full flow coverage via mixed-integer linear programming. Third, it dynamically reallocates or migrates sub-trees to adapt to changing resources by monitoring both classification accuracy and switch resource changes. Testbed experiments with 12.8 Tbps programmable switches show that Proteus improves classification accuracy, reduces switch resource consumption, and reduces classification latency. Longlong Zhu, Linying Zheng, Qing Shu, Zedi Chen, Jiashuo Yu, Shaopeng Zhou, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001, Xiang Chen 0017 |
WWW | 1 |
| 2026 | DeepConfig: A verifiable configuration generation framework for MAN overlays using LLMs
Longlong Zhu, Hongyan Liu 0001, Dong Zhang 0010, Jiashuo Yu, Lida Liao |
Comput. Networks | 2 |
| 2026 | Energy-Efficient Multi-UAV-Assistant Data Collection for Multisensor Marine NetworksabstractAs the marine economy continues to expand, the importance of efficient and reliable marine data collection has become increasingly evident. This paper investigates a multi-unmanned aerial vehicle (UAV)-assisted marine data collection network system, where multiple UAVs are deployed within a designated area to collect data from water buoy sensors (WBSs) and act as relays to offload the collected data to a central ship. The primary objective is to minimize the total system energy consumption, subject to constraints on access relationships, power scheduling, and movement trajectories. The formulated optimization problem is non-convex and highly complex due to the coupling of multiple variables. To address this challenge, we propose an alternating optimization algorithm that jointly optimizes the trajectories of multiple UAVs, the access selection of WBSs, the trajectory of the ship, the data offloading decisions, and the UAV transmit power scheduling in an iterative manner. The algorithm leverages techniques such as successive convex approximation (SCA) and greedy strategies to efficiently solve the decomposed sub-problems. Simulation results demonstrate that the proposed approach achieves significant performance improvements compared to several benchmark algorithms, highlighting its effectiveness in enhancing energy efficiency and system robustness in dynamic marine environments. Longlong Zhu, Rui Ming, Guolong Zheng, Jianshan Zhang |
IEEE Internet Things J. | 3 |
| 2025 | DHC: Distributed Homomorphic Compression for Gradient Aggregation in AllreduceabstractDistributed training is critical for efficiently developing deep neural networks (DNNs) on tasks like image classification and natural language processing. However, as model and dataset sizes continue to grow, high communication overhead during gradient exchanges has become a major bottleneck in distributed training. Although existing homomorphic compression frameworks effectively reduce communication overhead, their reliance on centralized architectures makes them unsuitable for the mainstream decentralized AllReduce architecture. To address this, we propose DHC, a framework for homomorphic gradient compression in AllReduce architectures. Its key idea is HG-Sketch, which leverages multi-level index tables for direct in-network aggregation of compressed gradients, thereby eliminating additional computational overhead. Additionally, DHC introduces an index-sharing method to optimize memory usage on programmable switches. Furthermore, we establish an Integer Linear Programming (ILP) model to optimize the deployment strategy of programmable switches, further enhancing in-network aggregation capabilities. Experimental results demonstrate that DHC achieves a$3.8 \times$increase in aggregation speed and a$4.2 \times$improvement in aggregation throughput. Lida Liao, Zhengli Lin, Longlong Zhu, Hongyan Liu 0001, Jiashuo Yu, Dong Zhang 0010, Chunming Wu 0001 |
ICC | 4 |
| 2025 | P4Alex: A Scalable Range Matching Approach for Programmable SwitchesabstractRange matching (RM), a flexible primitive for implementing network applications on programmable switches, is often subject to limited TCAM resources. Consequently, existing RM approaches rely on SRAM/ALU-assisted data structures to extend TCAM capacity. However, they consume significant SRAM, ALU, and pipeline stages, which hinders the implementation of other primitives (e.g., basic forwarding). In this paper, we propose P4Alex, a scalable RM framework for programmable switches. The key idea is leveraging emerging learned index structures to support RM, which replaces the storage by model inference for lightweight and fixed index depth. Unfortunately, the learning index structure cannot be directly implemented on programmable switches due to hardware limitations (e.g., floating-point computation and no-loop operations). In response, we design several optimizations in P4Alex to make it deployable. We successfully implemented P4Alex on Intel Tofino switches. Experimental results show that, compared to existing RM approaches, P4Alex extends RM capabilities from tens of thousands to millions. At the same scales, P4Alex reduces TCAM and SRAM resource consumption by up to 89.9% and 43.7%, respectively, while increasing latency by only$0.48 \mu ~\mathrm{s}$. Jiashuo Yu, Longlong Zhu, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001 |
ICC | 2 |
| 2025 | Carrera: Enabling High-Performance eBPF-based Sketches in Network MeasurementabstractTo achieve dynamic network measurement, trends build sketches on eBPF to avoid service interruptions. However, existing eBPF-based sketches suffer from high CPU consumption, leading to poor throughput and high latency and making them hard to measure high-speed traffic. Optimizing their performance requires users to refactor codes based on each sketch’s characteristics on eBPF, which is highly complex and time-consuming.In this paper, we argue that users should write sketches without concerning low-level eBPF performance optimizations, with the deployment automatically activating cross-sketch performance optimizations. We present Carrera, a library that offers domain-specific optimizations for eBPF-based sketches. Our contributions are (1) systematically analyzing the performance bottlenecks of eBPF-based sketches through microbenchmarks, (2) identifying practical optimizations, including hardware offloading, SIMD-accelerated hashing, traffic-aware flow index caching, prefetched randomization, and active data collection, to address the identified bottlenecks in eBPF-based sketches, (3) evaluating these optimizations with state-of-the-art sketches and demonstrating that Carrera improves throughput by up to 65% and reduces latency by up to 93% via testbed experiments. Xiang Chen 0017, Xin Yao 0008, Longlong Zhu, Linying Zheng, Hongyan Liu 0001, Jianshan Zhang, Dong Zhang 0010, Xuan Liu 0006, Qun Huang 0001, Haifeng Zhou, Chunming Wu 0001 |
ICNP | 4 |
| 2025 | EffiMatch: Enabling Fast and Accurate Learning-based Packet ClassificationabstractLearning-based Packet Classification methods reduce memory overhead by using lightweight Recursive Model Index(RMI) structures to limit the search range, followed by linear matching. However, they face a trade-off: complex RMI structures achieve smaller search ranges but slow down lookup, while simpler ones are faster but require larger scans. In this paper, we propose EffiMatch, a parallel multi-model lookup architecture aimed at resolving the trade-off between RMI complexity and linear search range in learning-based index systems. We propose two key designs: 1) We design a partitioning strategy called Distribution-Distance Partitioning (DDP), which groups data points with similar trends into the same segment. Combined with parallel lookup, this reduces the linear search range while maintaining high lookup speed. 2) We propose a more fine-grained binarization method, Base-Index Representation (BI), which approximates floating-point operations using integers. This method further reduces the search range without increasing model complexity. Experimental results show that EffiMatch reduces the linear search range by 26.84% using lower-complexity RMI models, which improves lookup speed by up to 6× and reduces construction time by up to 4 orders of magnitude compared to state-of-the-art LPC methods. Lida Liao, Jiashuo Yu, Longlong Zhu, Hongyan Liu 0001, Dong Zhang 0010, Xiang Chen 0017, Chunming Wu 0001 |
ICNP | 4 |
| 2025 | TurboCache: Empowering Switch-Accelerated Key-Value Caches with Accurate and Fast Cache UpdatesabstractRecent key-value (KV) caches are offloaded to programmable switches to offer high query processing performance. However, they suffer from both low accuracy in hot key detection and high latency in cache updates due to the strict limitations on switch registers. We propose TurboCache, a switch-accelerated KV cache with accurate hot key detection and fast cache updates. Our key idea is to leverage the switch recirculation capability to build a novel data structure that caches hot KV pairs. With this hardware-compatible cache data structure, TurboCache designs efficient data plane algorithms that accurately detects new hot keys and quickly updates its cache entirely within switch ASIC pipelines. We have implemented TurboCache on a${64}\times {100}$Gbps Tofino switch. Testbed results indicate that TurboCache improves the hot key detection accuracy and decreases the cache update latency of existing KV caches by several orders of magnitude. Xiang Chen 0017, Longlong Zhu, Linying Zheng, Lingfei Cheng, Jianshan Zhang, Xu Yang 0002, Dong Zhang 0010, Xuan Liu 0006, Xiaoming Lu, Xun Yi, Ibrahim Khalil 0001, Albert Y. Zomaya, Haifeng Zhou, Chunming Wu 0001 |
INFOCOM | 2 |
| 2025 | Scaling Learning-based Packet Classification Hardware with NeuTree
Jiashuo Yu, Longlong Zhu, Linying Zheng, Dong Zhang 0010, Xiang Chen 0017 |
INFOCOM | 2 |
| 2025 | Monica: Towards Scalable Distributed System Verification by Programmable Switch-Based TestingabstractData correctness in distributed systems is ensured by data consistency, where consistency is achieved by consensus algorithms. To safeguard data consistency, current testing tools use stress testing methods to examine consensus algorithms. However, existing tools are unable to simulate the situation under high traffic and suffer from excessive verification time. In this paper, we propose Monica, a scalable and efficient verification framework. Its key idea is to leverage the programmable switch to verify consensus algorithms. Specifically, Monica provides a set of primitives that researchers can invoke. Then, the control server recognizes the primitives and automatically configures the data plane. After that, the programmable switch collaborates with the control server to complete the verification. Experimental results show that Monica can generate traffic at the rate of Tbps level while keeping the computational and memory consumption of the programmable switch under 11.87%. Compared to existing testing tools, Monica increases the verification speed by up to 3.13 times. Further, Monica improved accuracy by 35.71% in high-traffic scenarios over other tools. Jiashuo Yu, Longlong Zhu, Dong Zhang 0010, Lida Liao, Rongbang Wu, Xiang Chen 0017, Chunming Wu 0001 |
IWQoS | 3 |
| 2025 | Polyx: Accelerating Verification of Traffic Migration in Large-Scale BGP NetworksabstractIn BGP networks, traffic migration verification ensures the scalability and reliability of the network during configuration changes. However, previous approaches suffer from low scalability and high computational overhead. In this poster, we propose Polyx, a framework for accelerating verification of traffic migration in large-scale BGP networks. Its key idea is to leverage hardware parallelism with a deterministic serialization algorithm to enhance state machine techniques. We implement the Polyx prototype and evaluate it on our built testbed. The experimental results demonstrate that Polyx achieves up to 46× overall speedup, 36× in state machine construction, and 131× in equivalence verification with minimal FPGA resource usage. Rongbang Wu, Longlong Zhu, Jiashuo Yu, Dong Zhang 0010, Hongyan Liu 0001, Zongye Lin, Lida Liao, Xiang Chen 0017, Chunming Wu 0001 |
IWQoS | 2 |
| 2025 | Handling Data Plane Program Deployment Dynamics with High-Quality Generative Diffusion ModelsabstractDeploying data plane programs across the network is typically formulated as a mixed-integer programming task, leading to a long execution time. In response, existing studies carefully tailor heuristics for specific task properties such as objectives. However, they suffer from poor solution quality under dynamic task deployment since they overfit specific task properties. Recently, generative diffusion models have been widely adopted in network optimizations due to their strong adaptability and generalization. Accordingly, in this poster, we propose a diffusion model-based framework for data plane program deployment tasks. Our key idea is to leverage the reverse denoising process of diffusion models to react to dynamic task changes at runtime while maintaining high solution quality. Preliminary results on our testbed show that we reduce latency by 66.67% and resource overhead by 58.62% during dynamic deployment. Longlong Zhu, Jiashuo Yu, Xiang Chen 0017, Qing Shu, Zedi Chen, Zhifan Jiang, Qun Huang 0001, Xuan Liu 0006, Dong Zhang 0010, Chunming Wu 0001 |
IWQoS | 1 |
| 2025 | UAVMamba: Elevating UAV's Crowd Counting Through a Synergistic Integration of Hybrid CNN and Mamba Paradigms
Longlong Zhu, Song Yuan, Mingjie Wang 0002 |
PRCV (15) | 1 |
| 2025 | Multistrategy Improved Particle Swarm Optimization Algorithm for Path Planning of UAV in 3-D Low Altitude Urban EnvironmentabstractThe Internet of Things (IoT) system and path planning algorithm provide a technological foundation for autonomous navigation of uncrewed aerial vehicles (UAVs). Geospatial data from the IoT system is transmitted to UAVs through lightweight protocols, and UAVs make optimal path decisions based on these data through optimization algorithms. The combination of the IoT, UAV, and path planning technology constitutes a UAV delivery system, which offers an efficient and economical solution for last-mile logistics in smart cities. Among these, rapid and accurate optimal path planning is crucial for the autonomous delivery of UAVs. Therefore, this article proposes a multistrategy improved particle swarm optimization (PSO) algorithm called MSIPSO. First, the algorithm incorporates a local deadlock jump strategy to increase the success rate of path planning in dense obstacle environments. Second, to mitigate the influence of parameter selection on the algorithm’s performance, adaptive nonlinear inertia weights and learning factors are introduced to improve the algorithm’s stability. Finally, multiple population differentiation evolution strategies are designed, with different position update equations tailored for populations of varying qualities, which enhances the search efficiency of the algorithm. The simulation results show that MSIPSO outperforms PSO, gray wolf optimizer (GWO), whale optimizer (WOA), elite archive-driven PSO (EAPSO) algorithm, and hybrid GWO and differential evolution (HGWODE) algorithm in terms of convergence speed, accuracy, and stability. Fazhan Tao, Zezheng Chen, Longlong Zhu, Jun Wang 0064 |
IEEE Internet Things J. | 4 |
| 2024 | MAReraser: Metal Artifact Reduction with Image Prior Using CNN and Transformer TogetherabstractThis paper presents a new dual domain network with image prior based on Convolutional Neural Network (CNN) and Transformer simultaneously for CT Metal Artifact Reduction (MAR). Challenges in MAR derive from the following aspects: firstly, the different morphologies of metal artifacts complexify resolving the issue just in a single domain; secondly, albeit many methods excel in quantitative metrics, yet the restored anatomical structures are over-smooth blurring reconstructed CT images; thirdly, MAR demands better performance as a clinical application, but the approaches relying on CNN or Transformer struggle due to CNN’s restricted spatial scope and Transformer’s ignorance to the local details, respectively, that is, CNN focuses on the local information while Transformer emphasizes the global information with higher computational complexity. To address these problems, we put forward MAReraser, a novel dual domain network, to deal with metal artifacts. MAReraser removes metal artifacts in both the projection and image domains, effectively reducing heteromorphic metal artifacts. Moreover, MAReraser introduces the image prior generated by an image prior subnet to refine the quality of reconstructed CT images. The prior subnet is pretrained in an expanded dataset which incorporates CT images corrected by diverse traditional MAR methods, providing extra potential prior knowledge from different perspectives. Further, the network backbone of MAReraser integrates CNN and Transformer, enabling complementary local and global feature extraction and balancing computational complexity. Extensive experiment results demonstrate that our method outperforms several other approaches whether in quantitative metrics or in qualitative visualization results. Songwei Zheng, Dong Zhang 0010, Chunyan Yu, Linghui Jia, Longlong Zhu, Zhanchao Huang, Danhong Zhu |
BIBM | 5 |
| 2024 | N4: Network for N Neural Network TrainingabstractAs the amount of data and complexity of neural network models continue to grow, distributed training has become increasingly crucial for improving training speed. However, the bottleneck of distributed training is the communication overheads among distributed workers. Recent research has shown that performing in-network aggregation using programmable switches is a good way to accelerate distributed training. However, previous work has only targeted specific neural network models and can only be applied in specified network topologies. Administrators may train different models and train them in different network topologies. In order to generalize the approach of using programmable switches to accelerate distributed training, we propose N4, a programmable intra-switch acceleration framework that supports distributed training of multiple neural networks. N4 also realizes the deployment of distributed workers based on any topology. Our experimental results show that N4 ensures high performance and isolation when training numerous neural networks. N4 outperforms state-of-the-art systems, accelerating training for existing methods by up to 3.4×. Shengrui Lin, Hongyan Liu 0001, Pengpai Shi, Longlong Zhu, Dong Zhang 0010 |
ICC | 6 |
| 2024 | OpenINT: Dynamic In-band Network Telemetry with Lightweight Deployment and Flexible PlanningabstractThe normal operation of data center network management tasks relies on accurate measurement of the network status. In-band Network Telemetry (INT) leverages programmable data planes to provide fine-grained and accurate network status. However, existing INT-related works have not considered the telemetry data required for dynamic adjustments of INT under uninterrupted conditions, including additions, deletions, and modifications. To address this issue, this paper proposes OpenINT, a lightweight and flexible In-band Network Telemetry system. The key innovation of OpenINT lies in decoupling telemetry operations in the data plane, using three generic sub-modules to achieve lightweight telemetry. Meanwhile, the control plane utilizes heuristic algorithms for dynamic planning to achieve near-optimal telemetry paths. Additionally, OpenINT provides primitives for defining network measurement tasks, which abstract the underlying telemetry architecture’s details, enabling network operator to conveniently access network status. A prototype of OpenINT is implemented on a programmable switch equipped with the Tofino chip. Experimental results demonstrate that OpenINT achieves highly flexible dynamic telemetry and significantly reduces network overhead. Jiayi Cai, Tingxin Sun, Zhengyan Zhou, Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001 |
INFOCOM | 5 |
| 2024 | TupleRadar: Accelerating Tuple Space Search in Packet Classification by Learned IndexabstractTuple space search(TSS)-based packet classification is the keystone of network system. Previous studies accelerate TSS by partitioning tuples, combining trees and tuples, and merging tuples. However, they do not scale with the number of rules, resulting in a high memory footprint or update time. In this paper, we propose TupleRadar, a framework for accelerating TSS while ensuring low memory footprint and fast rule updates. Our key idea is to construct learned indexes for tuples, which inherently improve the lookup speed but ensure the advantages of TSS. Specifically, TupleRadar builds orderly hash table-based tuples and then constructs the updatable learned index. It provides a bounded memory footprint of the index structure as well. We have evaluated TupleRadar on multiple scales rule-sets. Experimental results show that TupleRadar outperforms previous solutions, reducing 46.66% lookup time and 61.53% memory footprint on average, by up to 86.70% and 88.95%. It also performs a competitive rule update speed. Longlong Zhu, Jiashuo Yu, Kaiwei Huang, Zhengyan Zhou, Dong Zhang 0010, Xiang Chen 0010, Chunming Wu 0001 |
IWQoS | 1 |
| 2024 | CardSketch: Shift Attention for Network-wide Cardinality TelemetryabstractNetwork telemetry is an essential part of network management and infrastructure. Among them, cardinality telemetry provides statistics on network connectivity and distribution. Network-wide cardinality telemetry refers to the deployment of multiple telemetry nodes in network for cardinality estimate. This requires the deployed data structure to be mergeable, enabling the consolidation of data from different nodes. Unfortunately, existing mergeable data structures can’t simultaneously address two important criterions of cardinality telemetry: measurement accuracy and estimation interval. We propose CardSketch, aiming to adjust attention to cardinality telemetry based on changes of the network state. CardSketch incorporates a shift attention mechanism that leverages the randomness of hash functions to achieve unbiased transformations between data structures. This mechanism enables real-time selection of cardinality estimation methods based on the network’s state while preserving the original telemetry information as much as possible during the attention shift. We have implemented prototypes of CardSketch in software and hardware. Through extensive experimentation, the results demonstrate that CardSketch achieves excellent cardinality telemetry with minimal memory overhead. Even with a mere 50KB of memory space, it achieves a measurement precision of 87.75% and a measurement recall of 91.49%. Additionally, CardSketch supports multi-point aggregation and arbitrary partial key queries. Hanze Chen, Zhengyan Zhou, Pengpai Shi, Yanni Wu, Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001 |
LCN | 6 |
| 2024 | DOT: Towards Fast Decision Tree Packet Classification by Optimizing Rule PartitionsabstractPacket classification is a crucial component of modern networks. Existing decision tree-based algorithms alleviate the rule replication problem caused by overlapping rules in the ruleset via rule partitioning. They partition the ruleset into multiple subsets based on rule characteristics to reduce rule overlaps. However, existing algorithms fail to address the overlap between rules in the same set, seriously decreasing speed and memory performance. In this paper, we propose DOT, a framework for optimizing rule partitions before constructing decision trees. Its key idea is to migrate rules in subsets based on rule overlaps and the features of heuristics used to construct trees, as well as reorganize rules aided by tuples. DOT finds out the migrated rule candidates using rule dependency graphs and heuristic features, then transforms the rule migration problem into an integer linear programming problem and solves for the optimal migration strategy. Further, we employ a tuple-assisted approach to accelerate rule matching. Experiments show that DOT enhances existing decision tree-based algorithms, improving lookup speed by 1.69 ×, reducing average 24.85% memory consumption and 31.03% decision tree depth. Longlong Zhu, Jiashuo Yu, Linying Zheng, Dong Zhang 0010, Chunming Wu 0001 |
LCN | 2 |
| 2024 | TransTuple: Toward Fast Packet Classification via Adaptive Tuple ReplacementabstractOpen vSwitch (OVS) is a widely used software switch in virtualized environments and software-defined networks. OVS uses tuple space search (TSS) for packet classification in the datapath, allowing fast network rule updates, but the increasing number of rules poses a classification performance challenge. To address this, existing methods incorporate decision trees with TSS to form a hybrid structure, enhancing classification speed. However, decision trees tend to overfit the initial ruleset, becoming unbalanced after rule updates and leading to a sharp decline in classification performance. In this paper, we propose TransTuple, a framework to optimize hybrid structures for fast packet classification under rule updates. The core idea of TransTuple is to identify bottleneck branches in decision trees that degrade performance and to replace them with lightweight tuples, providing better throughput under rule updates. These tuples maintain rules using hash tables, enabling fast updating and packet matching on bottleneck branches. We use TransTuple to optimize three state-of-the-art hybrid structured methods, i.e., CutTSS, TabTree, and MBitTree, achieving up to a 3.1x improvement in classification speed during rule updates. Jiashuo Yu, Longlong Zhu, Rongbang Wu, Linying Zheng, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001 |
SECON | 2 |
| 2024 | P4Rex: Accelerating regular expression matching with programmable switches
Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001 |
Comput. Networks | 4 |
| 2024 | Optimizing fuel economy of fuel cell hybrid electric vehicle based on energy management strategy with integrated rapid thermal regulation
Xiaolong Tian, Fazhan Tao, Zhumu Fu, Longlong Zhu, Haochen Sun 0002, Shuzhong Song |
Eng. Appl. Artif. Intell. | 4 |
| 2023 | Vision Transformer with Progressive Tokenization for CT Metal Artifact ReductionabstractHigh-quality Computed Tomography(CT) plays a vital role in clinical diagnosis, but the presence of metallic implants will introduce severe metal artifacts on CT images and obstruct doctors’ decision-making. Many prior researches on Metal Artifact Reduction(MAR) are based on Convolutional Neural Network(CNN). Recently, Transformer has demonstrated phenomenal potential in computer vision. Also, transformer-based methods have been harnessed in CT image denoising. Nevertheless, these methods have been little explored in MAR. To fill the gap, we put forth, to the best of our knowledge, the first transformer-based architecture for MAR. Our method relies on a standard Vision Transformer(ViT). Furthermore, we tap into the progressive tokenization to refrain from the simple tokenization of ViT which gives rise to inability to model the local anatomical information. Additionally, for the sake of facilitating the interaction among tokens, we take advantage of cyclic shift from Swin Transformer. Finally, many experiment results reveal that the transformer-based technique is superior to those on the basis of CNN to some degree. Songwei Zheng, Dong Zhang 0010, Chunyan Yu, Danhong Zhu, Longlong Zhu, Zhongzheng Huang |
ICASSP | 5 |
| 2023 | MINT: Empowering Multiple Flow Definition Query for Network-Wide MeasurementabstractNetwork management tasks rely on precise and fine-grained network information to make correct and appropriate decisions. These tasks (e.g., DDoS detection) require network information with multiple flow definitions to better manage the network. However, the existing works mainly focus on the query of multiple flow definitions on a single switch, without a thoughtful solution for this query in network-wide measurement. In this paper, to address this problem, we overcome several challenges and propose MINT, a system that enables the query for multiple flow definitions in network-wide measurement. The key insights of MINT are: deploying MFSketch to measure multiple flow definitions information on the switch, cutting MFSketch into fixed-size slices, and using in-band telemetry (INT) to carry the slice to the analyzer. Therefore, after the analyzer collects and reorganizes the slices, network operators can query multiple flow definitions information of the whole network for various network management tasks. We implemented a prototype of MINT on a Barefoot Tofino switch. Experimental results show that MINT provides reliable transmission and consistency guarantees while only using switch resources comparable to state-of-the-art works, with less than 1% additional network overhead. Additionally, MFSketch provides accurate measurements for multiple flow definitions query, outperforming other solutions in both accuracy and F1 score. Jiayi Cai, Zhengyan Zhou, Tingxin Sun, Jiashuo Yu, Longlong Zhu, Chengze Li, Dong Zhang 0010, Chunming Wu 0001 |
ICC | 5 |
| 2023 | MiCuts: Combing Bit-Based Cutting and Splitting for Efficient Packet ClassificationabstractPacket classification is a crucial component in computer networking. To achieve high throughput and low memory consumption, existing solutions apply different heuristics in each construction stage to build efficient decision trees. However, previous studies divide the tree construction process based on the scale of rule subsets which is indirect to the performance goal, leading to massive rule replication and high tree depth. In this paper, we propose MiCuts, a fine-grained framework for packet classification with both high speed and low memory footprint. Its key idea is directly utilizing rule replication and tree depth to divide the tree-building process into three stages, each with suitable optimization goals. First, it partitions rules and builds shallow semi-trees without rule replication via selecting effective bits. Second, it transforms the switching problem of heuristics into an ILP problem and aims to minimize memory consumption while ensuring high lookup speed. Third, it merges some nodes to eliminate memory explosion caused by splitting, where MiCuts combines splitting and linear search. Extensive experimental results on ClassBench show that MiCuts outperforms state-of-the-art approaches, improving lookup speed by 1.71× while reducing memory footprint by 74.4% on average. Longlong Zhu, Jiashuo Yu, Linying Zheng, Jinfeng Pan, Zhengyan Zhou, Hanze Chen, Dong Zhang 0010, Xiang Chen 0010, Chunming Wu 0001 |
ICC | 1 |
| 2023 | A Network Function Virtualization Resource Allocation Model Based on Heterogeneous ComputingabstractWith the continuous increase in the speed and quantity of network traffic, higher performance requirements are put forward for the NFV system. Traditional virtualization technology is limited by slowing down of increase in CPU performance. There has some research on the use of hardware to accelerate network functions. However, existing method only consider use one hardware for acceleration, but hardware itself has limitation. It is difficult to match the network functions and hardware characteristics by considering the combination of each hardware and CPU discretely. In this paper, we propose a resource allocation model of network function virtualization(NFV) based on heterogeneous computing, which can maximize the resource utilization and obtain the global optimal solution. Our experiments prove that the genetic algorithm solution method we propose can take into account both the solution accuracy and the solution speed, and obtain an accurate Pareto curve under the multi-objectives optimization model. Hanze Chen, Lingfei Cheng, Longlong Zhu, Dong Zhang 0010 |
ISCC | 4 |
| 2023 | P4CTM: Compressed Traffic Pattern Matching Based on Programmable Data PlaneabstractPattern matching is an important technology applied to many security applications. Most network service providers choose to compress network traffic for better transmission, which brings the challenges of compressed traffic matching. However, existing works focus on improving the performance of uncompressed traffic matching or only realize the compressed traffic matching on end-host that can not keep pace with the dramatic increase in traffic. In this paper, we present P4CTM, a proof-of-concept method to conduct efficient compressed traffic matching on the programmable data plane. P4CTM uses the two-stage scan scheme to skip some bytes of compressed traffic, the 2-stride DFA combines with the compression algorithm to condense the state space, and the wildcard match to downsize the match action tables in the programmable data plane. The experiment indicates that P4CTM skips 83.10% bytes of compressed traffic, condenses the state space by order of magnitude, and reduces most of the table entries. Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001 |
ISCC | 4 |
| 2023 | DTRadar: Accelerating Search Process of Decision Trees in Packet ClassificationabstractPacket classification is an essential part of computer networks. Existing algorithms propose a partition process to address the memory explosion problem of the decision tree algorithm caused by the huge number of rules with multiple fields. However, the search process requires traversing multiple trees generated by the partition, which reduces the search efficiency. The existing algorithms take simple approaches to optimize the search process, which is low efficiency or high hardware overhead. In this paper, we propose DTRadar, a framework for expediting the decision tree packet lookup process. Its key idea is building an abstract One-Big-Tree(OBT) for multiple decision trees by establishing the middle data structure. DTRadar considers each decision tree as a splittable tree and organizes these subtrees by intermediate data structures. Extensive experiments show that DTRadar benefits existing decision tree-based solutions in classification time by 61.60%, and the memory footprint only increased by 4.21% on average. Jiashuo Yu, Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001 |
ISCC | 3 |
| 2022 | SketchGuide: Reconfiguring Sketch-based Measurement on Programmable SwitchesabstractSketches enable efficient and fine-grained network measurement results with configurable resource-performance trade-offs. While sketch configurations are guided by theories, the current theoretical guidelines are either impractical or deficient for sketch configurations on emerging programmable switches. To better configure sketches on programmable switches, we (1) systematically analyze the limitations of sketch configuration guidelines on programmable hardware switches (i.e., unguided parameters, accuracy profiles, and resource budgets); (2) propose a generic and practical framework called SketchGuide to automate efficient sketch configurations on programmable switches; (3) implement SketchGuide on a Barefoot Tofino switch and compare SketchGuide to the state-of-the-art sketches by conducting extensive experiments. Our evaluations demonstrate that SketchGuide can automatically configure unguided parameters given resource budgets. SketchGuide reduces the hardware resource footprint by 52.92%-99.28% compared with current guidelines without impacting fidelity. Zhengyan Zhou, Jingwen Lv, Lingfei Cheng, Xiang Chen 0017, Tianzhu Zhang 0002, Qun Huang 0001, Jiayu Luo, Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001 |
ICNP | 8 |
| 2022 | Libra: A Stateful Layer-4 Load Balancer with Fair Load DistributionabstractLayer-4 (L4) load balancers (LBs) are essential for data centers, dispatching incoming connections among thousands of servers. There are two critical requirements for L4 LBs: i) load balancing fairness, i.e., the ability to assign the load to servers in proportion to their capacity; ii) per-connection consistency (PCC), i.e., all the packets belonging to the same connection should be forwarded to the same server. Howbeit, existing LBs at large sacrifice load balancing fairness to mitigate PCC violations, which cannot satisfy both requirements in the meantime. In this paper, we present Libra, a stateful L4 LB that supports fair load distribution, PCC, memory efficiency, and resilience to resource depletion attacks. Libra makes load balancing decisions resorting to the proposed Weighted M-Least-Connection First (WMLCF) mechanism considering the real-time load and available processing capacity of servers, hence enabling load balancing fairness. We prototype Libra in a programmable software switch—BMv2 using P4 language and conduct extensive flow-level simulation to evaluate the performance. The evaluation indicates that Libra significantly improves load balancing fairness (over 95%), fully ensures PCC, and reduces the average flow completion time by 17.27–42.55% compared to existing mechanisms. Xingong Guo, Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001 |
IPCCC | 2 |
| 2022 | FROD: An Efficient Framework for Optimizing Decision Trees in Packet ClassificationabstractTo perform efficient packet classification, decision tree-based methods conduct decision trees via hand-tuned heuristics. Then the performance testing and optimization are executed to ensure an excellent searching speed and space overhead. Specifically, when the performance is below expectation, existing solutions attempt to optimize the algorithms, such as conducting more sophisticated heuristics. However, reconstruction or adjustment for algorithms produces an intolerable time overhead due to the long optimization period, caused by uncertain performance benefits and high pre-processing time. In this paper, we propose FROD, an efficient framework for optimizing the decision trees directly in packet classification. FROD raises a meticulous evaluation to accurately appraise decision trees constructed by different heuristics. It then seeks out the bottleneck components via a lightweight heuristic. After that, FROD searches the optimal division for inferior components considering structural constraints and characteristics of traffic distribution. Evaluation on ClassBench shows that FROD benefits existing decision tree-based solutions in classification time by 41% and memory footprint by 19% on average, and reduces classification time by up to 64%. Longlong Zhu, Jiashuo Yu, Jiayi Cai, Jinfeng Pan, Zhigao Li, Zhengyan Zhou, Dong Zhang 0010, Chunming Wu 0001 |
IWQoS | 1 |
| 2022 | Optimization Based Adaptive Cruise Control and Energy Management Strategy for Connected and Automated FCHEVabstractWith the development of vehicle electrification, automation and connectivity, collaborative optimization among the traffic throughput, driving comfort, fuel economy and driving safety targets is still a huge challenging barrier for a connected and automated fuel cell/battery hybrid electric vehicle. Hence, this paper proposes an optimal car-following energy management strategy (EMS) that combines energy management and adaptive cruise control considering the above targets. Specifically, based on vehicle-to-vehicle and vehicle-to-infrastructure information, an optimal following distance algorithm is developed to obtain the optimal following distance considering driving safety, driving comfort and traffic throughput. Then, based on the established vehicle longitudinal dynamics model, an adaptive cruise controller using back-stepping technique is designed to accurately track optimal following distance. Meantime, combining the obtained controller, optimal EMS based on equivalent consumption minimization strategy is proposed to coordinate the output power of fuel cell and battery to improve fuel economy. The simulations of short and long-term driving cycles indicate that the proposed method can reduce hydrogen consumption by 12.12%, jerk by 61.21%, and keep the desired following distance tracking error within 0.5m. Longlong Zhu, Fazhan Tao, Zhumu Fu, Nan Wang 0018, Baofeng Ji 0004, Yongsheng Dong 0004 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2020 | Energy Management Strategy Using Equivalent Consumption Minimization Strategy for Hybrid Electric VehiclesabstractIn this paper, an energy management strategy for electric vehicles equipped with fuel cell (FC), battery (BAT), and supercapacitor (SC) is considered, aiming at improving the whole performance under a framework of vehicle to network application. In detail, based on wavelet transform and equivalent consumption minimization strategy (ECMS), the demand power of vehicles is optimized to enhance the lifespan of fuel cell, fuel economy, and dynamic performance of electric vehicles. The wavelet transform is used to separate the high-frequency power in order to provide a peak power and recycle the braking energy. The equivalent consumption minimization strategy is used to distribute the low-frequency power to fuel cell and battery for minimizing the hydrogen consumption. Obtained results are studied using an advanced vehicle simulator, and its effectiveness of the strategy is confirmed, which provides a fundamental control method for the IOV application. Fazhan Tao, Longlong Zhu, Pengju Si, Zhumu Fu |
Secur. Commun. Networks | 2 |