Linying Zheng

dblp:359/6302 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
12since 2021 · last 2026
0009-0007-6346-8275ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 11 · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Turbolearn: Harnessing Accurate and Line-Rate Deep Learning on Programmable Switches
Zhifan Jiang, Longlong Zhu, Jiashuo Yu, Linying Zheng, Chunming Wu 0001, Xiang Chen 0017
APNet4
2026 TurboLearn: Harnessing Accurate and Line-Rate Deep Learning on Programmable Switches
abstract
The intelligent data plane (IDP) embeds deep learning (DL) models on switches for line-rate traffic analysis, but hardware constraints often force simplified models, reducing accuracy, while complex models like Transformers remain undeployable. We present TurboLearn, which achieves high accuracy and line-rate performance by co-designing inference across the switch ASIC and switch OS on the same switch. TurboLearn uses three key techniques: (1) a hardware fast path for lightweight classification and a software normal path for complex models, (2) confidence-based selective inference that escalates only low-confidence packets, and (3) confidence-calibrated knowledge distillation, where the normal path teaches the fast path. On Intel Tofino2 switches across three real-world traffic tasks, TurboLearn supports models that existing IDPs cannot deploy, improves macro-F1 by up to 31.31%, and keeps over 90% of traffic on the fast path with zero throughput loss.
Zhifan Jiang, Longlong Zhu, Jiashuo Yu, Linying Zheng, Chunming Wu 0001, Xiang Chen 0017
APNet4
2026 MonPlan: Taming Network Measurement with Accurate and Resource-Efficient Sketch-INT Co-Design
Xiang Chen 0017, Linying Zheng, Longlong Zhu, Zedi Chen, Qing Shu, Jialu Tian, Siqi Dong, Qun Huang 0001, Jianshan Zhang, Xuan Liu 0006, Haifeng Zhou, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001
INFOCOM2
2026 LTD: Low-Overhead Topology Discovery using Programmable Data Planes
Dezhang Kong, Minghao Li 0012, Shi Lin, Zhenhua Xu 0004, Longlong Zhu, Linying Zheng, Xiang Chen 0017, Changting Lin, Xuan Liu 0006, Dong Zhang 0010, Chunming Wu 0001
INFOCOM6
2026 SketchPipe: Toward Accurate Sketch-based Network Measurement on Multi-Pipeline Switches with Splitless Sketch Placement
Xiang Chen 0017, Longlong Zhu, Linying Zheng, Hongyang Du 0001, Dong Zhang 0010, Jianshan Zhang, Xuan Liu 0006, Qun Huang 0001, Dusit Niyato, Haifeng Zhou, Chunming Wu 0001, Hongyan Liu 0001, Kui Ren 0001
NSDI3
2026 Proteus: Towards Accurate and Low-overhead In-Network Malicious Traffic Detection
abstract
Network intrusion detection systems (NIDS) are essential for web security by identifying and dropping malicious traffic. Existing in-network NIDS leverage the Tbps-level packet processing capability of programmable switches to achieve high-speed flow classification. They translate complex trained machine learning models to decision trees (DTs), where DTs are deployed on programmable switches via single-DT or multiple-DT deployment. However, they face a fundamental trade-off: single-DT deployment suffers from low classification accuracy due to over-pruning of trees, while multiple-DT deployment suffers from high overhead due to deploying multiple tree replicas. In this paper, we propose Proteus, an in-network malicious traffic detection system that achieves both high classification accuracy and low overhead. Its key idea is to split the original DT into critical and normal sub-trees, where these sub-trees have different impacts on overall accuracy. More precisely, Proteus first splits a DT into one critical and several normal sub-trees for adapting to the accuracy requirement and switch resource budgets. Second, it minimizes coordination overhead between sub-trees while ensuring full flow coverage via mixed-integer linear programming. Third, it dynamically reallocates or migrates sub-trees to adapt to changing resources by monitoring both classification accuracy and switch resource changes. Testbed experiments with 12.8 Tbps programmable switches show that Proteus improves classification accuracy, reduces switch resource consumption, and reduces classification latency.
Longlong Zhu, Linying Zheng, Qing Shu, Zedi Chen, Jiashuo Yu, Shaopeng Zhou, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001, Xiang Chen 0017
WWW2
2025 Carrera: Enabling High-Performance eBPF-based Sketches in Network Measurement
abstract
To achieve dynamic network measurement, trends build sketches on eBPF to avoid service interruptions. However, existing eBPF-based sketches suffer from high CPU consumption, leading to poor throughput and high latency and making them hard to measure high-speed traffic. Optimizing their performance requires users to refactor codes based on each sketch’s characteristics on eBPF, which is highly complex and time-consuming.In this paper, we argue that users should write sketches without concerning low-level eBPF performance optimizations, with the deployment automatically activating cross-sketch performance optimizations. We present Carrera, a library that offers domain-specific optimizations for eBPF-based sketches. Our contributions are (1) systematically analyzing the performance bottlenecks of eBPF-based sketches through microbenchmarks, (2) identifying practical optimizations, including hardware offloading, SIMD-accelerated hashing, traffic-aware flow index caching, prefetched randomization, and active data collection, to address the identified bottlenecks in eBPF-based sketches, (3) evaluating these optimizations with state-of-the-art sketches and demonstrating that Carrera improves throughput by up to 65% and reduces latency by up to 93% via testbed experiments.
Xiang Chen 0017, Xin Yao 0008, Longlong Zhu, Linying Zheng, Hongyan Liu 0001, Jianshan Zhang, Dong Zhang 0010, Xuan Liu 0006, Qun Huang 0001, Haifeng Zhou, Chunming Wu 0001
ICNP5
2025 TurboCache: Empowering Switch-Accelerated Key-Value Caches with Accurate and Fast Cache Updates
abstract
Recent key-value (KV) caches are offloaded to programmable switches to offer high query processing performance. However, they suffer from both low accuracy in hot key detection and high latency in cache updates due to the strict limitations on switch registers. We propose TurboCache, a switch-accelerated KV cache with accurate hot key detection and fast cache updates. Our key idea is to leverage the switch recirculation capability to build a novel data structure that caches hot KV pairs. With this hardware-compatible cache data structure, TurboCache designs efficient data plane algorithms that accurately detects new hot keys and quickly updates its cache entirely within switch ASIC pipelines. We have implemented TurboCache on a${64}\times {100}$Gbps Tofino switch. Testbed results indicate that TurboCache improves the hot key detection accuracy and decreases the cache update latency of existing KV caches by several orders of magnitude.
Xiang Chen 0017, Longlong Zhu, Linying Zheng, Lingfei Cheng, Jianshan Zhang, Xu Yang 0002, Dong Zhang 0010, Xuan Liu 0006, Xiaoming Lu, Xun Yi, Ibrahim Khalil 0001, Albert Y. Zomaya, Haifeng Zhou, Chunming Wu 0001
INFOCOM3
2025 Scaling Learning-based Packet Classification Hardware with NeuTree
Jiashuo Yu, Longlong Zhu, Linying Zheng, Dong Zhang 0010, Xiang Chen 0017
INFOCOM5
2024 DOT: Towards Fast Decision Tree Packet Classification by Optimizing Rule Partitions
abstract
Packet classification is a crucial component of modern networks. Existing decision tree-based algorithms alleviate the rule replication problem caused by overlapping rules in the ruleset via rule partitioning. They partition the ruleset into multiple subsets based on rule characteristics to reduce rule overlaps. However, existing algorithms fail to address the overlap between rules in the same set, seriously decreasing speed and memory performance. In this paper, we propose DOT, a framework for optimizing rule partitions before constructing decision trees. Its key idea is to migrate rules in subsets based on rule overlaps and the features of heuristics used to construct trees, as well as reorganize rules aided by tuples. DOT finds out the migrated rule candidates using rule dependency graphs and heuristic features, then transforms the rule migration problem into an integer linear programming problem and solves for the optimal migration strategy. Further, we employ a tuple-assisted approach to accelerate rule matching. Experiments show that DOT enhances existing decision tree-based algorithms, improving lookup speed by 1.69 ×, reducing average 24.85% memory consumption and 31.03% decision tree depth.
Longlong Zhu, Jiashuo Yu, Linying Zheng, Dong Zhang 0010, Chunming Wu 0001
LCN5
2024 TransTuple: Toward Fast Packet Classification via Adaptive Tuple Replacement
abstract
Open vSwitch (OVS) is a widely used software switch in virtualized environments and software-defined networks. OVS uses tuple space search (TSS) for packet classification in the datapath, allowing fast network rule updates, but the increasing number of rules poses a classification performance challenge. To address this, existing methods incorporate decision trees with TSS to form a hybrid structure, enhancing classification speed. However, decision trees tend to overfit the initial ruleset, becoming unbalanced after rule updates and leading to a sharp decline in classification performance. In this paper, we propose TransTuple, a framework to optimize hybrid structures for fast packet classification under rule updates. The core idea of TransTuple is to identify bottleneck branches in decision trees that degrade performance and to replace them with lightweight tuples, providing better throughput under rule updates. These tuples maintain rules using hash tables, enabling fast updating and packet matching on bottleneck branches. We use TransTuple to optimize three state-of-the-art hybrid structured methods, i.e., CutTSS, TabTree, and MBitTree, achieving up to a 3.1x improvement in classification speed during rule updates.
Jiashuo Yu, Longlong Zhu, Rongbang Wu, Linying Zheng, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001
SECON4
2023 MiCuts: Combing Bit-Based Cutting and Splitting for Efficient Packet Classification
abstract
Packet classification is a crucial component in computer networking. To achieve high throughput and low memory consumption, existing solutions apply different heuristics in each construction stage to build efficient decision trees. However, previous studies divide the tree construction process based on the scale of rule subsets which is indirect to the performance goal, leading to massive rule replication and high tree depth. In this paper, we propose MiCuts, a fine-grained framework for packet classification with both high speed and low memory footprint. Its key idea is directly utilizing rule replication and tree depth to divide the tree-building process into three stages, each with suitable optimization goals. First, it partitions rules and builds shallow semi-trees without rule replication via selecting effective bits. Second, it transforms the switching problem of heuristics into an ILP problem and aims to minimize memory consumption while ensuring high lookup speed. Third, it merges some nodes to eliminate memory explosion caused by splitting, where MiCuts combines splitting and linear search. Extensive experimental results on ClassBench show that MiCuts outperforms state-of-the-art approaches, improving lookup speed by 1.71× while reducing memory footprint by 74.4% on average.
Longlong Zhu, Jiashuo Yu, Linying Zheng, Jinfeng Pan, Zhengyan Zhou, Hanze Chen, Dong Zhang 0010, Xiang Chen 0010, Chunming Wu 0001
ICC4