Lu Tang 0004

dblp:23/3111-4 · DBLP profile ↗
← Back
24ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0003-2923-6247ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 16 · 5 first-author · 12 since 2021Systems, architecture and hardware · 6 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AegisPath: Privacy-Preserving Interdomain Data-Plane Verification with Versioned Verifiable Evidence
Mingjun Fang, Shuhao Zheng, Zonglun Li, Letian Zhu, Qingyu Song 0002, Lizhao You, Lu Tang 0004, Wanjian Feng, Fei Yuan 0014, Qiao Xiang, Xue (Steve) Liu, Jiwu Shu
APNet8
2026 ParaSync: Exploiting Fine-Grained Parallelism for Efficient File Synchronization
Lu Tang 0004, Huiba Li, Yue Yu 0001, Guangtao Xue, Jiwu Shu, Yiming Zhang 0003
FAST2
2026 SkySync: Accelerating File Synchronization with Collaborative Delta Generation
Huiba Li, Lu Tang 0004, Guangtao Xue, Jiwu Shu, Yiming Zhang 0003
FAST3
2026 REACT: Toward Real-Time, End-to-End, Adaptive Cross-Layer Restoration for IP-Over-Optical Networks
Siyong Huang, Mochun Long, Qingyu Song 0002, Lizhao You, Lu Tang 0004, Wanjian Feng, Fei Yuan 0001, Qiao Xiang, Jiwu Shu
IWQoS7
2026 RepLLM: Toward Automatically Reproducing Network Research Results
abstract
Result reproduction of computer networking research is challenging as the scarcity of open-source implementations and the complexity of heterogeneous system architectures. Even though Large Language Models have demonstrated potential in code generation, existing code generation frameworks often fail to address the long-context constraints and intricate logical dependencies, which are vital in reproducing network systems from academic papers. Thus, we introduce RepLLM, an end-to-end multi-agent framework designed to automate code reproduction from paper content. RepLLM features a collaborative architecture comprising four specialized agents—Content Parsing, Architecture Design, Code Generation, and Audit & Repair, which are coordinated through Shared Memory mechanism to ensure global context consistency. With the enhancement of Structured Chain-of-Thought LLM reasoning and a sandbox-isolated static-dynamic debugging methodology, our framework effectively resolves semantic discrepancies and runtime errors, thereby improving reliable reproductions. Extensive evaluations on representative papers in top conferences demonstrate that RepLLM outperforms state-of-the-art system-level LLM frameworks in generating compile-ready and logically correct systems. Our results show that, with the aid of RepLLM, we can reproduce 95% of the original benchmarks within approximately two hours while reducing token consumption by up to 10% compared with state-of-the-art baselines.
Yining Jiang, Yunxin Xu, Wenyun Xu, Yufan Zhu, Tangtang He, Letian Zhu, Qingyu Song 0002, Lizhao You, Lu Tang 0004, Wanjian Feng, Yuchao Zhang 0004, Linghe Kong, Qiao Xiang, Jiwu Shu
SIGCOMM12
2026 Maat: A fair Layer-4 load balancer with per-connection consistency
Ju Huang, Dongzhan Zhang, Lu Tang 0004
Comput. Networks4
2025 Maat: A Fair Layer-4 Load Balancer With Per-Connection Consistency
Ju Huang, Lu Tang 0004
APNet2
2025 Unmasking Vulnerabilities of HyperLogLog: Security via Parameter Extraction
Shishi Zhang, Lu Tang 0004
APNet3
2025 SieveSketch: A Fine-grained and Adaptive Sketch Framework for Accurate Frequency Estimation
abstract
Estimating item frequencies in data streams is a fundamental task that supports a wide range of applications. To improve accuracy, existing algorithms typically employ filters to handle cold (infrequent) and hot (frequent) items separately. However, their accuracy often degrades across different data streams due to fixed parameter settings. Once the filter reaches its capacity, it can no longer effectively distinguish target items, resulting in a significant drop in accuracy. To achieve higher accuracy and better adaptability to data streams, we propose SieveSketch, a novel framework for frequency estimation in data stream processing. Inspired by two observations of narrow cold-item frequency range and different sensitivity of items to hash collisions, SieveSketch proposes adaptive scaling to adjust the count range of each counter with few bits (e.g. 4 bits) to record massive cold items efficiently, and takes a frequency-oriented counting method to process items at a more fine-grained level to improve the accuracy. We theoretically analyze the error bound of SieveSketch. We conduct extensive experiments on real-world and synthetic datasets, and the results show that, compared to the state-of-the-art, SieveSketch reduces the estimation error by up to 222.1 times.
Shishi Zhang, Lu Tang 0004
Proc. ACM Manag. Data3
2025 Toward Distributed Write-Back Caching in Programmable Switches
abstract
Skewed write-intensive key-value storage workloads are increasingly observed in modern data centers, yet they also incur server overloads due to load imbalance. Programmable switches provide viable solutions for realizing load-balanced caching on the I/O path, and hence implementing write-back caching in programmable switches is a natural approach to absorb frequent writes and improve write performance. However, enabling in-switch write-back caching is challenged by not only the strict programming rules and limited stateful memory of programmable switches, but also the need for reliable protection against data loss due to switch failures. We first propose FarReach, a new caching framework that supports fast, available, and reliable in-switch write-back caching. FarReach carefully co-designs both the control and data planes for cache management in programmable switches, so as to achieve high data-plane performance with lightweight control-plane management. We further extend FarReach into DistReach, which reduces the reliability maintenance overhead via distributed switch deployment. Our experimental results on a Tofino-switch testbed show that FarReach achieves a throughput gain of up to$6.6\times $over a state-of-the-art in-switch caching approach under skewed write-intensive workloads. Also, DistReach reduces the crash recovery time of FarReach by 77.4%.
Siyuan Sheng, Jiazhen Cai, Qun Huang 0001, Lu Tang 0004, Patrick P. C. Lee
IEEE Trans. Netw.4
2024 Reducing Write Tail Latency of Distributed Key-Value Stores Using In-Network Chasing
abstract
Multiple systems have explored how to use programmable switch ASICs to improve the performance of distributed systems. However, they focus on accelerating read operations and perform poorly under write-intensive workloads. In this paper, we present Gecko, a system that accelerates write operations in distributed key-value store systems using switch ASICs. The core idea of Gecko is to offload the client-side chasing mechanism, a technique deployed by production storage networks, to the programmable switch to simultaneously reduce the perceived and actual write-tail latency in distributed key-value stores. The perceived latency is the interval between the user sending a write request and the user receiving the write success, and there may be replicas that have not yet completed the write, but the actual latency requires all replicas to be successfully written. Gecko not only reduces the perceived write tail latency by deploying the chasing mechanism, but the actual write tail latency is also reduced by utilizing the capabilities of programmable switches. Specifically, Gecko’s in-network chasing design caches a write request at the switch data plane and reports success to the client when only m out of n (usually set to 2 and 3 in production networks, respectively) replicas have been successfully written to the server, and retries the cached write request if the remaining n − m replicas are not successfully written. In addition to caching the write request, Gecko also introduces novel designs to fully implement the chasing controller and a corresponding timer controller in the switch data plane, minimizing the interaction overhead between the switch control and data plane. Extensive experiments on a testbed of Barefoot Tofino switch and commodity servers show that Gecko not only substantially reduces the write tail latency by more than 2.16x caused by transient glitches at servers, but also maintains the same level of reliability as the classic three-replica write operation in distributed key-value stores.
Jinghui Jiang, Xiwen Fan, Zhenpei Huang, Kairui Zhou, Qiao Xiang, Lu Tang 0004, Qiang Li 0045, Jiwu Shu
IWQoS6
2024 RVCC: Congestion Control to Reduce Victim Flows in Data Center Networks
Lang Dai, Lu Tang 0004
NPC (2)2
2023 EasyQuantile: Efficient Quantile Tracking in the Data Plane
abstract
Quantile tracking is an essential component of network measurement, where the tracked quantiles of the key performance metrics allow operators to better understand network performance. Given the high network speed and huge volume of traffic, the line-rate packet-processing performance and network visibility of programmable switches make it a trend to track quantiles in the programmable data plane. However, due to the rigorous resource constraints of programmable switches, quantile tracking is required to be both memory and computation efficient to be deployed in the data plane. In this paper, we present EasyQuantile, an efficient quantile tracking approach that has small constant memory usage and involves only hardware-friendly computations. EasyQuantile adopts an adjustable incremental update approach and calculates a pre-specified quantile with high accuracy entirely in the data plane. We implement EasyQuantile on Intel Tofino switches with small resource usage. Trace-driven experiments show that EasyQuantile achieves higher accuracy and lower complexities compared with state-of-the-art approaches.
Bo Wang 0081, Rongqiang Chen, Lu Tang 0004
APNet3
2023 Accelerating SAT Solving Using Switching ASICs
abstract
People have been leveraging the capabilities of programmable switches, which are programmable in the data plane and process packets at the line rate, to improve the performance of distributed systems. However, few have explored whether programmable switches can speed up problem-solving. In this paper, we select the SAT problem, one of the most fundamental problems in computer science, as a case study to first explore the feasibility and benefits of this line of research. Our intuition is that by exploiting the parallel lookup capability of programmable switches, we can substantially speed up the process of checking whether an assignment is a solution to a SAT problem. Consequently, we design conflict tables using TCAM to quickly check assignment satisfiability. Building on the conflict tables, we propose two SAT solvers, P4-DPLL and Antler. Specifically, P4-DPLL is based on the classical DPLL algorithm and implements a stack data structure using registers and SRAM to efficiently make variable search decisions in the data plane, while Antler utilizes the mirroring capability of the programmable switch to further accelerate the SAT solving, which improves the non-parallel depth-first search-based P4-DPLL into a parallel breadth-first search-based SAT solver. We implement the prototypes of P4-DPLL and Antler on the Tofino switch, and evaluate their performance extensively. Results show that P4-DPLL and Antler improve the solving time by 2x and 169x on 90% quantile of test cases, compared to a CPU-based DPLL implementation. Besides, compared with the most popular SAT solvers, MathSAT and Z3, Antler improves the solving time by 42x and 35x, respectively.
Zhenpei Huang, Xiwen Fan, Jinghui Jiang, Mingyuan Song, Lu Tang 0004, Qiao Xiang, Jiwu Shu
ICPADS5
2023 Poster: P4-DPLL: Accelerating SAT Solving Using Switching ASICs
abstract
People have been leveraging the capabilities of programmable switches, which are programmable in the data plane and process packets at the line rate, to improve the performance of distributed systems. However, few have explored whether programmable switches can speed up problem-solving. In this demonstration, we take a first step to explore the feasibility and benefits of this line of research. Specifically, we select the SAT problem, one of the most fundamental problems in computer science, as a case study. Our intuition is that by exploiting the parallel lookup capability of programmable switches, we can substantially speed up the process of checking whether an assignment is a solution to a SAT problem. In particular, we base on the classical DPLL algorithm and design P4-DPLL [5], which consists of (1) match action tables using TCAM to quickly check assignment satisfiability, and (2) a stack data structure using register and SRAM to efficiently make variable search decisions in the data plane. We implement a prototype of P4-DPLL and evaluate its performance extensively. Results show that P4-DPLL improves the solving time by 101x speedup on 90% quantile of all test cases, compared with a CPU-based DPLL implementation.
Jinghui Jiang, Zhenpei Huang, Qiao Xiang, Lu Tang 0004, Jiwu Shu
SIGCOMM4
2023 FarReach: Write-back Caching in Programmable Switches
Siyuan Sheng, Huancheng Puyang, Qun Huang 0001, Lu Tang 0004, Patrick P. C. Lee
USENIX ATC4
2023 MVPipe: Enabling Lightweight Updates and Fast Convergence in Hierarchical Heavy Hitter Detection
abstract
Finding hierarchical heavy hitters (HHHs) (i.e., hierarchical aggregates with exceptionally huge amounts of traffic) is critical to network management, yet it is often challenged by the requirements of fast packet processing, real-time and accurate detection, as well as resource efficiency. Existing HHH detection schemes either incur expensive packet updates for multiple aggregation levels in the IP address hierarchy, or need to process sufficient packets to converge to the required detection accuracy. We present MVPipe, an invertible sketch that achieves both lightweight updates and fast convergence in HHH detection. MVPipe builds on the skewness property of IP traffic to process packets via a pipeline of majority voting executions, such that most packets can be updated for only one or few aggregation levels in the IP address hierarchy. We show how MVPipe can be feasibly deployed in P4-based programmable switches subject to limited switch resources. We also theoretically analyze the accuracy and coverage properties of MVPipe. Evaluation with real-world Internet traces shows that MVPipe achieves high accuracy, high throughput, and fast convergence compared to six state-of-the-art HHH detection schemes. It also incurs low resource overhead in the Tofino switch deployment.
Lu Tang 0004, Qun Huang 0001, Patrick P. C. Lee
IEEE/ACM Trans. Netw.1
2023 A High-Performance Invertible Sketch for Network-Wide Superspreader Detection
abstract
Superspreaders (i.e., hosts with numerous distinct connections) remain severe threats to production networks. How to accurately detect superspreaders in real-time at scale remains a non-trivial yet challenging issue. We present SpreadSketch, an invertible sketch data structure for network-wide superspreader detection with the theoretical guarantees on memory space, performance, and accuracy. SpreadSketch tracks candidate superspreaders and embeds estimated fan-outs in binary hash strings inside small and static memory space, such that multiple SpreadSketch instances can be readily merged to provide a network-wide measurement view for recovering superspreaders and their estimated fan-outs. We present formal theoretical analysis on SpreadSketch in terms of space and time complexities as well as error bounds. We further extend SpreadSketch with a fast and small data structure that filters out the packets of high-frequency connections from sketch processing, so as to improve the update performance of SpreadSketch while maintaining the accuracy guarantees. Trace-driven evaluation shows that SpreadSketch achieves higher accuracy and performance over state-of-the-art sketches and remains accurate in detecting real-world worms and DDoS attacks. Furthermore, we prototype SpreadSketch in P4 and show its feasible deployment in commodity hardware switches.
Lu Tang 0004, Qun Huang 0001, Patrick P. C. Lee
IEEE/ACM Trans. Netw.1
2021 Enabling Low-Redundancy Proactive Fault Tolerance for Stream Machine Learning via Erasure Coding
abstract
Machine learning for continuous data streams, or stream machine learning in short, is increasingly adopted in real-time big data applications. Fault tolerance is a critical requirement for stream machine learning applications in large-scale distributed deployment. However, existing reactive fault tolerance mechanisms, which trigger failure recovery upon the detection of failures, inevitably incur high recovery overhead and compromise the low-latency requirement of stream machine learning. We design StreamLEC, a stream machine learning system that leverages erasure coding to provide low-redundancy proactive fault tolerance for immediate failure recovery. StreamLEC supports general stream machine learning applications, and incorporates different techniques to mitigate erasure coding overhead. Evaluation on a local cluster and Amazon EC2 shows that StreamLEC achieves much higher throughput than both reactive fault tolerance and replication-based proactive fault tolerance, with negligible failure recovery overhead.
Zhinan Cheng, Lu Tang 0004, Qun Huang 0001, Patrick P. C. Lee
SRDS2
2020 SpreadSketch: Toward Invertible and Network-Wide Detection of Superspreaders
abstract
Superspreaders (i.e., hosts with numerous distinct connections) remain severe threats to production networks. How to accurately detect superspreaders in real-time at scale remains a non-trivial yet challenging issue. We present SpreadSketch, an invertible sketch data structure for network-wide superspreader detection with the theoretical guarantees on memory space, performance, and accuracy. SpreadSketch tracks candidate super-spreaders and embeds estimated fan-outs in binary hash strings inside small and static memory space, such that multiple SpreadSketch instances can be merged to provide a network-wide measurement view for recovering superspreaders and their estimated fan-outs. We present formal theoretical analysis on SpreadSketch in terms of space and time complexities as well as error bounds. Trace-driven evaluation shows that SpreadSketch achieves higher accuracy and performance over state-of-the-art sketches. Furthermore, we prototype SpreadSketch in P4 and show its feasible deployment in commodity hardware switches.
Lu Tang 0004, Qun Huang 0001, Patrick P. C. Lee
INFOCOM1
2020 A Fast and Compact Invertible Sketch for Network-Wide Heavy Flow Detection
abstract
Fast detection of heavy flows (e.g., heavy hitters and heavy changers) in massive network traffic is challenging due to the stringent requirements of fast packet processing and limited resource availability. Invertible sketches are summary data structures that can recover heavy flows with small memory footprints and bounded errors, yet existing invertible sketches incur high memory access overhead that leads to performance degradation. We present MV-Sketch, a fast and compact invertible sketch that supports heavy flow detection with small and static memory allocation. MV-Sketch tracks candidate heavy flows inside the sketch data structure via the idea of majority voting, such that it incurs small memory access overhead in both update and query operations, while achieving high detection accuracy. We present theoretical analysis on the memory usage, performance, and accuracy of MV-Sketch in both local and network-wide scenarios. We further show how MV-Sketch can be implemented and deployed on P4-based programmable switches subject to hardware deployment constraints. We conduct evaluation in both software and hardware environments. Trace-driven evaluation in software shows that MV-Sketch achieves higher accuracy than existing invertible sketches, with up to 3.38× throughput gain. We also show how to boost the performance of MV-Sketch with SIMD instructions. Furthermore, we evaluate MV-Sketch on a Barefoot Tofino switch and show how MV-Sketch achieves line-rate measurement with limited hardware resource overhead.
Lu Tang 0004, Qun Huang 0001, Patrick P. C. Lee
IEEE/ACM Trans. Netw.1
2019 MV-Sketch: A Fast and Compact Invertible Sketch for Heavy Flow Detection in Network Data Streams
abstract
Fast detection of heavy flows (e.g., heavy hitters and heavy changers) in massive network traffic is challenging due to the stringent requirements of fast packet processing and limited resource availability. Invertible sketches are summary data structures that can recover heavy flows with small memory footprints and bounded errors, yet existing invertible sketches incur high memory access overhead that leads to performance degradation. We present MV-Sketch, a fast and compact invertible sketch that supports heavy flow detection with small and static memory allocation. MV-Sketch tracks candidate heavy flows inside the sketch data structure via the idea of majority voting, such that it incurs small memory access overhead in both update and query operations, while achieving high detection accuracy. We present theoretical analysis on the memory usage, performance, and accuracy of MV-Sketch. Trace-driven evaluation shows that MVSketch achieves higher accuracy than existing invertible sketches, with up to 3.38× throughput gain. We also show how to boost the performance of MV-Sketch with SIMD instructions.
Lu Tang 0004, Qun Huang 0001, Patrick P. C. Lee
INFOCOM1
2017 SketchVisor: Robust Network Measurement for Software Packet Processing
abstract
Network measurement remains a missing piece in today's software packet processing platforms. Sketches provide a promising building block for filling this void by monitoring every packet with fixed-size memory and bounded errors. However, our analysis shows that existing sketch-based measurement solutions suffer from severe performance drops under high traffic load. Although sketches are efficiently designed, applying them in network measurement inevitably incurs heavy computational overhead.
Qun Huang 0001, Xin Jin 0008, Patrick P. C. Lee, Runhui Li, Lu Tang 0004, Yi-Chao Chen 0001, Gong Zhang 0001
SIGCOMM5
2015 A Priority-Based Scheduling Heuristic to Maximize Parallelism of Ready Tasks for DAG Applications
abstract
In practical Cloud/Grid computing systems, DAG scheduling may be faced with challenges arising from severe uncertainty about the underlying platform. For instance, it could be hard to have explicit information about task execution time and/or the availability of resources, both may change dynamically, in difficult to predict ways. In such a setting, the development of various kinds of just-in-time scheduling schemes, which aim at maximizing the parallelism of ready tasks of DAG, seems to be a promising approach to cope with the lack of environment information and achieve efficient DAG execution. Although many attempts have been tried to develop such just-in-time scheduling heuristics, most of them are based on DAG decomposition, which results in complicated and suboptimal solutions for general DAGs. This paper presents a priority-based heuristic, which is not only easy to apply to arbitrary DAGs, but also exhibits comparable or better performance than the existing solutions.
Wei Zheng 0002, Lu Tang 0004, Rizos Sakellariou
CCGRID2