VLDB 2026 Research / reviewers in the wild / expert
Penglai Cui
dblp:294/8803
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Patronum: In-network Volumetric DDoS Detection and Mitigation with Programmable Switches
Penglai Cui, Jianer Zhou, Peng He 0003, Yanbiao Li 0001, Zhenyu Li 0001, Gaogang Xie |
ESORICS (4) | 3 |
| 2024 | Zebra: Accelerating Distributed Sparse Deep Training With in-Network Gradient Aggregation for Hot ParametersabstractDistributed sparse deep learning has been widely used in many Internet-scale applications. Network communication is one of the major hurdles for training performance. In-network gradient aggregation on programmable switches is a promising solution for speeding up the performance. Nevertheless, existing in-network aggregation solutions are designed for the dense deep training, and fall short when used for the sparse training. To address this gap, we present Zebra based on our key observation on the extremely biased update frequency of parameters in distributed sparse deep training. Specifically, Zebra offloads only the aggregation for “hot” parameters that are updated frequently onto programmable switches. To enable this offloading and achieve high aggregation throughput, we propose solutions to address the challenges related to hot parameter identification, parameter orchestration and gradient aggregation as well as system reliability. We implemented Zebra on Intel Tofino switches and integrated it with PS-lite. Finally, we evaluate Zebra's performance through extensive experiments and show that it can speed up the gradient aggregation by$1.5 \sim 4 \times$and the end-to-end performance by$1.4 \sim 2.6 \times$. Penglai Cui, Zhenyu Li 0001, Ru Jia, Penghao Zhang, Mathy Lauren, Gaogang Xie |
ICNP | 2 |
| 2024 | VAKY: Scheduling In-network Aggregation for Distributed Deep Training AccelerationabstractDistributed machine learning (DML) has recently experienced widespread application. A major performance bottleneck is the costly communication for gradients synchronization. Recently, researchers have explored the use of programmable switches for in-network synchronous aggregation of gradients to mitigate the communication overhead. Nevertheless, the performance of in-network synchronous aggregation is significantly impacted by the stragglers. Unfortunately, the schedulers in existing DML systems are no longer effective in dealing with stragglers because of the ignorance of the aggregation progress that is offloaded from the parameter servers to the programmable switches. To address this gap, this paper presents VAKY, an adaptive scheduler specifically designed for in-network aggregation. At the heart of VAKY is the variable K-block sync method, where the aggregators stop waiting for updates from more workers once having received updates from the fastest K workers for each block of gradients. We propose an efficient solution that can dynamically choose the optimal values of K during the training process, in order to minimize the expected training completion time. We have integrated VAKY into PyTorch, and our experiments show that compared to the state-of-the-art in-network aggregation systems, VAKY improves the aggregation throughput by up to $40 \%$ and reduces the training time by $25 \%$. Penglai Cui, Jianer Zhou, Qinghua Wu 0004, Zhaohua Wang, Zhenyu Li 0001 |
ICPADS | 1 |
| 2023 | Misconfiguration-Free Compositional SDN for Cloud NetworksabstractCloud computing provides a new paradigm to offer flexible IT infrastructures. In IaaS clouds, tenants deploy software-defined networking (SDN) policies to simplify network management and customize network behaviors. However, programming SDN networks is error-prone no matter using low-level APIs or high-level programming languages. Specifically, SDN policies may contain misconfigurations that do not break the pre-defined network invariants (e.g., black holes), but either degrade the deployment efficiency or mistakenly translate tenants intents. Prior studies for checking either traditional access control policies or network-wide invariants, are thus fail to detect these misconfigurations. To address this gap, this paper presents PMM, a misconfiguration checking tool for compositional SDN that works at the data plane of cloud networks. We first propose a new data structure, minimal interval set, to represent the match patterns of rulesets. This representation serves the basis for composition algebra construction and misconfiguration checking. We then propose the principles, algorithms and also optimisations for fast and accurate checking. We finally implement PMM in Covisor. Experiments with both real-world rulesets and synthetic rulesets show that PMM can detect misconfigurations of SDN policies in cloud networks within hundreds of milliseconds. Zhenyu Li 0001, Penghao Zhang, Penglai Cui, Kavé Salamatian, Gaogang Xie |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2022 | MD-Roofline: A Training Performance Analysis Model for Distributed Deep LearningabstractDue to the bulkiness and sophistication of the Distributed Deep Learning (DDL) systems, it leaves an enormous challenge for AI researchers and operation engineers to analyze, diagnose and locate the performance bottleneck during the training stage. Existing performance models and frameworks gain little insight on the performance reduction that a performance straggler induces. In this paper, we introduce MD-Roofline, a training performance analysis model, which extends the traditional rooftine model with communication dimension. The model considers the layer-wise attributes at application level, and a series of achievable peak performance metrics at hardware level. With the assistance of our MD-Roofline, the AI researchers and DDL operation engineers could locate the system bottleneck, which contains three dimensions: intra-GPU computation capacity, intra-GPU memory access bandwidth and inter-GPU communication bandwidth. We demonstrate that our performance analysis model provides great insights in bottleneck analysis when training 12 classic CNNs. Tianhao Miao, Qinghua Wu 0004, Penglai Cui, Zhenyu Li 0001, Gaogang Xie |
ISCC | 4 |
| 2022 | Enabling In-Network Floating-Point Arithmetic for Efficient Computation OffloadingabstractProgrammable switches are recently used for accelerating data-intensive distributed applications. Some computational tasks, traditionally performed on servers in data centers, are offloaded into the network on programmable switches. These tasks may require the support of on-the-fly floating-point operations. Unfortunately, programmable switches are restricted to simple integer arithmetic operations. Existing systems circumvent this restriction by converting floats to integers or relying on local CPUs of switches, incurring extra processing delayed and accuracy loss. To address this gap, we propose NetFC, a table-lookup method to achieve on-the-fly in-network floating-point arithmetic operations nearly without accuracy loss. Specifically, NetFC utilizes logarithm projection and transformation to convert the original huge table enumerating all operands and results into several much smaller tables that can fit into the data plane of programmable switches. To cope with the table inflation problem on 32-bit floats, we also propose an approximation method that further breaks the large tables into smaller ones. In addition, NetFC leverages two optimizations to improve accuracy and reduce on-chip memory consumption. We use both synthetic and real-life datasets to evaluate NetFC. The experimental results show that the average accuracy of NetFC is above 99.9% with only 448KB memory consumption for 16-bit floats and 99.1% with 496KB memory consumption for 32-bit floats. Furthermore, we integrate NetFC into two distributed applications and two in-network telemetry systems to show its effectiveness in further improving the performance. Penglai Cui, Zhenyu Li 0001, Penghao Zhang, Tianhao Miao, Jianer Zhou, Hongtao Guan, Gaogang Xie |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | NetSHa: In-Network Acceleration of LSH-Based Distributed SearchabstractLocality Sensitive Hashing (LSH) is widely adopted to index similar data in high-dimensional space for approximate nearest neighbor search. Demanding applications (e.g. web search) mean that LSH must exhibit low response times and high throughput. To achieve this, they tend to load balance between multiple machines. However, as the scale of concurrent queries and the volume of data grow, large numbers of index messages are required. Hence, the network is a key bottleneck. To address this gap, we propose NetSHa, which exploits the computational capacity of programmable switches. Specifically, we introduce a heuristic sort-reduce approach to drop potentially poor candidate answers while preserving search quality. Then, NetSHa aggregates good candidate answers from different index messages when transmitting them. Through this, it reduces the network communication cost. Furthermore, we introduce a best-effort replacement mechanism to improve its concurrency. We implement NetSHa on a Barefoot Tofino programmable switch and evaluate it using 7 real-world datasets. The experimental results show that NetSHa reduces the packet volume by$4\sim 10$times and improves the search efficiency by least 3× in comparison with typical LSH-based distributed search frameworks. Penghao Zhang, Zhenyu Li 0001, Penglai Cui, Ru Jia, Peng He 0003, Gareth Tyson, Gaogang Xie |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | NetFC: Enabling Accurate Floating-point Arithmetic on Programmable SwitchesabstractProgrammable switches are recently used for accelerating data-intensive distributed applications. Some computational tasks, traditionally performed on servers in data centers, are offloaded to the network on programmable switches. These tasks may require the support of on-the-fly floatingpoint operations. Unfortunately, the computational capacity of programmable switches is limited to simple integer arithmetic operations. To address this issue, prior approaches either adopt a float-to-integer method or rely on local CPUs of switches, incurring accuracy loss and delayed processing.To this end, we propose NetFC, a table-lookup method to achieve on-the-fly in-network floating-point arithmetic operations nearly without accuracy loss. NetFC adopts a divide-and-conquer mechanism that converts the original huge table into several much smaller tables that are operated by the built-in integer operations. NetFC further leverages a scaling-factor mechanism for improving computational accuracy, and a prefix-based lossless table compression method to reduce memory consumption. We use both synthetic and real-life datasets to evaluate NetFC. The experimental results show that the average accuracy of NetFC is above 99.94% with only 448KB memory consumption. Furthermore, we integrate NetFC into Sonata [12] for detecting Slowloris attack, yielding significant decrease of detection delay. Penglai Cui, Zhenyu Li 0001, Jiaoren Wu, Shengzhuo Zhang, Xingwu Yang, Hongtao Guan, Gaogang Xie |
ICNP | 1 |