VLDB 2026 Research / reviewers in the wild / expert
Peng He 0003
dblp:84/6016-3
· DBLP profile ↗
17ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-7472-9529ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 11 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gryphon: Scaling Hyperscale Multi-Tenant Gateways Beyond the Petabit-Era via DPU-Augmented Hierarchical Co-OffloadingabstractAt ByteDance, cloud gateway clusters orchestrate petabit-scale aggregate traffic. Traditional ASIC-only gateways fail to meet these escalating demands due to severe on-chip resource constraints and limited programmable flexibility, while pure software solutions or alternatives like disaggregated SmartNICs struggle to match terabit-scale line-rate throughput. To bridge this gap, we present Gryphon, a hyperscale cloud gateway built on a hybrid architecture that integrates DPUs directly into the switching ASIC's forwarding path. This design resolves the fundamental tension between capacity and speed, expanding table scale by up to 1000× and augmenting programmability, while sustaining 1.6 Tbps line-rate throughput at a cost of only ~8 μs in additional average latency. To manage this hardware heterogeneity, we introduce Hierarchical Co-Offloading (HLCO) in the data plane, achieving >99.9% fast path hit rate, while retaining software fallback for complex operations. In the control plane, we develop an abstraction layer (P4Bridge) that decouples hardware specifics from policy configuration. Gryphon has been operating at production scale for over a year, deployed on hundreds of nodes across multiple Availability Zones. We also share production measurements and operational experiences that serve as the first hyperscale-proven guidelines for next-generation DPU-augmented cloud gateways. Yuemeng Xu, Jiarui Guo, Mingwei Cui, Qiuheng Yin, Peng He 0003, Chenmin Sun, Yangyujia Wang, Daxiang Kang, Lirong Lai, Zhuochen Fan, Tong Yang 0003 |
SIGCOMM | 7 |
| 2026 | Traffic-Aware Design for Multi-Dimensional Lookup and Forwarding: From IP Routing to Packet ClassificationabstractPacket processing in modern routers and switches relies on rule matching, primarily performed by two core modules: IP prefix lookup for next-hop determination and packet classification for multi-field policy enforcement. However, most existing algorithms are rule-centric and assume uniform rule access, overlooking the highly skewed nature of real-world network traffic. Such mismatch between static rule organization and dynamic traffic behavior leads to inefficiency in both lookup and classification. To address this limitation, we propose a Traffic-aware Lookup and Forwarding (TLF) framework that leverages traffic measurement with lookup operations, enabling online adaptation to dynamic traffic patterns and frequent rule updates. Experimental results demonstrate that TLF provides 1.04×–3.37× speedups for lookup and forwarding over state-of-the-art algorithms, while substantially reducing both memory overhead and construction time. Furthermore, integrating TLF into Vector Packet Processor (VPP) and Open vSwitch (OVS) results in throughput improvements of 2.61× and 4.88×, respectively. Xinyi Zhang 0004, Qianrui Qiu, Peng He 0003, Guangxing Zhang, Luyiyun Li, Jianer Zhou, Kavé Salamatian, Gaogang Xie |
IEEE Trans. Netw. | 4 |
| 2025 | Byte vSwitch: A High-Performance Virtual Switch for Cloud NetworkingabstractVirtual switch is a fundamental component of cloud computing as it provides core networking functionalities for VMs and containers. Open vSwitch (OVS) is widely adopted in cloud environments due to its open-source nature, programmability, and rich set of features. At ByteDance, we initially adopted OVS in our public cloud, but as our cloud business grew, its generic design along with its complex code-base quickly became obstacles to improvements. Hence, we developed Byte vSwitch (BVS), a high-performance virtual switch that was specifically designed to address the performance, scalability, serviceability, and operational efficiency needs of our cloud services. More specifically, BVS adopts a simple architecture with an optimized hash table to maximize forwarding performance. In addition, we introduced several optimizations to improve BVS scalability, operability, and serviceability in cloud environments. Our evaluations show that BVS achieves up to 3.3× higher PPS and 25% lower latency compared to OVS. BVS has been deployed at scale across all regions of the ByteDance public cloud for over four years, and this paper presents our experience in designing, deploying, and operating BVS in production. Xin Wang 0261, Deguo Li, Lidong Jiang, Shubo Wen, Daxiang Kang, Engin Arslan, Peng He 0003, Xinyu Qian, Jianwen Pi, Xiaoning Ding, Hao Luo 0013 |
EuroSys | 8 |
| 2025 | NPC: Rethinking Dataplane through Network-aware Packet ClassificationabstractPacket classification is a critical component for accurately categorizing traffic in network systems. The efficiency of packet classification algorithms is primarily determined by two key factors: the classifier's data structure and the characteristics of the traffic being classified. While significant efforts have been made to optimize data structures, the potential of leveraging traffic characteristics remains underexplored. In this study, we revisit the network dataplane by integrating the network measurement module with the packet classification module. We propose an innovative Network-aware Packet Classification system (NPC) that utilizes sketch techniques to extract network traffic features. These features guide the construction of decision trees, enabling efficient and adaptable packet classification across diverse network environments. Experimental results demonstrate that the NPC achieves speedups ranging from 1.86× to 23.88× over state-of-the-art algorithms, while significantly reducing memory overhead and construction time, highlighting its practical value in real-world scenarios. Furthermore, integrating NPC into Open vSwitch (OVS) yields throughput improvements of 10.71× to 13.01× compared to the native OVS. Xinyi Zhang 0004, Qianrui Qiu, Peng He 0003, Xilai Liu, Kavé Salamatian, Changhua Pei, Gaogang Xie |
SIGCOMM | 4 |
| 2024 | Patronum: In-network Volumetric DDoS Detection and Mitigation with Programmable Switches
Penglai Cui, Jianer Zhou, Peng He 0003, Yanbiao Li 0001, Zhenyu Li 0001, Gaogang Xie |
ESORICS (4) | 6 |
| 2024 | Hoda: a High-performance Open vSwitch Dataplane with Multiple Specialized Data PathsabstractOpen vSwitch (OvS) has been widely used in cloud networks in view of its programmability and flexibility. However, we observe a huge performance drop when it loads practical cloud networking services (e.g., tunneling and firewalling). Our further analysis reveals that the root cause lies in the gap between the needs of supporting various selections of packet header fields and the one-size-fits-all data path in the vanilla OvS. Motivated by this, we design Hoda, a high-performance OvS dataplane with multiple specialized data paths. Specifically, Hoda constructs the specialized parser and microflow cache for each OpenFlow program so as to achieve lightweight parsing and caching. We also propose a configurable version of Hoda that introduces configuration knobs in the data path to ease specialization. The experiments with real-life OpenFlow rules show that Hoda achieves up to 1.7× speed up over the state-of-the-art OvS and 1.5× speed up over mSwitch. Hoda has also been deployed in a large cloud to serve various online services; the A/B test in the cloud reveals a 20% request process time reduction for Ngnix services. Peng He 0003, Zhenyu Li 0001, Junjie Wan, Xiongchun Duan, Yu Zhang 0209, Gaogang Xie |
EuroSys | 2 |
| 2023 | Dissecting Overheads of Service Mesh SidecarsabstractService meshes play a central role in the modern application ecosystem by providing an easy and flexible way to connect microservices of a distributed application. However, because of how they interpose on application traffic, they can substantially increase application latency and its resource consumption. We develop a tool called MeshInsight to help developers quantify the overhead of service meshes in deployment scenarios of interest and make informed trade-offs about their functionality vs. overhead. Using MeshInsight, we confirm that service meshes can have high overhead---up to 269% higher latency and up to 163% more virtual CPU cores for our benchmark applications---but the severity is intimately tied to how they are configured and the application workload. IPC (inter-process communication) and socket writes dominate when the service mesh operates as a TCP proxy, but protocol parsing dominates when it operates as an HTTP proxy. MeshInsight also enables us to study the end-to-end impact of optimizations to service meshes. We show that not all seemingly-promising optimizations lead to a notable overhead reduction in realistic settings. Xiangfeng Zhu, Guozhen She, Yu Zhang 0209, Yongsu Zhang, Xuan Kelvin Zou, Xiongchun Duan, Peng He 0003, Arvind Krishnamurthy, Matthew Lentz, Danyang Zhuo, Ratul Mahajan |
SoCC | 8 |
| 2022 | PextCuts: A High-performance Packet Classification Algorithm with Pext CPU InstructionabstractPacket classification is the most essential component for switches and firewalls to perform network functions. In Software Defined Network, the growing scale of traffic requires the packet classification algorithm to perform high-speed lookup. Even though a lot of algorithms are proposed, the lookup performance is still the bottleneck because of the inefficient and unscientific schemes to cut rules and split trees. In this paper, we propose a novel decision-tree-based algorithm PextCuts. First, to efficiently cut rules, PextCuts applies one pext CPU instruction to select discontiguous bits rather than contiguous bits. Second, to scientifically split trees, PextCuts applies the dynamic programming method to split each field into multiple sizes rather than large and small sizes. Compared to ten representative algorithms, PextCuts has the highest lookup speed with the minimal numbers of average memory accesses, maximal memory accesses, and tree height simultaneously. It also consumes the least memory cost and the shortest construction time. For the state-of-the-art algorithm ByteCuts, PextCuts achieves 2.1x lookup speed with only 57% memory cost and 10% construction time. In addition, we implement PextCuts in DPDK to perform packet classification with optional fields and achieve 3.0x lookup speed. Gaogang Xie, Peng He 0003 |
IWQoS | 3 |
| 2022 | NetSHa: In-Network Acceleration of LSH-Based Distributed SearchabstractLocality Sensitive Hashing (LSH) is widely adopted to index similar data in high-dimensional space for approximate nearest neighbor search. Demanding applications (e.g. web search) mean that LSH must exhibit low response times and high throughput. To achieve this, they tend to load balance between multiple machines. However, as the scale of concurrent queries and the volume of data grow, large numbers of index messages are required. Hence, the network is a key bottleneck. To address this gap, we propose NetSHa, which exploits the computational capacity of programmable switches. Specifically, we introduce a heuristic sort-reduce approach to drop potentially poor candidate answers while preserving search quality. Then, NetSHa aggregates good candidate answers from different index messages when transmitting them. Through this, it reduces the network communication cost. Furthermore, we introduce a best-effort replacement mechanism to improve its concurrency. We implement NetSHa on a Barefoot Tofino programmable switch and evaluate it using 7 real-world datasets. The experimental results show that NetSHa reduces the packet volume by$4\sim 10$times and improves the search efficiency by least 3× in comparison with typical LSH-based distributed search frameworks. Penghao Zhang, Zhenyu Li 0001, Penglai Cui, Ru Jia, Peng He 0003, Gareth Tyson, Gaogang Xie |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2021 | Accelerating LSH-based Distributed Search with In-network ComputationabstractLocality Sensitive Hashing (LSH) is widely adopted to index similar data in high-dimensional space for approximate nearest neighbor search. With the rapid increase of datasets, recent interests in LSH have moved to the implementation of distributed search systems with low response time and high throughput. However, as the scale of the concurrent queries and the volume of available data grow, large amounts of index messages still need to be transmitted to centralized servers for the candidate answer reducing and resorting. Hence, the network remains the bottleneck in distributed search systems.To address this gap, we turn our efforts to the network itself and propose NetSHa. NetSHa exploits the in-network computational capacity provided by programmable switches. Specially, NetSHa designs a sort-reduce approach to drop the potential poor candidate answers and aggregates the good candidate answers on programmable switches, while preserving the search quality. We implement NetSHa on Barefoot Tofino switches and evaluate it using 3 datasets (i.e., Random, Wiki and Image). The experimental results show that NetSHa reduces the packet volume by 10 times at most and improves the search efficiency by 3x at least, in comparison with typical LSH-based distributed search frameworks. Penghao Zhang, Zhenyu Li 0001, Peng He 0003, Gareth Tyson, Gaogang Xie |
INFOCOM | 4 |
| 2018 | Efficient Action Computation for Compositional SDN PoliciesabstractSoftware-defined networking envisions the support of multiple applications collaboratively operating on the same traffic. Policies of applications therefore require composition into a rule list that represents the union of application intents. In this context, ensuring the correctness and efficiency of composition for match fields as well as the associated actions is the fundamental requirement. Prior work however focuses only on the composition of match fields and assumes simple concatenation for action composition. We show in this paper that simple concatenation can result in incorrect behavior and inefficiency of packet processing. To address this issue, we formalize the action composition problem and propose two graph-based computation models to facilitate efficient composition of action lists. Our proposed approach has been integrated into the CoVisor code base and the evaluation results show its fitness for purpose. Zhenyu Li 0001, Gaogang Xie, Peng He 0003, Hongtao Guan, Laurent Mathy |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2018 | Partial Order Theory for Fast TCAM UpdatesabstractTernary content addressable memories (TCAMs) are frequently used for fast matching of packets against a given ruleset. While TCAMs can achieve fast matching, they are plagued by high update costs that can make them unusable in a high churn rate environment. We present, in this paper, a systematic and in-depth analysis of the TCAM update problem. We apply partial order theory to derive fundamental constraints on any rule ordering on TCAMs, which ensures correct checking against a given ruleset. This theoretical insight enables us to fully explore the TCAM update algorithms design space, to derive the optimal TCAM update algorithm (though it might not be suitable to be used in practice), and to obtain upper and lower bounds on the performance of practical update algorithms. Having lower bounds, we checked if the smallest update costs are compatible with the churn rate observed in practice, and we observed that this is not always the case. We therefore developed a heuristic based on ruleset splitting, with more than a single TCAM chip, that achieves significant update cost reductions (1.05~11.3×) compared with state-of-the-art techniques. Peng He 0003, Hongtao Guan, Kavé Salamatian, Gaogang Xie |
IEEE/ACM Trans. Netw. | 1 |
| 2017 | FlowConvertor: Enabling portability of SDN applicationsabstractSoftware-Defined Networking (SDN) provides network administrators opportunities to control network devices more simply and easily than in traditional networking. However, heterogeneity in switch hardware, especially in forwarding pipeline architecture, renders the task of network application developers and network administrators tedious, by hampering portability across switch models. In this paper, we propose FlowConvertor, an algorithm capable of converting rules from any forwarding pipeline to any other different forwarding pipeline, as long as both pipelines offer compatible operations. More precisely, FlowConvertor is an online algorithm that operates on flow updates issued to the origin pipeline and computes the corresponding updates for the target pipeline in real time. Performance evaluation shows that the latency introduced by FlowConvertor on the path between the SDN controller and the target switch is of the order of 1ms in most cases, and is thus acceptable for practical deployment. Gaogang Xie, Zhenyu Li 0001, Peng He 0003, Laurent Mathy |
INFOCOM | 4 |
| 2016 | Transparent flow migration for NFVabstractNFV together with SDN provides the flexibility for NFs in the way that they are deployed and managed. The flexibility enables dynamical scale in and scale out through migrating in-process flows among NFs. Due to stateful packet processing in NFs, flow migration has to guarantee loss-free and order-preserving for both flow states and packets. Existing frameworks closely coupled state transfer and packets migration, and thus fail to achieve safe and efficient migration with low overhead. This paper presents our design and implementation of a distributed flow migration framework, Transparent Flow Migration (TFM). TFM completely decouples the state transfer and packets migrations. The decoupling allows us to optimize the two processes separately and run them in parallel. TFM implements various optimizations through the TFM box, a shim layer providing transparent packet migration to NFs. Our evaluation shows that TFM guarantees loss-free and order-preserving for both scale-in and scale-out flow migration, and outperforms existing approaches with 3× smaller migration time. Besides, TFM uses small overhead and has very limited impacts on throughput of live TCP flows. Yang Wang 0147, Gaogang Xie, Zhenyu Li 0001, Peng He 0003, Kavé Salamatian |
ICNP | 4 |
| 2014 | Meta-algorithms for Software-Based Packet ClassificationabstractWe observe that a same rule set can induce very different memory requirement, as well as varying classification performance, when using various well known decision tree based packet classification algorithms. Worse, two similar rule sets, in terms of types and number of rules, can give rise to widely differing performance behaviour for a same classification algorithms. We identify the intrinsic characteristics of rule sets that yield such performance differences, allowing us to understand and predict the performance behaviour of a rule set for various modern packet classification algorithms. Indeed, from our observations, we are able to derive a memory consumption model and an offline algorithm capable of quickly identifying which packet classification is suited to a give rule set. By splitting a large rule set in several subsets and using different packet classification algorithms for different subsets, our Smart Split algorithm is shown to be capable of configuring a multi-component packet classification system that exhibits up to 11 times less memory consumption, as well as up to about 4× faster classification speed, than the state-of-art work [20] for large rule sets. Our Auto PC framework obtains further performance gain by avoiding splitting large rule sets if the memory size of the built decision tree is shown by the memory consumption model to be small. Peng He 0003, Gaogang Xie, Kavé Salamatian, Laurent Mathy |
ICNP | 1 |
| 2013 | Toward predictable performance in decision tree based packet classification algorithmsabstractPacket classification has been studied extensively in the past decade. While many efficient algorithms have been proposed, the lack of deterministic performance has hindered the adoption and deployment of these algorithms: the expensive and power-hungry TCAM is still the de facto standard solution for packet classification. In this work, in contrast to proposing yet another new packet classification algorithm, we present the first steps to understand this unpredictability in performance for the existing algorithms. We focus on decision-tree based algorithms in this paper. In order to achieve the predictability, we firstly revisit the classical and many state-of-art packet classification algorithms. Through a detailed analysis, we conclude that two features of ruleset usually dominate the performance results: 1) the uniformity of the range distribution in different dimensions of the rules; 2) the existence and the number of “orthogonal structure” and wildcard rules in the ruleset. We conduct experiments to show the correctness of these observations, and discribe some potential applications for those results. Our work provides some insight to make the packet classification algorithms a credible alternative to the TCAM-only solutions. Peng He 0003, Hongtao Guan, Laurent Mathy, Kavé Salamatian, Gaogang Xie |
LANMAN | 1 |
| 2012 | Evaluating and Optimizing IP Lookup on Many Core ProcessorsabstractIn recent years, there has been a growing interest in multi/many core processors as a target architecture for high performance software router. This is a clear difference from the previous trend to use dedicated network processors and hardware components. Because of its key position in routers, hardware IP lookup implementation has been intensively studied with TCAM and FPGA based architecture. However, increasing interest in software implementation has also been observed. In this paper, we evaluate the performance of software only IP lookup on a many core chip, the TILEPro64 processor. For this purpose we have implemented two widely used IP lookup algorithms, DIR-24-8-BASIC and Tree Bitmap. We evaluate the performance of these two algorithms over the TILEPro64 processor with both synthetic and real-world traces. After a detailed analysis, we propose a hybrid scheme which provides high lookup speed and low worst case update overhead. Our work shows how to exploit the architectural features of TILEPro64 to improve the performance, including many optimization in both single-core and parallelism aspects. Experiment results show by using only 18 cores, we can achieve a lookup throughput of 60Mpps (almost 40Gbps) with low power consumption, which demonstrates great performance potentials in many core processor. Peng He 0003, Hongtao Guan, Gaogang Xie, Kavé Salamatian |
ICCCN | 1 |