VLDB 2026 Research / reviewers in the wild / expert
Chao Pei
dblp:214/9364
· DBLP profile ↗
9ranked-venue papers
1as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | XFir: Accelerating New-Flow Setup on Host Servers of a Large Cloud NetworkabstractIn today's cloud networks, host servers widely deploy Data Processing Units (DPUs) as network accelerators under the "Sep-Path" paradigm. However, as server capabilities scale with increasing CPU cores and network bandwidth, the software slow path (executed on a DPU's CPU) has become a critical bottleneck for workloads with high new-flow rates. Meanwhile, new-flow setup logic on host servers must continuously evolve to meet diverse and changing customer demands, making flexibility a key requirement alongside performance. To address this gap, we present XFir, the first hardware-accelerated new-flow setup system for cloud host servers that delivers high CPS throughput while preserving sufficient flexibility. XFir leverages a next-generation DPU equipped with a Cloud Network co-Processor (CNP) to execute the host server's new-flow setup logic. XFir redesigns the host-server flow-setup datapath and table layout, optimizes LPM lookups, and introduces CPU-CNP collaboration mechanisms to further improve performance and reliability. Our evaluation shows that XFir achieves over 776K new-flow CPS on a single host server with 11.7μs slow-path latency. Compared to prior work (Fornax), XFir achieves 4.8x CPS and reduces latency by 69.2%. Moreover, XFir is cost-effective to deploy, requiring only a single DPU per host. Overall, XFir improves new-flow throughput while maintaining development flexibility at low financial cost. Shihan Lin, Shunqiao Jiang, Chao Pei, Jian Zhao 0006, Wenjun Wu 0001, Lijun Zhuang, Qingmin Liu, Heng Yu 0005, Yibo Huang 0005, Yifei Zhu 0001, Yunming Xiao, Ang Chen 0001, Linghe Kong, Congcong Miao |
SIGCOMM | 5 |
| 2026 | DistDPU: A Disaggregated DPU Architecture for High-Performance and Cost-Efficient AI CloudsabstractAI training and inference are driving cloud networks toward terabit-per-second (Tbps) bandwidth per server, challenging the scalability and efficiency of today's cloud network architectures. A prevalent design scales bandwidth by stacking monolithic Data Processing Units (DPUs), but this approach tightly couples control and data plane resources, leading to excessive cost, power consumption, and operational complexity. We identify a fundamental control-data plane divergence in AI clouds: while data plane bandwidth demand grows rapidly, control plane demand remains largely flat due to the dominance of elephant flows. As a result, monolithic DPUs become systematically over-provisioned when used as bandwidth scaling primitives. Lizhou Gao, Yuanyi Zhu, Chao Pei, Chuhao Chen 0001, Zijian Li 0003, Jian Zhao 0006, Dongbo Gu, Hongchen Ren, Jiyuan Chen, Yunpeng Guan, Jianye Yuan, Yibo Huang 0005, Yang Xu 0010 |
SIGCOMM | 6 |
| 2026 | CubeTrace: Microscopic Network Tracing for Heterogeneous Cloud Gateways
Yunming Xiao, Yinchao Yang, Jiaqi Zheng 0001, Xuqian Li, Dongbo Gu, Jun Zhang 0014, Miantao Wan, Chao Pei, Chen Tian 0001, Mingwei Xu 0001, Ang Chen 0001, Congcong Miao |
SIGCOMM | 8 |
| 2026 | Dorado: Scaling SmartNIC Session Tables on Commodity DDRs
Heng Yu 0005, Jiajun Liang, Baozeng Zhang, Guozhi Lin, Xinyi Zhang 0004, Jian Zhao 0006, Ziyue Zhai, Chao Pei, Jilong Wang 0001, Gaogang Xie, Ang Chen 0001, Congcong Miao |
SIGCOMM | 11 |
| 2026 | Pegasus: A Data Center Network for Bare-Metal AI CloudabstractToday, AI cloud is key to serving diverse users with AI services, where cloud networking forms the basis. In this paper, we share our experience in designing, deploying, and operating Pegasus, a data center network tailored for the AI cloud, along with operational lessons learned from its deployment. The key designs of Pegasus include: 1) Network virtualization: a DPU-RNIC decoupled collaborative hardware architecture to enable a single DPU to virtualize multiple RNICs while reducing the power consumption. We design two-level flow tables on both DPU and RNICs to support underlay-overlay IP address translation and ensure isolation. For DPU-RNIC communication, we introduce a per-RNIC communication state machine to reduce communication overhead. 2) Network transport: customized and transparent transport offloading in the RNIC for low-latency and high-throughput communication performance for various AI workloads. We carefully offload per-packet load balancing and credit-based congestion control in RNICs, optimizing reorder delay and eliminating the impacts of hardware jitter. Pegasus has been deployed in production for over two years, currently covering 8K GPUs and supporting a wide range of tenants' AI applications. Xianneng Zou, Zhaoxun Zhou, Xingda Wei, Zhaohe Chen, Yinben Xia, Lizhou Gao, Jiajun Liang, Chunxu Zhao, Jiewei Yang, Yunpeng Guan, Dongbo Gu, Chao Pei, Zekun He, Yachen Wang |
SIGCOMM | 26 |
| 2025 | P4-IDet: A Programmable Switch-Based Framework for Real-Time and High-Accuracy Traffic Anomaly Detection in ICPSsabstractThe rise of Industry 4.0 exposes traditionally isolated Industrial Cyber-Physical Systems (ICPSs) to increasing network attacks, posing serious security threats and potential damage. Traffic anomaly detection is essential for identifying such attacks. Nevertheless, existing work faces a dilemma between high accuracy and real-time performance. In this paper, we resolve this dilemma through P4-IDet, a novel traffic anomaly detection framework based on programmable switches, achieving both high accuracy and real-time performance. P4-IDet first deploys a low-complexity detector in the data plane to stamp timestamps, extract traffic features, and perform line-rate preliminary detection. Only suspicious packets and their features are uploaded to a server for fine-grained analysis by a high-accuracy machine learning model. To further reduce the upload and accelerate detection, a Bayesian optimizer adaptively tunes detection rules based on differences between detection results of the switch and the server. Moreover, P4-IDet can be integrated with existing detection models to enhance accuracy and real-time performance. Finally, we implement the prototype on a Barefoot Tofino 2.0 switch using the P4 language and an x86 server, and validate it on a large-scale ICPS platform with real-world industrial systems. Experiments show 5.6–41.1% accuracy gains, a 28.93% reduction in machine learning model workload, and 8.31–25.90% improvements in real-time performance. Jiayu Luo, Zhengyan Zhou, Qiaoxiong Tang, Ruohan Chen, Xiang Chen 0017, Chao Pei, Qiang Yang 0004, Wenhai Wang, Haifeng Zhou |
IECON | 8 |
| 2025 | Fornax: A Hardware-Centric Session Management in Large Public Cloud NetworkabstractSmartNIC is increasingly utilized to accelerate cloud network components. The effectiveness and correctness of hardware acceleration heavily rely on its management mechanism. Unfortunately, traditional management mechanisms adopt software-centric architecture, which treats flow as the basic management unit and completely relies on one-way commands to manage the flow table, making it challenging to support various cloud network scenarios while managing extremely large tables. In this paper, we advocate for a radical new mechanism to shift the management paradigm from software-centric architecture to hardware-centric architecture, which adopts session as the basic management unit and designs two-way protocols to facilitate the management process. We propose and implement a first-of-its-kind system, called Fornax, a novel management architecture for large public cloud networks. At the core of Fornax is leveraging a session-empowered hardware engine to provide various management capabilities. Besides, Fornax utilizes a light-weight software manager to enhance system scalability, and hardware-driven management protocols to improve resource efficiency. Our testbed evaluations demonstrate that Fornax can reduce the software storage usage by 80% and CPU usage by 77% with little hardware resource overhead. Our large-scale production results show that Fornax can manage up to 16M session entries while significantly reducing the resource overhead by over 79%. Heng Yu 0005, Jian Zhao 0006, Guozhi Lin, Baozeng Zhang, Yunpeng Guan, Jiajun Liang, Chao Pei, Yachen Wang, Xin Jin 0008, Jilong Wang 0001, Congcong Miao |
SIGCOMM | 12 |
| 2017 | Integrated metric learning with adaptive constraints for person re-identificationabstractPerson re-identification is an important technique to search a probe person against a set of gallery persons and metric learning methods have shown their effectiveness in matching person images. In this paper, an Integrated Metric Learning with Adaptive Constraints (IMLAC) method is proposed to promote the performance for person re-identification. In the method, the difference and commonness of an image pair are combined to define a novel integrated metric. Considering the complex variations of pedestrian images, a rule of adaptive pairwise constraints is extended for the integrated metric to further enhance separation and reunion between image pairs. Extensive experiments conducted on three person re-identification datasets including VIPeR, PRID450S and GRID indicate that the proposed method outperforms the state-of-the-art methods. Wenbin Yao, Chao Pei, Yuesheng Zhu |
ICIP | 3 |
| 2017 | Pedestrian detection with dynamic iterative bootstrappingabstractRecent years have seen the increasing importance of pedestrian detection, which is a key problem in computer vision. In this paper, we propose a novel pedestrian detection approach based on Faster R-CNN. In order to obtain high-quality candidate regions, relevant adjustments with more precise anchors are made for region proposal network. To resolve the data imbalance issue in the classifier training, we propose a dynamic iterative bootstrapping method where the hard negative examples are automatically selected and the weights of the network are updated iteratively by them to make the training more effective. The square method is used to optimize the multi-task loss in our approach, which can accelerate convergence and reduce sensitivity. Experimental results on different widely used benchmark datasets show that the proposed approach achieves better performance in comparison with other common methods. Chao Pei, Yuesheng Zhu |
ICIP | 1 |