Zongpu Zhang

dblp:228/1374 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-2548-5732ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorArtificial intelligence and machine learning · 1Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 P4KVS: A Role-Replica Separation Offloading Method to Achieve In-Network Consistency for KV Stores Based on P4 Switches
abstract
Strong consistency, particularly linearizability, is essential for distributed DBMSs deployed in correctness-critical domains such as finance and defense. In general, an optimal linearizability DBMS system focus on two key principles: (1) matching single-node (no-consistency cost) Read/Write performance under strong consistency, and (2) practical deployability via general database compatibility. Unfortunately, existing solutions fall short on both fronts. %However, achieving strong consistency often comes with steep performance penalties. For example, etcd-a widely-used Raft-based system-achieves only ~5.9% of the throughput of LevelDB, a single-node store without consistency overhead. To achieve higher performance, software approaches adopt weaker consistency models (e.g., ZAB), rely on narrow network assumptions (e.g., NOPaxos), or expose protocol internals to clients (e.g., CURP), yet still fail to close the performance gap. Recent programmable networking hardware offers promising advances, yet current hardware solutions face practical limitations, including minimal storage and incompatibility with general-purpose databases. We propose P4KVS, the first practical Raft-based in-network consensus offloading solution leveraging programmable switches (P4) for distributed key-value stores. P4KVS offloads only the Leader role to the switch while retaining Followers on servers. Under linearizability, it achieves 74% of single-node LevelDB's throughput for write-heavy workloads, and up to 222.4% for read-heavy workloads by distributing reads across three replicas. This demonstrates that, even under strong consistency, P4KVS can match or exceed the performance of a single-node system. Compared to etcd (which also uses Raft), P4KVS delivers 37.5× higher read throughput and 3520× lower write latency. These results validate our hardware role-replica separation design in eliminating software Raft bottlenecks, while preserving compatibility via standard database interfaces (e.g., LevelDB, etcd) and scaling beyond typical switch memory constraints.
Haojuan Li, Zongpu Zhang, Chenzhen Ye, Ruohan Tang, Jian Li 0021, Haibing Guan, Qiaoling Wang, Pengpeng Zhou
Proc. ACM Manag. Data2
2025 Ephemera: Accelerating I/O-Intensive Serverless Workloads with a Harvested In-memory File System
abstract
Serverless computing has gained popularity for its ability to shift the burden of server management from developers to cloud providers, which allows providers to exercise greater control over resource management, optimizing configurations to enhance efficiency and performance. The diversity of serverless computing tasks, from short-lived, event-driven tasks to more complex workloads, highlights the growing importance of efficient file I/O performance for I/O-intensive workloads, yet effectively handling ephemeral storage for I/O-intensive tasks remains a challenge. Traditional file system approaches often introduce substantial latency and fail to fully leverage available memory resources within the execution environment, limiting performance and efficiency. Our work stems from the observation of the under-utilization of memory resources in serverless computing platforms and the potential efficiency improvement of I/O operations using an in-memory file system. Based on this observation, we propose Ephemera , a system designed to enhance ephemeral storage efficiency and memory utilization. Ephemera satisfies three design goals: transparent memory I/O integration , heterogeneous tasks resource synergy , and harmonized cluster workload orchestration . Ephemera integrates three components: the Runtime Daemon, responsible for managing a container’s in-memory file system; the Tenant Manager, facilitating memory configuration sharing across containers; and the Cluster Controller, optimizing workload balancing. Our experiments demonstrate that Ephemera significantly improves performance for I/O-intensive tasks compared to traditional file systems. Specifically, Ephemera decreases I/O processing time by 50% on average and reduces latency by up to 95.73% in certain scenarios with negligible overhead.
Lingxiao Jin, Zinuo Cai, Haoxin Wang 0005, Zongpu Zhang, Ruhui Ma, Haibing Guan, Yuan Liu 0021, Rajkumar Buyya
ACM Trans. Archit. Code Optim.4
2025 SMore: Enhancing GPU Utilization in Deep Learning Clusters by Serverless-Based Co-Location Scheduling
Junhan Liu, Zinuo Cai, Yumou Liu, Hao Li 0142, Zongpu Zhang, Ruhui Ma, Rajkumar Buyya
IEEE Trans. Parallel Distributed Syst.5
2024 HD-IOV: SW-HW Co-designed I/O Virtualization with Scalability and Flexibility for Hyper-Density Cloud
abstract
As the resource density of cloud servers increases, cloud providers deploy hundreds of VMs concurrently on a single server, requiring a high-performance, scalable, flexible and high-density I/O virtualization method. Hardware assisted virtualization such as device pass-through with SR-IOV can achieve near-native performance, however, at the expense of flexibility and a limited device count. Traditional software-based I/O virtualization systems tend to dedicate additional computing cores for higher performance, but suffer from critical scalability problems especially in high-density cloud.
Zongpu Zhang, Jiangtao Chen, Banghao Ying, Yahui Cao, Lingyu Liu, Jian Li 0021, Weigang Li 0002, Haibing Guan
EuroSys1
2024 vCrypto: a Unified Para-Virtualization Framework for Heterogeneous Cryptographic Resources
abstract
Transport Layer Security (TLS) connections involve costly cryptographic operations which incur significant resource consumption in the cloud. Hardware accelerators are affordable substitutes of expensive CPU cores to accommodate with the constantly increasing security requirements of datacenters. Existing accelerators virtualization mainly relies on passthrough of Single Root I/O Virtualization (SR-IOV) devices. However, deficiency of service accessibility, functionality and availability make device passthrough not an optimal solution for heterogeneous accelerators with different capabilities. To make up the gap, we propose vCrypto, a unified para-virtualization framework for heterogeneous cryptographic resources. vCrypto supports stateful crypto requests offloading and result retrieval with session lifecycle management and event driven notification. vCrypto transparently integrates virtual crypto device capabilities into the OpenSSL framework to benefit existing applications that are based on crypto library APIs without modification. Multiple physical resources can be partitioned flexibly and scheduled cooperatively to enhance the functionality, performance and robustness of virtual crypto service. Finally, vCrypto achieves an optimized performance with two layers polling and memory sharing mechanism. The comprehensive experiments show that with the same cryptographic resources used, vCrypto framework can provide 2.59x to 3.36x higher AES-CBC-HMAC-SHA1 throughput compared to passthrough SR-IOV device.
Chao Zhang 0115, Zongpu Zhang, Hubin Zhang, Weigang Li 0002, Yibin Shen, Jian Li 0021, Haibing Guan
INFOCOM3
2024 Un-IOV: Achieving Bare-Metal Level I/O Virtualization Performance for Cloud Usage With Migratability, Scalability and Transparency
abstract
I/O virtualization is utilized by cloud platforms to provide tenants with efficient, scalable, and manageable network and storage services. The de-facto industrial standard, paravirtualization, offers rich cloud functionality by introducing split front-end and back-end drivers in the guest and host operating systems, respectively. Given this fact, paravirtualization incurs host inefficiency and performance overhead. Thus, emerging hardware virtio accelerators (i.e., SRIOV-capable devices that conform to virtio specification) with device passthrough technologies mitigate the performance issue. However, adopting these devices presents the challenge of insufficient support for live migration.This paper proposes Un-IOV, a novel I/O virtualization system that simultaneously achieves bare-metal level I/O performance and migratability. The key idea is to develop a new hybrid virtualization stack with: (1) a host-bypassed direct data path for virtio accelerators, and (2) a relayed control path guaranteeing seamless live migration support. Un-IOV achieves high scalability by consuming minimum host resources. Extensive experiment results demonstrate that Un-IOV achieves superior network and storage virtualization performance than software implementations with comparable performance of direct passthrough I/O virtualization, while imposing zero guest modification (i.e., guest transparency).
Zongpu Zhang, Chenbo Xia, Cunming Liang, Jian Li 0021, Chen Yu 0003, Tiwei Bie, Roberts Martin, Dan Daly, Xiao Wang 0084, Haibing Guan
IEEE Trans. Computers1
2023 QKPT: Securing Your Private Keys in Cloud With Performance, Scalability and Transparency
abstract
Private key (e.g., RSA key) protection is a significant issue for cloud but existing keyless or keyguard solutions suffer from performance, elasticity or applicability limitations. Recently, represented by Intel KPT, a novel keyguard architecture emerges to combine trusted platform module and crypto accelerator for achieving both security and performance. However, the straight use of KPT for private key protection may not be a good fit in cloud as it incurs challenges on protection capacity, key provisioning latency and transparency. Based on KPT-like hardware, we propose QKPT, a comprehensive key management system to bring your own private keys (BYOPK) into multi-tenant clouds. QKPT introduces a carefully-designed key wrapping layer to overcome these challenges. A small symmetric wrapping key (SWK) is generated for each tenant as the master key to resolve the former two challenges, while a special private key wrapping scheme is adopted to resolve the transparency limitation. Additionally, QKPT incorporates certificate trust to enhance the security of the SWK lifecycle and provides a hardened key server solution without expensive HSM. The evaluation shows that QKPT has a low runtime overhead ($\leq$1.2% for SSL/TLS handshakes) and still greatly outperforms the software baseline (3.5x-17x) owing to the crypto offloading.
Zongpu Zhang, Hubin Zhang, Xiaokang Hu, Jian Li 0021, Weigang Li 0002, Guodong Zhu, Kapil Sood, Brian Will, Haibing Guan
IEEE Trans. Dependable Secur. Comput.1
2022 Towards Ubiquitous Intelligent Computing: Heterogeneous Distributed Deep Neural Networks
abstract
For the pursuit of ubiquitous computing, distributed computing systems containing the cloud, edge devices, and Internet-of-Things devices are highly demanded. However, existing distributed frameworks do not tailor for the fast development of Deep Neural Network (DNN), which is the key technique behind many intelligent applications nowadays. Based on prior exploration on distributed deep neural networks (DDNN), we propose Heterogeneous Distributed Deep Neural Network (HDDNN) over the distributed hierarchy, targeting at ubiquitous intelligent computing. While being able to support basic functionalities of DNNs, our framework is optimized for various types of heterogeneity, including heterogeneous computing nodes, heterogeneous neural networks, and heterogeneous system tasks. Besides, our framework features parallel computing, privacy protection and robustness, with other consideration for the combination of heterogeneous distributed system and DNN. Extensive experiments demonstrate that our framework is capable of utilizing hierarchical distributed system better for DNN and tailoring DNN for real-world distributed system properly, which is with low response time, high performance, and better user experience.
Zongpu Zhang, Tao Song 0003, Yang Hua 0001, Xufeng He, Zhengui Xue, Ruhui Ma, Haibing Guan
IEEE Trans. Big Data1
2019 Object Guided External Memory Network for Video Object Detection
abstract
Video object detection is more challenging than image object detection because of the deteriorated frame quality. To enhance the feature representation, state-of-the-art methods propagate temporal information into the deteriorated frame by aligning and aggregating entire feature maps from multiple nearby frames. However, restricted by feature map's low storage-efficiency and vulnerable content-address allocation, long-term temporal information is not fully stressed by these methods. In this work, we propose the first object guided external memory network for online video object detection. Storage-efficiency is handled by object guided hard-attention to selectively store valuable features, and long-term information is protected when stored in an addressable external data matrix. A set of read/write operations are designed to accurately propagate/allocate and delete multi-level memory feature under object guidance. We evaluate our method on the ImageNet VID dataset and achieve state-of-the-art performance as well as good speed-accuracy tradeoff. Furthermore, by visualizing the external memory, we show the detailed object-level reasoning process across frames.
Hanming Deng, Yang Hua 0001, Tao Song 0003, Zongpu Zhang, Zhengui Xue, Ruhui Ma, Neil Robertson 0002, Haibing Guan
ICCV4
2019 Unsupervised Video Summarization with Attentive Conditional Generative Adversarial Networks
abstract
With the rapid growth of video data, video summarization technique plays a key role in reducing people's efforts to explore the content of videos by generating concise but informative summaries. Though supervised video summarization approaches have been well studied and achieved state-of-the-art performance, unsupervised methods are still highly demanded due to the intrinsic difficulty of obtaining high-quality annotations. In this paper, we propose a novel yet simple unsupervised video summarization method with attentive conditional Generative Adversarial Networks (GANs). Firstly, we build our framework upon Generative Adversarial Networks in an unsupervised manner. Specifically, the generator produces high-level weighted frame features and predicts frame-level importance scores, while the discriminator tries to distinguish between weighted frame features and raw frame features. Furthermore, we utilize a conditional feature selector to guide GAN model to focus on more important temporal regions of the whole video frames. Secondly, we are the first to introduce the frame-level multi-head self-attention for video summarization, which learns long-range temporal dependencies along the whole video sequence and overcomes the local constraints of recurrent units, e.g., LSTMs. Extensive evaluations on two datasets, SumMe and TVSum, show that our proposed framework surpasses state-of-the-art unsupervised methods by a large margin, and even outperforms most of the supervised methods. Additionally, we also conduct the ablation study to unveil the influence of each component and parameter settings in our framework.
Xufeng He, Yang Hua 0001, Tao Song 0003, Zongpu Zhang, Zhengui Xue, Ruhui Ma, Neil Robertson 0002, Haibing Guan
ACM Multimedia4
2018 Tracking-assisted Weakly Supervised Online Visual Object Segmentation in Unconstrained Videos
abstract
This paper tackles the task of online video object segmentation with weak supervision, i.e., labeling the target object and background with pixel-level accuracy in unconstrained videos, given only one bounding box information in the first frame. We present a novel tracking-assisted visual object segmentation framework to achieve this. On the one hand, initialized with a given bounding box in the first frame, the auxiliary object tracking module guides the segmentation module frame by frame by providing motion and region information, which is usually missing in semi-supervised methods. Moreover, compared with the unsupervised approach, our approach with such minimum supervision can focus on the target object without bringing unrelated objects into the final results. On the other hand, the video object segmentation module also improves the robustness of the visual object tracking module by pixel-level localization and objectness information. Thus, segmentation and tracking in our framework can mutually help each other in an online manner. To verify the generality and effectiveness of the proposed framework, we evaluate our weakly supervised method on two cross-domain datasets, i.e., the DAVIS and VOT2016 datasets, with the same configuration and parameter setting. Experimental results show the top performance of our method, which is even better than the leading semi-supervised methods. Furthermore, we conduct the extensive ablation study on our approach to investigate the influence of each component and main parameters.
Zongpu Zhang, Yang Hua 0001, Tao Song 0003, Zhengui Xue, Ruhui Ma, Neil Robertson 0002, Haibing Guan
ACM Multimedia1