VLDB 2026 Research / reviewers in the wild / expert
Zhen Jin 0008
dblp:344/5672
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0002-9763-707XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 2 first-author · 9 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OS2G: A High-Performance DPU Offloading Architecture for GPU-based Deep Learning with Object StorageabstractObject storage is increasingly attractive for deep learning (DL) applications due to its cost-effectiveness and high scalability. However, it exacerbates CPU burdens in DL clusters due to intensive object storage processing and multiple data movements. Data processing unit (DPU) offloading is a promising solution, but naively offloading the existing object storage client leads to severe performance degradation. Besides, only offloading the object storage client still involves redundant data movements, as data must first transfer from the DPU to the host and then from the host to the GPU, which continues to consume valuable host resources. Zhen Jin 0008, Yiquan Chen, Mingxu Liang, Guoju Fang, Keyao Zhang, Jiexiong Xu, Wenhai Lin, Yiquan Lin, Shushu Zhao, Wenkai Shi, Zhenhua He, Shishun Cai, Wenzhi Chen |
ASPLOS (2) | 1 |
| 2025 | rInfer: A Generic and High-Performance Framework for Remote Inference with Heterogeneous AcceleratorsabstractInference applications leverage various heterogeneous accelerators, including GPUs, TPUs, and FPGAs, to achieve remarkable performance. Remote inference, which enables applications and accelerators to reside on different nodes, can effectively address low hardware resource utilization and enhance flexibility in resource allocation. However, existing APIremoting solutions face compatibility issues, as they demand extensive development efforts to support diverse heterogeneous accelerators and adapt to rapidly evolving device runtimes. On the other hand, current inference service systems demand cumbersome cross-server configuration and suffer from suboptimal data transmission performance, overlooking data copy overhead in the network stack and not leveraging high-performance RDMA. In this paper, we present rInfer, a generic and highperformance framework for remote inference with heterogeneous accelerators. The key idea of rInfer is to abstract the entire remote inference process and optimize network transmission. On the client node, rInfer offers users an inference-oriented rInfer Device and an ease-of-use rInfer programming model, shielding users from the complexities of server configuration and network communication. On the server node, rInfer integrates with diverse inference frameworks, such as TensorRT and Torch, ensuring high compatibility with various accelerators. Furthermore, we optimize network data transmission by eliminating the need for serialization and deserialization in Protobuf to minimize data copying and adopt RDMA to enhance data transfer efficiency. Experimental results demonstrate that rInfer achieves performance nearly equivalent to local execution for the BERT model, with only a minor difference of 1.5 %. Compared to TorchServe, rInfer-TCP achieves a 51.6 % reduction in execution time for the ResNet models and consumes 37.3% less CPU and 61.2% less memory bandwidth on the server node. Zhen Jin 0008, Yiquan Chen, Yin Du, Keyao Zhang, Jiexiong Xu, Wenhai Lin, Jingchang Qin, Kanghua Fang, Wenzhi Chen |
CCGrid | 1 |
| 2025 | NVMePass: A Lightweight, High-performance and Scalable NVMe Virtualization Architecture with I/O Queues PassthroughabstractMost data-intensive applications currently run on NVMe storage, and virtualization is essential in cloud computing. Existing NVMe virtualization technologies include software-based and hardware-assisted. Virtio suffers from severe performance degradation, and polling-based solutions consume too many valuable CPU resources. Hardware-assisted solutions provide high performance and no CPU usage but have the challenges of developing dedicated hardware.In this paper, we propose NVMePass, a novel software-hardware co-design NVMe passthrough virtualization architecture designed to achieve high performance and no CPU overhead while maintaining high scalability. The key ideas of NVMePass are NVMe I/O queues passthrough for VMs and a mechanism to ensure security. The NVMePass supports DMA and interrupts remapping for VMs without hypervisor involvement, eliminating virtualization overhead and providing near-native performance. Isolation is achieved by I/O queues and logical block address resources exclusively allocated to VMs. We propose NVMe Resource Domain (NRD) and implement it in the NVMe controller to intercept illegal I/O requests. Thus, isolation and security are fully achieved. Results from our experiments show that NVMePass can provide comparable performance to VFIO, with an IOPS of $\mathbf{1 0 0. 1 \% - 1 0 0. 5 \%}$ of VFIO. Furthermore, compared to SPDK-Vhost, NVMePass achieves $\mathbf{4 0. 0 \%}$ lower latency when running 150 VMs, and NVMePass has an improvement of $\mathbf{6 8. 0 \%}$ OPS performance in a real-world application when running 100 VMs. Yiquan Chen, Zhen Jin 0008, Jiexiong Xu, Hao Yu 0016, Wenhai Lin, Kanghua Fang, Keyao Zhang, Chengkun Wei, Yuan Xie 0001, Wenzhi Chen |
HPCA | 2 |
| 2024 | CINDA: Don't Ignore Instructions When Cloning Memory Access BehaviorabstractExisting workload cloning methods suffer from low accuracy as they primarily focus on data access patterns and ignore instruction access. This limitation reduces the accuracy of shared L2 cache design exploration and impedes processor designers from optimizing Icache and ITLB designs. In this paper, we propose CINDA, a novel workload cloning technique that can Capture both INstruction and DAta access patterns of applications. In particular, CINDA separates the instruction and data traces of applications to generate proxy instruction and proxy data traces, subsequently merging them. The results show that CINDA can accurately replicate memory access behavior with 99.1%, 99.9%, and 96.2% accuracy in replicating L1 Icache, ITLB and L2 cache performance, respectively. Furthermore, CINDA outperforms the state-of-the-art methods by reducing 7.7% L2 cache miss error. Wenhai Lin, Yiquan Chen, Jiexiong Xu, Zhen Jin 0008, Peiyu Liu 0003, Shishun Cai, Yuzhong Zhang, Jingchang Qin, Yiquan Lin, Wenzhi Chen |
CCGrid | 4 |
| 2024 | BlueJay: A Platform to Quantifying the Impact of Memory Latency on Datacenter Application PerformanceabstractUnderstanding the impact of memory latency on datacenter application performance can provide decision support to memory subsystem designers. Currently, various methods are available to quantify this impact, including cycle-accurate simulators, memory-level parallelism models, and software delay injection techniques. However, these methods suffer from several limitations, such as slow simulation speed, inaccuracy, and insufficient compatibility that requires application modification.This paper proposes BlueJay, a novel platform to quantify the impact of memory latency on the end-to-end performance of datacenter applications, avoiding slow simulation and providing high accuracy and compatibility. The key idea of BlueJay is to control the consumed memory bandwidth and read/write ratio, thereby manipulating memory latency to achieve quantification. Experiment shows that BlueJay provides accurate quantification with an average error of 3.04%. In addition, we built regression models for five applications deployed at scale in Alibaba data centers. The results reveal that a 10 ns increase in memory latency results in a performance decrease of 2.61%-3.31% for enterprise Java applications and databases, while the elastic block storage service experiences a more modest performance decrease of 0.73%-0.92%. Jingchang Qin, Yiquan Chen, Shishun Cai, Wenhai Lin, Jiexiong Xu, Zhen Jin 0008, Lifa Cao, Yuzhong Zhang, Wenzhi Chen |
CCGrid | 6 |
| 2024 | LightPool: A NVMe-oF-based High-performance and Lightweight Storage Pool Architecture for Cloud-Native Distributed DatabaseabstractEmerging cloud-native distributed databases rely on local NVMe SSDs to provide high-performance and highavailable data services to many cloud applications. However, the database clusters suffer from low utilization of local storage because of the imbalance between CPU and storage capacities within each node. For instance, the OceanBase distributed database cluster, with hundreds of PB local storage capacity, only utilizes around 40% of its local storage. Although disaggregated storage (EBS) can enhance storage utilization by provisioning the CPU and storage independently on demand, they suffer from performance bottlenecks and high costs. In this paper, we propose LightPool, a high-performance and lightweight storage pool architecture large-scale deployed in the OceanBase clusters, enhancing storage resource utilization. The key idea of LightPool is aggregating cluster storage into a storage pool and enabling unified management. In particular, LightPool adopts NVMe-oF to enable high-performance storage resource sharing among cluster nodes and integrate the storage pool with Kubernetes to achieve flexible management and allocation of storage resources. Furthermore, we design the hot-upgrade and hot-migration mechanisms to enhance the availability of LightPool. We have deployed LightPool on over 8500 nodes in production clusters. Statistics show that LightPool can improve storage resource utilization from about 40% to 65%. Experimental results show that the extra latency from LightPool is only about 2.1 μs compared to local storage. Compared to OpenEBS, LightPool enhances bandwidth up to 190.9% in microbenchmarks and throughput up to 6.9% in real-world applications. LightPool is the best practice to deploy NVMe-oF (NVMe/TCP) in the production environment. We also discuss important lessons and experiences learned from the development of LightPool. Jiexiong Xu, Yiquan Chen, Wenhui Shi, Guoju Fang, Huasheng Liao, Zhen Jin 0008, Wenzhi Chen |
HPCA | 10 |
| 2024 | PARS: A Pattern-Aware Spatial Data Prefetcher Supporting Multiple Region SizesabstractHardware data prefetching is a well-studied technique to bridge the processor-memory performance gap. Bit-pattern-based prefetchers are one of the most promising spatial data prefetchers that achieve substantial performance gains. In bit-pattern-based prefetchers, the region size is a crucial parameter, which denotes the memory size that can be recorded by a pattern or prefetched by a prediction. However, existing bit-pattern-based prefetchers only support one fixed region size. Our experiment shows that the fixed region size cannot meet the requirements for numerous applications and leads to suboptimal performance and high hardware overhead. In this article, we propose PARS, a pattern-aware spatial data prefetcher supporting multiple region sizes. The key idea of PARS is that it supports multiple region sizes, enabling it to simultaneously enhance application performance while reducing the hardware overhead. Moreover, PARS supports dynamically switching appropriate region sizes for different patterns through an adaptive RS-switching mechanism. We evaluated PARS on numerous workloads and results show that PARS provides an average performance improvement of 40.6% over a baseline with no data prefetchers and outperforms the two state-of-the-art prefetchers Bingo by 2.1% (up to 24.4%) and Pythia by 3.9% (up to 111.2%) in the single-core system. In the four-core system, PARS outperforms Bingo by 5.0% (up to 66.0%) and Pythia by 5.4% (up to 177.9%). Yiquan Lin, Wenhai Lin, Jiexiong Xu, Yiquan Chen, Zhen Jin 0008, Jingchang Qin, Shishun Cai, Yuzhong Zhang, Zonghui Wang, Wenzhi Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | HyQ: Hybrid I/O Queue Architecture for NVMe over Fabrics to Enable High- Performance Hardware OffloadingabstractNVMe over Fabrics (NVMe-oF) has been widely applied as a remote storage protocol in cloud computing. The existing NVMe-oF software stack consumes a large number of CPU resources. Emerging devices, such as Smart NICs and DPUs, have supported hardware offloading of NVMe-oF to free these valuable CPU cores. However, NVMe-oF offloading capacity is always compromised because of limited hardware resources on design. Additionally, from thorough evaluations, we found that NVMe-oF inevitably suffers from severe performance degradation on complex application I/O patterns when using hardware offloading. It is challenging to achieve high performance and fully utilize NVMe-oF offloading simultaneously. In this paper, we propose HyQ, a novel hybrid I/O queue architecture for NVMe-oF, to achieve high performance while gaining the advantages of hardware offloading. HyQ realizes the coexistence of hardware offloading and software non-offloading queues, thus enabling the dynamic dispatching of I/O requests to appropriate processing queues according to user-defined I/O scheduling policies. Additionally, HyQ provides a request scheduling framework to support customized schedulers that select appropriate queues for I/O requests. In our evaluation, HyQ achieves up to 1.91x IOPS and 8.36x bandwidth performance improvement over the original hardware offloading scheme. Yiquan Chen, Zhen Jin 0008, Jiexiong Xu, Guoju Fang, Wenhai Lin, Chengkun Wei, Wenzhi Chen |
CCGrid | 5 |
| 2023 | JACO: JAva Code Layout Optimizer Enabling Continuous Optimization without Pausing Application ServicesabstractMany Java applications in data centers suffer from severe processor pipeline frontend bottlenecks, which can be mitigated by profile-guided code layout optimizations (PGCLO). To maximize optimization opportunities, state-of-the-art PGCLO solutions adopt continuous optimization to ensure that the code layout consistently matches ever-changing application control flow characteristics. However, existing continuous optimizations inevitably pause the application to execute the new code completely, which leads to high response latency and significantly deteriorates user experience.In this paper, we propose JACO, a novel profile-guided Java code layout optimizer, enabling continuous optimization without pausing application services. The key idea of JACO is to enable the execution of both the old and new code simultaneously rather than completely switching to the new code. In particular, JACO is composed of three components: (1) A lightweight profiler captures the control flow information of the application and then generates an optimized function order. (2) A control flow switcher generates new code based on optimized function order and switches the application to execute the new code without pausing the application services. (3) A selective code reclaimer only frees the memory occupied by the inactive old code. We evaluated JACO on both open-source applications and real-world applications from a world-leading company. JACO achieved up to a 16.36% performance improvement for real-world applications. The state-of-the-art approach introduces up to 37.93x latency overhead that will interrupt application services, while JACO only introduces a negligible 7% latency overhead. Wenhai Lin, Jingchang Qin, Yiquan Chen, Zhen Jin 0008, Jiexiong Xu, Yuzhong Zhang, Shishun Cai, Lirong Fu, Wenzhi Chen |
CLUSTER | 4 |