Wenzhe Zhu

dblp:283/8152 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0002-3965-2597ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Mitigating Dual Load Imbalance via Dynamic Cooperative Scheduling in Distributed Key-Value Stores
Jiakun Zhang, Patrick P. C. Lee, Wenzhe Zhu, Yongkun Li, Yinlong Xu 0001
ICDE3
2026 CoCache: Accelerating Reads in KV Stores via Cooperative Metadata and Data Cache Management
abstract
Modern LSM-based KV stores reduce read amplification with two in-memory caches, namely a table cache for metadata such as index and Bloom-filter blocks, and a data cache for value blocks. These caches draw from a shared memory budget and are jointly exercised on the read path, so the metadata-data split is inherently coupled and end-to-end read latency can be non-monotonic in the allocation. Allocating more memory to one cache may improve its hit rate but evict blocks from the other, yielding hard-to-predict performance especially under dynamic workloads. Most systems therefore rely on fixed, ratio-based heuristics, which can be far from optimal. We present CoCache, a cooperative cache-management framework that continuously tunes the metadata-data cache split. CoCache combines lightweight online hotness tracking with a unified latency model that captures the coupled impact of metadata and data caching on the read path. Using these signals, CoCache efficiently searches candidate splits and applies the one with the lowest predicted latency, then re-optimizes as access patterns shift. We implement CoCache in RocksDB and evaluate it on synthetic, benchmark, and production-derived workloads. Compared to state-of-the-art baselines, CoCache improves cache hit rates by up to 1.56 ×, increases read throughput by up to 1.43 ×, and reduces read latency by up to 31.2%.
Haoting Tang, Wenzhe Zhu, Jiakun Zhang, Junlin Jiang, Yongkun Li 0001, Yinlong Xu 0001
ICS2
2025 Split-LSM-Tree: High-Performance KV Storage via Dynamic Tree Splitting
abstract
Many key-value storage systems employ LSM-tree-based architectures, which often suffer from significant read and write amplification issues that can substantially impact overall performance. While existing solutions mitigate amplification through key-value separation or relaxed ordering, they retain deep hierarchical structures that perpetuate scalability bottlenecks. We present Split-LSM-Tree, a novel storage engine that dynamically partitions monolithic LSM-trees into a forest of capacity-bounded subtrees when thresholds are breached. Our design introduces three key innovations: (1) An adaptive splitting trigger mechanism that initiates partitioning based on storage pressure and performance degradation; (2) A non-blocking splitting algorithm enabling seamless subtree migration; and (3) I/O optimizations preserving foreground operation continuity. Our experiments show that compared to the widely deployed LevelDB and RocksDB, Split-LSM-Tree reduces write amplification by$\mathbf{2 5. 7} \boldsymbol{\%}$and$\mathbf{1 8. 1 \%}$respectively, while achieving$\mathbf{2. 9} \times$and$\mathbf{2. 4} \times$higher write throughput.
Yu-Ang Cao, Fan Guo 0003, Deming Ren, Ruida Xu, Wenzhe Zhu, Yongkun Li 0001, Yinlong Xu 0001
ICPADS6
2025 Double-Metadata-Cache: An Efficient Metadata Caching Architecture for Distributed File Systems
abstract
Current distributed file systems manage metadata through a flattened metadata service, which requires accessing metadata across multiple metadata servers. To improve performance, file systems use client-side cache to reduce frequent server access and employ a lease mechanism to maintain client metadata cache consistency. However, the lease mechanism blocks metadata updates and requires periodic renewal of cache entries. To address these challenges, we propose Double-Metadata-Cache, an efficient two-layer metadata caching architecture. This solution incorporates two key designs: a validity period-based lazy update mechanism to maintain client metadata cache consistency, a segmented caching strategy uses hot prefixes to reduce accesses to metadata servers. Experimental results show that Double-Metadata-Cache achieves higher performance in high-proportion metadata operations such as creation and stat. Moreover, as the file depth increases, it reduces path resolution latency by 50 % compared to the client-side caching strategy.
Fan Guo 0003, Yu-Ang Cao, Deming Ren, Wenzhe Zhu, Yongkun Li 0001, Yinlong Xu 0001
ICPADS5
2025 PTWalker: Cache-Efficient Random Walks via Alternating Dual-Subgraph Walker Updating
abstract
Random walks on graphs are essential for various applications such as network analysis, recommendation systems, and graph embedding. However, existing frameworks for random walks face challenges in efficiently managing and updating walkers. These challenges arise from inefficient walker updating due to the subgraph-based iterative updating strategy, which makes walkers update only one step per iteration, and the high management and storage costs associated with the walker array-based management strategy. This paper presents a novel cache-efficient random walk framework called PTWalker, which sets out to have the majority of walkers updated at least two steps per iteration. This is achieved through an alternating dual-subgraph walk updating scheme, which involves dividing the graph into two sets of subgraphs with maximized edge cuts between them. These subgraph sets are then loaded into the cache for walk updating in an alternating manner, prompting most walkers to transition to the other subgraph set and update additional steps. Besides, a thread-level walker pool management strategy is designed to reduce management and storage costs for walkers. Experimental results demonstrate that PTWalker outperforms state-of-the-art random walk systems, achieving a speedup ranging from 0.18 × to 8.65 ×.
Rui Wang 0076, Long Deng, Wenzhe Zhu, Yongkun Li 0001, Yinlong Xu 0001
ICPP5
2025 Amber: Towards Fast and Space-Efficient Incremental Checkpointing in Large Language Model Training
abstract
Large Language Models have transformed numerous domains with their exceptional capabilities. However, their training processes are inherently prone to failures due to the massive scale of training clusters, involving thousands to tens of thousands of GPUs and requiring several months of uninterrupted computation. Checkpointing serves as a critical mechanism for ensuring fault tolerance. However, the substantial size of LLM checkpoints, attributed to their vast number of parameters, poses significant challenges in terms of time and storage overheads. While existing approaches offer partial solutions, they remain inadequate in addressing the unique demands and complexities of LLM training at scale. To this end, we propose Amber, a novel LLM training framework that significantly improves checkpointing speed and storage efficiency by leveraging selective incremental checkpointing, a technique that selectively checkpoints significant parameter updates while omitting minor changes. We implement Amber atop PyTorch, a widely adopted deep learning framework, and demonstrate that Amber achieves up to 10–155 × faster checkpointing compared to state-of-the-art full checkpointing methods across models of varying scales on a single GPU, while maintaining storage overhead below 3%.
Wenzhe Zhu, Chaomei Yan, Fan Guo 0003, Yongkun Li 0001, Yinlong Xu 0001
ICPP2
2025 FineMem: Breaking the Allocation Overhead vs. Memory Waste Dilemma in Fine-Grained Disaggregated Memory Management
Yongkun Li 0001, Wenzhe Zhu, Yinlong Xu 0001
OSDI4
2024 MinFlow: High-performance and Cost-efficient Data Passing for I/O-intensive Stateful Serverless Analytics
Yongkun Li 0001, Wenzhe Zhu, Yinlong Xu 0001, John C. S. Lui
FAST3
2022 On Optimizing Traffic Imbalance in Large-scale Block-based Cloud Storage: Trace Analysis and Algorithm Design
abstract
Cloud block storage (CBS) serves as the fundamental infrastructure of modern cloud computing services like the cloud disk service. Large-scale cloud block storage usually adopts a layered architecture, including a forwarding layer with a cluster of proxy servers as proxies to provide cloud disk abstraction, and a unified distributed storage engine providing persisted data storage. However, as all I/O traffics go through the proxy servers in the forwarding layer, there may be a severe traffic imbalance between the proxy servers, which finally degrades the performance of cloud disks. To investigate the traffic imbalance problem in the forwarding layer, we first conduct an in-depth analysis on the workload traces of a large-scale cloud block storage system in production. We find that both the traffic of individual cloud disks and the consolidated traffic of cloud disks at proxy servers are highly skewed and fluctuate violently and frequently at a fine-grained time granularity, and thus causing severe traffic imbalance. To address the traffic imbalance issue, we then develop a low-cost migration algorithm, weighted partial migration (WPM), and conduct simulation analysis via trace replay to study its effectiveness. Experiments under real-world workloads show that for 84.3% of clusters, WPM can make the imbalance factor be smaller than 3 (i.e., the maximum traffic at a proxy server is within 3$\times$ of the median traffic), with a very small migration cost by migrating only 0.1% segments.
Haoyu Mao, Yongkun Li 0001, Wenzhe Zhu, Fei Li 0040, Yinlong Xu 0001
ICPADS3
2020 UGNet: Underexposed Images Enhancement Network based on Global Illumination Estimation
abstract
This paper proposes a new neural network for enhancing underexposed images. Instead of the decomposition method based on Retinex theory, we introduce smooth dilated convolution to estimate global illumination of the input image, and implement an end-to-end learning network model. Based on this model, we formulate a multi-term loss function that combines content, color, texture and smoothness losses. Our extensive experiments demonstrate that this method is superior to other methods in underexposed image enhancement. It can cover more color details and be applied to various underexposed images robustly.
Wenzhe Zhu, Qing Zhu 0004
VCIP2
2020 Icon Colorization Based On Triple Conditional Generative Adversarial Networks
abstract
Current automatic colorization systems have many defects such as "contour blur", "color overflow"and "color miscellaneous", especially when they are coloring the images with hollowed-out structure. We propose a model based on triple conditional generative adversarial networks, for generator we provide contour image, colored icon and colorization mask as inputs, our network has three discriminators, structure discriminator is trained to judge if the generated icon has similar contour to the input icon, color discriminator anticipates generated icon and the input icon has the similar color style, the function of mask discriminator is to distinguish whether the output has the similar colorization area to the input mask. For the evaluation, we compared with some existing colorization models, also we made a questionnaire to obtain the evaluation of generated icons from different models. The results showed that our colorization model obtain better results comparing to the other models both in generating hollowed-out and solid structure icons.
Qinru Han, Wenzhe Zhu, Qing Zhu 0004
VCIP2