Yuanzhang Wang

dblp:226/3666 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0003-2646-1887ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 A general lightweight and adaptive cache space allocation scheme
Ke Liu 0014, Hua Wang 0008, Yajun Tan, Peng Wang 0037, Yuanzhang Wang, Ke Zhou 0001, Quan Fu
Future Gener. Comput. Syst.5
2025 SpeedSketch: An Ultra-Fast Sketch Generation and Delta Encoding Framework for Delta Compression
abstract
The exponential growth of data poses significant challenges to low-cost and efficient data management. Delta compression has attracted considerable attention as it dramatically reduces storage costs by eliminating redundant data. However, the prohibitive computational overhead incurred during its two critical phases—Sketch Generation and Delta Encoding—impedes its broader industry adoption. In this work, we present SpeedSketch, an ultra-fast framework that bridges Sketch Generation and Delta Encoding through a novel Bloom filter-inspired sketch structure. Based on classification and masking, we devise an innovative Sketch Generation scheme. The sketch generated by this scheme can not only be used to identify similar data but also accelerate the Delta Encoding process by bypassing data that cannot be reduced.
Fengkui Yang, Yuanzhang Wang, Chunhua Li 0002, Ke Zhou 0001
ICPP2
2024 LoADM: Load-Aware Directory Migration Policy in Distributed File Systems
abstract
Distributed file systems often suffer from load imbalance when encountering skewed workloads. A few directories can become hotspots due to frequent access. Failure to migrate these high-load directories promptly will result in node overload, which can seriously degrade the performance of the system. To solve this challenge, in this paper, we propose a novel load-aware directory migration policy named LoADM to alleviate the load imbalance caused by hot directories. LoADM consists of three parts, i.e. learning-based directory hotness model, urgency analysis and multidimensional directory migration model. Specifically, we use a directory hotness model to identify potentially high-load directories in advance. Second, by combining the predicted directory hotness and system node status, the urgency analysis determines when to trigger a migration or tolerate an imbalance. Then, peer directory co-migration is proposed to better exploit data locality. Finally, we migrate high-load directories to appropriate storage nodes through a Particle Swarm Optimization based directory migration model. Extensive experiments show that our approach provides a promising data migration policy and can greatly improve performance compared to the state-of-the-art.
Yuanzhang Wang, Fengkui Yang, Ke Zhou 0001, Chunhua Li 0002
DATE1
2024 An optimized learning-based directory placement policy with two-rounds selection in distributed file systems
Yuanzhang Wang, Fengkui Yang, Ke Zhou 0001, Chunhua Li 0002, Ji Zhang 0010
Future Gener. Comput. Syst.1
2022 LDPP: A Learned Directory Placement Policy in Distributed File Systems
abstract
Load balance is a critical problem in distributed file systems. Previous works focus on how to distribute data evenly on different nodes or storage devices from the perspective of file level, but neglect to effectively take advantage of the directory’s locality and the long duration of the directory’s hotness, which may affect the degree of balance and cause performance degradation. To overcome this shortcoming, in this paper, we propose a learning-based directory placement policy, called LDPP, which determines the data layout by predicting the load. We first establish a relationship between directory request characteristics and state information to predict the state information of the directory (storage capacity, bandwidth, and IOPS). Then, the new directory is placed on different nodes in a multi-dimensional manner based on the Manhattan distance according to the predicted multidimensional state information. In addition, we also take into account the trade-off between the same category directory classified by the load prediction module and the peer directories and explore their influence on the balance. Extensive experiments demonstrate that LDPP not only efficiently alleviates load imbalance and increases the utilization of the resources but also improves DFS performance in practice, which can reduce service latency by up to 36 and increase IOPS and bandwidth by 8 and 9, respectively.
Yuanzhang Wang, Fengkui Yang, Ji Zhang 0010, Chunhua Li 0002, Ke Zhou 0001, Jinhu Liu
ICPP1
2020 Tier-Scrubbing: An Adaptive and Tiered Disk Scrubbing Scheme with Improved MTTD and Reduced Cost
abstract
Sector errors are a common type of error in modern disks. A sector error that occurs during I/O operations might cause inaccessibility of an application. Even worse, it could result in permanent data loss if the data is being reconstructed, and thereby severely affects the reliability of a storage system. Many disk scrubbing schemes have been proposed to solve this problem. However, existing approaches have several limitations. First, schemes use machine learning (ML) to predict latent sector errors (LSEs), but only leverage a single snapshot of training data to make a prediction, and thereby ignore sequential dependencies between different statuses of a hard disk over time. Second, they accelerate the scrubbing at a fixed rate based on the results of a binary classification model, which may result in unnecessary increases in scrubbing cost. Third, they naively accelerate the scrubbing of the full disk which has LSEs based on the predictive results, but neglect partial high-risk areas (the areas that have a higher probability of encountering LSEs). Lastly, they do not employ strategies to scrub these high-risk areas in advance based on I/O accesses patterns, in order to further increase the efficiency of scrubbing.We address these challenges by designing a Tier-Scrubbing (TS) scheme that combines a Long Short-Term Memory (LSTM) based Adaptive Scrubbing Rate Controller (ASRC), a module focusing on sector error locality to locate high-risk areas in a disk, and a piggyback scrubbing strategy to improve the reliability of a storage system. Our evaluation results on realistic datasets and workloads from two real world data centers demonstrate that TS can simultaneously decrease the Mean-Time-To-Detection (MTTD) by about 80% and the scrubbing cost by 20%, compared to a state-of-the-art scrubbing scheme.
Ji Zhang 0010, Yuanzhang Wang, Yangtao Wang, Ke Zhou 0001, Sebastian Schelter, Ping Huang 0001, Yong-guang Ji
DAC2