Fengkui Yang

dblp:337/7507 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
0009-0008-8611-2826ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 SpeedSketch: An Ultra-Fast Sketch Generation and Delta Encoding Framework for Delta Compression
abstract
The exponential growth of data poses significant challenges to low-cost and efficient data management. Delta compression has attracted considerable attention as it dramatically reduces storage costs by eliminating redundant data. However, the prohibitive computational overhead incurred during its two critical phases—Sketch Generation and Delta Encoding—impedes its broader industry adoption. In this work, we present SpeedSketch, an ultra-fast framework that bridges Sketch Generation and Delta Encoding through a novel Bloom filter-inspired sketch structure. Based on classification and masking, we devise an innovative Sketch Generation scheme. The sketch generated by this scheme can not only be used to identify similar data but also accelerate the Delta Encoding process by bypassing data that cannot be reduced.
Fengkui Yang, Yuanzhang Wang, Chunhua Li 0002, Ke Zhou 0001
ICPP1
2025 Achieving Better Benefits via Flexible Feature Matching in Post-Deduplication Delta Compression
abstract
Cloud or distributed storage systems characterized by high data redundancy necessitate effective data reduction techniques to reduce storage costs. Post-deduplication delta compression has proven effective by eliminating both duplicated and similar yet non-duplicated chunks. However, existing approaches often rely on fixed-feature matching for resemblance detection, which, while fast, may lead to lower reduction ratios and not robust benefits across various datasets. In this paper, we introduce BePro, a novel system that integrates Flexible Feature Matching (§IV-A) to achieve better benefits in post-deduplication delta compression. BePro employs Gain Filtering (§IV-B) to identify high-gain chunks while discarding low-gain similar chunks, ensuring robust benefits across different datasets. Additionally, BePro implements a new indexing structure, LSH-Delta (§IV-C), to search for similar chunks and utilizes Index Load Balancer (§IV-D) for efficient resemblance detection by exploiting the distribution characteristics of similar chunks. Furthermore, the Index Manager (§IV-E) skillfully manages memory space overhead, ensuring memory efficiency. We implemented a pipeline prototyping framework to facilitate the evaluation of BePro and other leading techniques. Extensive experiments demonstrate that BePro improves the data-reduction ratios by up to$1.15 \times-2.35 \times$while achieving comparable speed.
Fengkui Yang, Bo Mao 0003, Liang Bao, Dongying Zhang, Chunhua Li 0002, Ke Zhou 0001
IPDPS1
2024 LoADM: Load-Aware Directory Migration Policy in Distributed File Systems
abstract
Distributed file systems often suffer from load imbalance when encountering skewed workloads. A few directories can become hotspots due to frequent access. Failure to migrate these high-load directories promptly will result in node overload, which can seriously degrade the performance of the system. To solve this challenge, in this paper, we propose a novel load-aware directory migration policy named LoADM to alleviate the load imbalance caused by hot directories. LoADM consists of three parts, i.e. learning-based directory hotness model, urgency analysis and multidimensional directory migration model. Specifically, we use a directory hotness model to identify potentially high-load directories in advance. Second, by combining the predicted directory hotness and system node status, the urgency analysis determines when to trigger a migration or tolerate an imbalance. Then, peer directory co-migration is proposed to better exploit data locality. Finally, we migrate high-load directories to appropriate storage nodes through a Particle Swarm Optimization based directory migration model. Extensive experiments show that our approach provides a promising data migration policy and can greatly improve performance compared to the state-of-the-art.
Yuanzhang Wang, Fengkui Yang, Ke Zhou 0001, Chunhua Li 0002
DATE3
2024 An optimized learning-based directory placement policy with two-rounds selection in distributed file systems
Yuanzhang Wang, Fengkui Yang, Ke Zhou 0001, Chunhua Li 0002, Ji Zhang 0010
Future Gener. Comput. Syst.2
2022 LDPP: A Learned Directory Placement Policy in Distributed File Systems
abstract
Load balance is a critical problem in distributed file systems. Previous works focus on how to distribute data evenly on different nodes or storage devices from the perspective of file level, but neglect to effectively take advantage of the directory’s locality and the long duration of the directory’s hotness, which may affect the degree of balance and cause performance degradation. To overcome this shortcoming, in this paper, we propose a learning-based directory placement policy, called LDPP, which determines the data layout by predicting the load. We first establish a relationship between directory request characteristics and state information to predict the state information of the directory (storage capacity, bandwidth, and IOPS). Then, the new directory is placed on different nodes in a multi-dimensional manner based on the Manhattan distance according to the predicted multidimensional state information. In addition, we also take into account the trade-off between the same category directory classified by the load prediction module and the peer directories and explore their influence on the balance. Extensive experiments demonstrate that LDPP not only efficiently alleviates load imbalance and increases the utilization of the resources but also improves DFS performance in practice, which can reduce service latency by up to 36 and increase IOPS and bandwidth by 8 and 9, respectively.
Yuanzhang Wang, Fengkui Yang, Ji Zhang 0010, Chunhua Li 0002, Ke Zhou 0001, Jinhu Liu
ICPP2