EDBT 2026 Demo / reviewers in the wild / expert
Fengkui Yang
dblp:337/7507
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2025
0009-0008-8611-2826ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SpeedSketch: An Ultra-Fast Sketch Generation and Delta Encoding Framework for Delta CompressionabstractThe exponential growth of data poses significant challenges to low-cost and efficient data management. Delta compression has attracted considerable attention as it dramatically reduces storage costs by eliminating redundant data. However, the prohibitive computational overhead incurred during its two critical phases—Sketch Generation and Delta Encoding—impedes its broader industry adoption. In this work, we present SpeedSketch, an ultra-fast framework that bridges Sketch Generation and Delta Encoding through a novel Bloom filter-inspired sketch structure. Based on classification and masking, we devise an innovative Sketch Generation scheme. The sketch generated by this scheme can not only be used to identify similar data but also accelerate the Delta Encoding process by bypassing data that cannot be reduced. Fengkui Yang, Yuanzhang Wang, Chunhua Li 0002, Ke Zhou 0001 |
ICPP | 1 |
| 2025 | Achieving Better Benefits via Flexible Feature Matching in Post-Deduplication Delta CompressionabstractCloud or distributed storage systems characterized by high data redundancy necessitate effective data reduction techniques to reduce storage costs. Post-deduplication delta compression has proven effective by eliminating both duplicated and similar yet non-duplicated chunks. However, existing approaches often rely on fixed-feature matching for resemblance detection, which, while fast, may lead to lower reduction ratios and not robust benefits across various datasets. In this paper, we introduce BePro, a novel system that integrates Flexible Feature Matching (§IV-A) to achieve better benefits in post-deduplication delta compression. BePro employs Gain Filtering (§IV-B) to identify high-gain chunks while discarding low-gain similar chunks, ensuring robust benefits across different datasets. Additionally, BePro implements a new indexing structure, LSH-Delta (§IV-C), to search for similar chunks and utilizes Index Load Balancer (§IV-D) for efficient resemblance detection by exploiting the distribution characteristics of similar chunks. Furthermore, the Index Manager (§IV-E) skillfully manages memory space overhead, ensuring memory efficiency. We implemented a pipeline prototyping framework to facilitate the evaluation of BePro and other leading techniques. Extensive experiments demonstrate that BePro improves the data-reduction ratios by up to$1.15 \times-2.35 \times$while achieving comparable speed. Fengkui Yang, Bo Mao 0003, Liang Bao, Dongying Zhang, Chunhua Li 0002, Ke Zhou 0001 |
IPDPS | 1 |
| 2024 | LoADM: Load-Aware Directory Migration Policy in Distributed File SystemsabstractDistributed file systems often suffer from load imbalance when encountering skewed workloads. A few directories can become hotspots due to frequent access. Failure to migrate these high-load directories promptly will result in node overload, which can seriously degrade the performance of the system. To solve this challenge, in this paper, we propose a novel load-aware directory migration policy named LoADM to alleviate the load imbalance caused by hot directories. LoADM consists of three parts, i.e. learning-based directory hotness model, urgency analysis and multidimensional directory migration model. Specifically, we use a directory hotness model to identify potentially high-load directories in advance. Second, by combining the predicted directory hotness and system node status, the urgency analysis determines when to trigger a migration or tolerate an imbalance. Then, peer directory co-migration is proposed to better exploit data locality. Finally, we migrate high-load directories to appropriate storage nodes through a Particle Swarm Optimization based directory migration model. Extensive experiments show that our approach provides a promising data migration policy and can greatly improve performance compared to the state-of-the-art. Yuanzhang Wang, Fengkui Yang, Ke Zhou 0001, Chunhua Li 0002 |
DATE | 3 |
| 2024 | An optimized learning-based directory placement policy with two-rounds selection in distributed file systems
Yuanzhang Wang, Fengkui Yang, Ke Zhou 0001, Chunhua Li 0002, Ji Zhang 0010 |
Future Gener. Comput. Syst. | 2 |
| 2022 | LDPP: A Learned Directory Placement Policy in Distributed File SystemsabstractLoad balance is a critical problem in distributed file systems. Previous works focus on how to distribute data evenly on different nodes or storage devices from the perspective of file level, but neglect to effectively take advantage of the directory’s locality and the long duration of the directory’s hotness, which may affect the degree of balance and cause performance degradation. To overcome this shortcoming, in this paper, we propose a learning-based directory placement policy, called LDPP, which determines the data layout by predicting the load. We first establish a relationship between directory request characteristics and state information to predict the state information of the directory (storage capacity, bandwidth, and IOPS). Then, the new directory is placed on different nodes in a multi-dimensional manner based on the Manhattan distance according to the predicted multidimensional state information. In addition, we also take into account the trade-off between the same category directory classified by the load prediction module and the peer directories and explore their influence on the balance. Extensive experiments demonstrate that LDPP not only efficiently alleviates load imbalance and increases the utilization of the resources but also improves DFS performance in practice, which can reduce service latency by up to 36 and increase IOPS and bandwidth by 8 and 9, respectively. Yuanzhang Wang, Fengkui Yang, Ji Zhang 0010, Chunhua Li 0002, Ke Zhou 0001, Jinhu Liu |
ICPP | 2 |