Xinyang Shao

dblp:256/0887 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-8694-3979ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Towards Agile and Judicious Metadata Load Balancing for Ceph File System via Matrix-based Modeling
abstract
To scale out the massive metadata access, the Ceph distributed file system (CephFS) adopts a dynamic subtree partitioning method, splitting the hierarchical namespace and distributing subtrees across multiple metadata servers. However, this method suffers from a severe imbalance problem that may result in poor performance due to its inaccurate imbalance prediction, ignorance of workload characteristics, and unnecessary/invalid migration activities. To eliminate these inefficiencies, we propose Lunule, a novel CephFS metadata load balancer, which employs an imbalance factor model for accurately determining when to trigger re-balance and tolerate unharmful imbalanced situations. Lunule further adopts a workload-aware migration planner to appropriately select subtree migration candidates. Finally, we extend Lunule to Lunule + , which models metadata accesses into matrices, and employs matrix-based formulas for more accurate load prediction and re-balance decision. Compared to baselines, Lunule achieves better load balance, increases the metadata throughput by up to 315.8%, and shortens the tail job completion time by up to 64.6% for five real-world workloads and their mixture, respectively. Besides, Lunule is capable of handling the metadata cluster expansion and the workload growth, and scales linearly on a 16-node cluster. Compared to Lunule, Lunule + achieves up to 64.96% better metadata load balance, and 13.53-86.09% higher throughput.
Xinyang Shao, Yiduo Wang 0002, Cheng Li 0001, Hengyu Liang, Chenhan Wang, Feng Yan 0001, Yinlong Xu 0001
ACM Trans. Storage1
2023 MUSE: A Programmable Metadata Load Estimation Interface for Ceph File System
abstract
CephFS represents a prominent distributed file system that utilizes directory fragment migration to achieve improved runtime balance. However, its imprecise imbalance model and subtree selection algorithms can result in suboptimal performance. Our prior work, Lunule, enhances CephFS by introducing an imbalance factor model and a workload-aware load estimation policy. Nevertheless, Lunule’s built-in workload-aware planner still relies on a unified formula with adjustable coefficients, representing a one-size-fits-all approach. In this study, we introduce MUSE, a novel and user-friendly programmable interface that specifically focuses on subtree migration planning. MUSE effectively separates the complex and challenging task of evaluating subtree loads for different workloads, enabling designers to manipulate expected loads in the subsequent epoch. This facilitates the selection of appropriate subtrees for migration and opens up possibilities for implementing true isolated workloadaware policies. Through the utilization of two small Lua scripts, we demonstrate that MUSE achieves comparable load balancing effects and performance to CephFS and Lunule in workloads characterized by temporal and spatial locality.
Xinyang Shao, Cheng Li 0001, Yinlong Xu 0001
ICPADS1
2021 Lunule: an agile and judicious metadata load balancer for CephFS
abstract
For a decade, the Ceph distributed file system (CephFS) has been widely used to serve the ever-growing big data in many key fields ranging from Internet services to AI computing. To scale out the massive metadata access, CephFS adopts a dynamic subtree partitioning method, splitting the hierarchical namespace and distributing subtrees across multiple metadata servers. However, this method suffers from a severe imbalance problem that may result in poor performance due to its inaccurate imbalance prediction, ignorance of workload characteristics, and unnecessary/invalid migration activities. To eliminate these inefficiencies, we propose Lunule, a novel CephFS metadata load balancer, which employs an imbalance factor model for accurately determining when to trigger re-balance and tolerate benign imbalanced situations. Lunule further adopts a workload-aware migration planner to appropriately select subtree migration candidates. Compared to baselines, Lunule achieves better load balance, increases the metadata throughput by up to 315.8%, and shortens the tail job completion time by up to 64.6% for five real-world workloads and their mixture, respectively. Besides, Lunule is capable of handling the metadata cluster expansion and the client workload growth, and scales linearly on a cluster of 16 MDSs.
Yiduo Wang 0002, Cheng Li 0001, Xinyang Shao, Youxu Chen, Feng Yan 0001, Yinlong Xu 0001
SC3
2019 Explicit Data Correlations-Directed Metadata Prefetching Method in Distributed File Systems
abstract
Metadata performance in distributed file systems (DFS) is critical, due to the following trends: (a) the growing size of modern storage systems is expected to exceed billions of files and most files are small; (b) over half of the file accesses are metadata operations. In this work, we present SMeta, a metadata prefetching method that is seamlessly integrated into DFS for easy-of-use and significantly scales the metadata performance. Previous prefetching proposals primarily focus on mining groups of files that tend to be accessed together from the access history. Nevertheless, our study discovered that these solutions likely miss a huge number of correlated files whose co-occurrence frequency is not high enough. Unlike access correlations, we take a novel and completely different approach to explore explicit data correlations by understanding the reference relationships between files encoded in some forms of hyperlinks, which naturally exist in many applications. To embrace this new concept, SMeta explores correlations upon files are written via a light-weight pattern matching algorithm, stores correlations in the reserved extended attributes of file metadata to avoid changes in DFS APIs, and collapses multiple I/O rounds for accessing metadata of the target file and its data-correlated files into one round. A cost-efficient adaptive feedback mechanism is introduced to improve prefetching accuracy. We implemented SMeta atop of Ceph and evaluated it using synthetic and real system workloads. Compared to baselines, SMeta provides better metadata performance in terms of latency, throughput and scalability.
Youxu Chen, Cheng Li 0001, Min Lv, Xinyang Shao, Yongkun Li 0001, Yinlong Xu 0001
IEEE Trans. Parallel Distributed Syst.4