EDBT 2026 Demo / reviewers in the wild / expert
Yiduo Wang 0002
dblp:221/2416-2
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0002-8787-0134ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Origami: Efficient ML-Driven Metadata Load Balancing for Distributed File SystemsabstractModern distributed file systems (DFSs) rely on metadata server clusters to manage large-scale files and achieve scalability. However, the hierarchical namespace structure and dynamic user workloads pose severe challenges for efficient metadata partitioning and load balancing. Existing approaches primarily focus on identifying and redistributing hot metadata to address imbalances. While these load-balancing strategies offer potential benefits, they often reduce metadata locality, ultimately failing to improve the end-to-end job completion time—a key metric prioritized by users. Although recent research reveals that learning-based approaches are effective in predicting hotspots, they have been shown to be less effective in improving metadata performance. We revisit metadata load balancing strategies and propose a learning-based metadata load balance framework Origami, which focuses on minimizing end-to-end job completion time rather than equalizing loads. Origami first decomposes the overhead of metadata operations and assesses the impact of migration decisions on user requests, allowing us to compute the benefits of migration decisions for job completion time when future requests are known. Subsequently, Origami propose the Meta-OPT algorithm to determine near-optimal migration decisions. Finally, we implemented OrigamiFS, on which we collected statistical data to train and validate ML-models capable of predicting migration benefits. By predicting the benefits of migration decisions and employing Meta-OPT to quickly explore nearly optimal migration decisions, Origami makes a better trade-off between load balancing and namespace locality. Our evaluation shows that compared to state-of-the-art methods, Origami increases aggregated metadata throughput by 1.12-2.51 × across three real-world workloads, and enhances end-to-end throughput by 1.11-2.02 ×. Yiduo Wang 0002, Wenda Tang, Linghang Meng, Liang Li 0016, Jie Wu 0001 |
ICPP | 1 |
| 2025 | Leave No One Behind: Fair and Efficient Tiered Memory Management for Multi-ApplicationsabstractThe emergence of byte-addressable memory technologies, such as CXL-attached memory, has catalyzed extensive research into tiered memory management. Existing tiering solutions optimize system-wide performance by migrating frequently accessed (“hot”) data to fast-tier memory via page migration, which serves as the de facto mechanism in modern OS. However, these strategies often fail in the multi-tenant environment, where diverse workloads interfere with each other. For example, latency-critical workloads co-located with throughput-oriented ones may face the “cold page dilemma,” where critical pages are misclassified as “cold” and migrated to slower tiers, leading to significant performance degradation. Moreover, current methods often assume negligible migration overhead, which becomes problematic in multi-core systems handling write-intensive workloads that incur substantial costs. This paper proposes Vulcan, a workload-aware tiered memory management framework that targets fair and efficient tiering in multi-tenant environments. Vulcan introduces four key innovations: (1) workload-dependent migration mechanism, which decouples page migration from the OS kernel to enhance operational flexibility for multi-workloads; (2) QoS-aware fair resource partitioning, which dynamically optimizes fast memory distribution through per-workload fast tier hit ratios and fairness-oriented allocation policies; (3) per-thread page table replication, which minimizes TLB coherence overhead during migration; and (4) biased page migration policy, which optimizes efficiency by considering both access characteristics (read-intensive vs. write-intensive) and thread-level page ownership (private vs. shared). We evaluated Vulcan using multiple representative cloud applications with realistic working sets in co-location scenarios. Vulcan improves performance by 12.4% on average and achieves a 75.3% improvement in fairness compared to existing state-of-the-art memory tiering solutions. Wenda Tang, Yiduo Wang 0002, Yanwen Wang 0002, Jie Wu 0001 |
ICPP | 2 |
| 2025 | Mantle: Efficient Hierarchical Metadata Management for Cloud Object Storage Services
Biao Cao, Jielong Jian, Cheng Li 0001, Sen Han, Yiduo Wang 0002, Yufei Wu 0011, Kang Chen 0001, Zhihui Yin, Jiwei Xiong, Jie Zhao 0020, Liguo Duan, Miao Yu 0029, Feng Wu 0001, Xianjun Meng |
SOSP | 6 |
| 2025 | Towards Agile and Judicious Metadata Load Balancing for Ceph File System via Matrix-based ModelingabstractTo scale out the massive metadata access, the Ceph distributed file system (CephFS) adopts a dynamic subtree partitioning method, splitting the hierarchical namespace and distributing subtrees across multiple metadata servers. However, this method suffers from a severe imbalance problem that may result in poor performance due to its inaccurate imbalance prediction, ignorance of workload characteristics, and unnecessary/invalid migration activities. To eliminate these inefficiencies, we propose Lunule, a novel CephFS metadata load balancer, which employs an imbalance factor model for accurately determining when to trigger re-balance and tolerate unharmful imbalanced situations. Lunule further adopts a workload-aware migration planner to appropriately select subtree migration candidates. Finally, we extend Lunule to Lunule + , which models metadata accesses into matrices, and employs matrix-based formulas for more accurate load prediction and re-balance decision. Compared to baselines, Lunule achieves better load balance, increases the metadata throughput by up to 315.8%, and shortens the tail job completion time by up to 64.6% for five real-world workloads and their mixture, respectively. Besides, Lunule is capable of handling the metadata cluster expansion and the workload growth, and scales linearly on a 16-node cluster. Compared to Lunule, Lunule + achieves up to 64.96% better metadata load balance, and 13.53-86.09% higher throughput. Xinyang Shao, Yiduo Wang 0002, Cheng Li 0001, Hengyu Liang, Chenhan Wang, Feng Yan 0001, Yinlong Xu 0001 |
ACM Trans. Storage | 2 |
| 2023 | CFS: Scaling Metadata Service for Distributed File System via Pruned Scope of Critical SectionsabstractThere is a fundamental tension between metadata scalability and POSIX semantics within distributed file systems. The bottleneck lies in the coordination, mainly locking, used for ensuring strong metadata consistency, namely, atomicity and isolation. CFS is a scalable, fully POSIX-compliant distributed file system that eliminates the metadata management bottleneck via pruning the scope of critical sections for reduced locking overhead. First, CFS adopts a tiered metadata organization to scale file attributes and the remaining namespace hierarchies independently with appropriate partitioning and indexing methods, eliminating cross-shard distributed coordination. Second, it further scales up the single metadata shard performance by single-shard atomic primitives, shortening the metadata requests' lifespan and removing spurious conflicts. Third, CFS drops the metadata proxy layer but employs the light-weight, scalable client-side metadata resolving. CFS has been running in the production environment of Baidu AI Cloud for three years. Our evaluation with a 50-node cluster and microbenchmarks shows that CFS simultaneously improves the throughput of baselines like HopsFS and InfiniFS by 1.76--75.82× and 1.22--4.10×, and reduces their average latency by up to 91.71% and 54.54%, respectively. Under cases with higher contention and larger directories, CFS' throughput benefits expand by one order of magnitude. For three real-world workloads with data accesses, CFS introduces 1.62--2.55× end-to-end throughput speedups and 35.06--62.47% tail latency reductions over InfiniFS. Yiduo Wang 0002, Yufei Wu 0011, Cheng Li 0001, Biao Cao, Yinlong Xu 0001, Guangjun Xie |
EuroSys | 1 |
| 2021 | Lunule: an agile and judicious metadata load balancer for CephFSabstractFor a decade, the Ceph distributed file system (CephFS) has been widely used to serve the ever-growing big data in many key fields ranging from Internet services to AI computing. To scale out the massive metadata access, CephFS adopts a dynamic subtree partitioning method, splitting the hierarchical namespace and distributing subtrees across multiple metadata servers. However, this method suffers from a severe imbalance problem that may result in poor performance due to its inaccurate imbalance prediction, ignorance of workload characteristics, and unnecessary/invalid migration activities. To eliminate these inefficiencies, we propose Lunule, a novel CephFS metadata load balancer, which employs an imbalance factor model for accurately determining when to trigger re-balance and tolerate benign imbalanced situations. Lunule further adopts a workload-aware migration planner to appropriately select subtree migration candidates. Compared to baselines, Lunule achieves better load balance, increases the metadata throughput by up to 315.8%, and shortens the tail job completion time by up to 64.6% for five real-world workloads and their mixture, respectively. Besides, Lunule is capable of handling the metadata cluster expansion and the client workload growth, and scales linearly on a cluster of 16 MDSs. Yiduo Wang 0002, Cheng Li 0001, Xinyang Shao, Youxu Chen, Feng Yan 0001, Yinlong Xu 0001 |
SC | 1 |