VLDB 2026 Research / reviewers in the wild / expert
Ruiming Lu
dblp:313/3962
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0002-4236-289XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scheduling Cloud Block Storage Proactively and Reactively with Omar
Xinqi Chen, Weidong Zhang 0011, Erci Xu, Junping Wu, Ruiming Lu, Yaheng Song, Chaolei Hu, Lijun Ding, Guangtao Xue, Patrick P. C. Lee |
EuroSys | 9 |
| 2026 | Here, There and Everywhere: The Past, the Present and the Future of Local Storage in Cloud
Leping Yang, Yanbo Zhou, Gong Zeng, Saisai Zhang, Ruilin Wu, Chaoyang Sun, Shiyi Luo, Keqiang Niu, Junping Wu, Jiaji Zhu, Jiesheng Wu, Mariusz Barczak, Wayne Gao, Ruiming Lu, Erci Xu, Guangtao Xue |
FAST | 17 |
| 2025 | One-Size-Fits-None: Understanding and Enhancing Slow-Fault Tolerance in Modern Distributed Systems
Ruiming Lu, Yunchi Lu, Yuxuan Jiang 0016, Guangtao Xue, Peng Huang 0005 |
NSDI | 1 |
| 2025 | MasterPlan: A Reinforcement Learning Based Scheduler for Archive StorageabstractWith the sheer volume of data in today’s world, archive storage systems play a significant role in persisting the cold data. Due to stringent cost concerns, one popular design is to organize disks into groups and periodically switch them to be powered on for serving user requests. Scheduling thus becomes critical for both CapEx and performance. Unfortunately, field results indicate that existing schedulers can be often suboptimal. Our further analysis suggests that the main reason is the mismatch between the ever-changing workloads and the fixed set of coarsely-configured parameters in current heuristic-based schedulers. In this article, we propose MasterPlan , a reinforcement learning (RL) based scheduler for archive storage systems. By identifying the unique characteristics of archive storage service, we design a state space and reward function for the RL agent. MasterPlan includes a continuous action encoding approach to guarantee efficient exploration, and a meta adaptation module to extract features of workload series. Experiments show that MasterPlan can achieve 1.25× throughput, 2.16× 99 th latency and 1.47× power draw improvement compared to existing solutions. Xinqi Chen, Erci Xu, Dengyao Mo, Ruiming Lu, Dian Ding, Guangtao Xue |
ACM Trans. Archit. Code Optim. | 4 |
| 2024 | CSAL: the Next-Gen Local Disks for the CloudabstractCloud local disks are attractive for their affordable price and high performance. The recent advancement in CPUs motivates cloud vendors to further multiplex the computing resources to serve more users. Unfortunately, such proposals are constrained by the limited offerings of cloud local disks per server as the underlying storage devices are either large but slow (e.g., HDDs) or fast yet small (e.g., NVMe SSDs). Yanbo Zhou, Erci Xu, Kapil Karkra, Mariusz Barczak, Wayne Gao, Wojciech Malikowski, Mateusz Kozlowski, Lukasz Lasek, Ruiming Lu, Lilong Huang, Keqiang Niu, Jiaji Zhu, Jiesheng Wu |
EuroSys | 10 |
| 2023 | Perseus: A Fail-Slow Detection Framework for Cloud Storage Systems
Ruiming Lu, Erci Xu, Yiming Zhang 0003, Fengyi Zhu, Zhaosheng Zhu, Mengtian Wang, Zongpeng Zhu, Guangtao Xue, Jiwu Shu, Minglu Li 0001, Jiesheng Wu |
FAST | 1 |
| 2023 | From Missteps to Milestones: A Journey to Practical Fail-Slow DetectionabstractThe newly emerging “fail-slow” failures plague both software and hardware where the victim components are still functioning yet with degraded performance. To address this problem, this article presents Perseus , a practical fail-slow detection framework for storage devices. Perseus leverages a light regression-based model to quickly pinpoint and analyze fail-slow failures at the granularity of drives. Within a 10-month close monitoring on 248K drives, Perseus managed to find 304 fail-slow cases. Isolating them can reduce the (node-level) 99.99th tail latency by 48%. We assemble a large-scale fail-slow dataset (including 41K normal drives and 315 verified fail-slow drives) from our production traces, based on which we provide root cause analysis on fail-slow drives covering a variety of ill-implemented scheduling, hardware defects, and environmental factors. We have released the dataset to the public for fail-slow study. Ruiming Lu, Erci Xu, Yiming Zhang 0003, Fengyi Zhu, Zhaosheng Zhu, Mengtian Wang, Zongpeng Zhu, Guangtao Xue, Jiwu Shu, Minglu Li 0001, Jiesheng Wu |
ACM Trans. Storage | 1 |
| 2022 | NVMe SSD Failures in the Field: the Fail-Stop and the Fail-Slow
Ruiming Lu, Erci Xu, Yiming Zhang 0003, Zhaosheng Zhu, Mengtian Wang, Zongpeng Zhu, Guangtao Xue, Minglu Li 0001, Jiesheng Wu |
USENIX ATC | 1 |