VLDB 2026 Research / reviewers in the wild / expert
Ruoyi Ruan
dblp:352/6912
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0004-2786-0956ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TierBase: A Workload-Driven Cost-Optimized Key-Value StoreabstractIn the current era of data-intensive applications, the demand for high-performance, cost-effective storage solutions is paramount. This paper introduces a Space-Performance Cost Model for key-value store, designed to guide cost-effective storage configuration decisions. The model quantifies the trade-offs between performance and storage costs, providing a framework for optimizing resource allocation in large-scale data serving environments. Guided by this cost model, we present Tier-Base, a distributed key-value store developed by Ant Group that optimizes total cost by strategically synchronizing data between cache and storage tiers, maximizing resource utilization and effectively handling skewed workloads. To enhance cost-efficiency, TierBase incorporates several optimization techniques, including pre-trained data compression, elastic threading mechanisms, and the utilization of persistent memory. We detail TierBase's architecture, key components, and the implementation of cost optimization strategies. Extensive evaluations using both synthetic benchmarks and real-world workloads demonstrate TierBase's superior cost-effectiveness compared to existing solutions. Furthermore, case studies from Ant Group's production environments showcase TierBase's ability to achieve up to 62% cost reduction in primary scenarios, highlighting its practical impact in large-scale online data serving. Zhitao Shen, Shiyu Yang 0002, Weibo Chen, Kunming Wang 0001, Jiabao Jin, Yuan Su, Xiaoxia Duan, Ruoyi Ruan, Xuemin Lin 0001 |
ICDE | 14 |
| 2024 | Log-based anomaly detection for distributed systems: State of the art, industry experience, and open issuesabstractAbstract Distributed systems have been widely used in many safety‐critical areas. Any abnormalities (e.g., service interruption or service quality degradation) could lead to application crashes or decrease user satisfaction. These things may cause serious economic losses. Among the various quality assurance approaches for distributed systems, log‐based anomaly detection (LAD) has become a popular research topic. Its popularity relates to system logs being able to record and reveal important run‐time information. This paper presents a general LAD framework for distributed systems. Log grouping and feature‐pattern mining are two crucial LAD components that impact on the anomaly‐detection effectiveness. We also present a systematic survey of techniques in these two directions; propose classification frameworks for log grouping and feature patterns; and summarize four log‐grouping techniques and five feature patterns (which refer to invariant relationships among logs that can be used for anomaly detection). To evaluate their applicability, we report on the findings when applying existing techniques to Ray, a popular industrial distributed system. Based on these findings, several open issues are identified, which provide potential guidance for future research and development. Xinjie Wei, Chang-Ai Sun, Dave Towey, Shoufeng Zhang, Wanqing Zuo, Yiming Yu, Ruoyi Ruan, Guyang Song |
J. Softw. Evol. Process. | 8 |
| 2023 | Cougar: A General Framework for Jobs Optimization In CloudabstractIn the cloud environment, different kinds of jobs (Flink, PyTorch, TensorFlow, AI-Serving) are running in the same cluster with different service-level agreements (SLA). To manage large amounts of jobs in a cloud environment efficiently, it is critical to build a system to optimize the job performance in consideration of multiple predefined objectives. For example, one kind of optimization target is improving the resource utilization of jobs, other kinds of objectives are to guarantee the system SLA (e.g., system throughput, response time, and so on). Currently, most of the existing frameworks are working on one aspect of optimization, and can not support different kinds of optimization targets via a uniform framework or system. In Antgroup, we have designed and implemented a general framework (named Cougar) to improve jobs and cluster performance to meet such requirements. Cougar provides the ability to support different optimization scenarios like the initial and runtime optimization for one job, and cross-job optimization for multiple jobs. Nowadays, Cougar has widely used in the production environment of Antgroup including 110,000 jobs and 800,000 Pods daily, and has successfully improved the CPU/Memory/GPU utilization by more than 20% and performance (i.e., throughput or completion time or latency) by around 10%. In the end, we also like to share our best practice on how to tune Flink and Deep Learning Job (GPU collocate) in the production environment. Bo Sang, Shuwei Gu, Xiaojun Zhan, MingJie Tang, Haoyuan Ge, Ke Zhang 0048, Ruoyi Ruan |
ICDE | 10 |
| 2023 | Discovering Parallelisms in Python ProgramsabstractParallelization is a promising way to improve the performance of Python programs. Unfortunately, developers may miss parallelization possibilities, because they usually do not concentrate on parallelization. Many approaches have been proposed to parallelize Python programs automatically, however, they are either domain-specific or require manual annotation. Thus they cannot solve the problem well in general. In this paper, we propose PyPar, an effective tool aiming at discovering parallelization possibilities in real-world Python programs. PyPar doesn’t need manual annotation and is universally applicable. It first drives a data-dependence analysis to determine whether two pieces of code can run concurrently. The key is the use of a graph-theoretic approach. Next, it adopts a dynamic selection strategy to eliminate inefficient parallelisms. Finally, PyPar produces a parallelism report as well as a referential parallelized program, which is built by PyPar using one of the three parallelization methods (thread-based, processbased, and Ray-based). We have implemented a prototype of PyPar and evaluated it on six well-designed widely-used real-world Python packages: Scikit-Image, SciPy, librosa, trimesh, Scikit-learn and seaborn. In total, 1,240 functions are tested, and PyPar found 127 parallelizable functions among them. Based on manual filtering, only 7 of them are false positives (i.e., a 94.5% precision). The remaining 120 are parallelizable (almost 10% among all functions under test), and most of them can be efficiently sped up by gaining an acceleration of up to 90% , with an average of 44%. The acceleration in practice is close to theoretical estimation. The results show that even well-designed practical Python programs can be further parallelized for speeding up, and PyPar can bring effective and efficient parallelization on real-world Python programs. Siwei Wei, Guyang Song, Senlin Zhu, Ruoyi Ruan, Yan Cai 0001 |
ESEC/SIGSOFT FSE | 4 |