EDBT 2026 Demo / reviewers in the wild / expert
Ranhao Jia
dblp:303/3807
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FIFO-MEP: An Efficient Multi-Eviction-Point FIFO Cache with Stable Demotion for Burst-Oriented Access MitigationabstractCaching technology is widely used in multiple areas particularly in distributed computing, where its performance is highly dependent on the cache efficiency. The cache eviction algorithm serves as the core component of a cache, primarily aimed at improving cache efficiency by reducing the cache miss ratio. Numerous eviction algorithms are proposed in recent decades and state-of-the-art methods tend to adopt lazy promotion and quick demotion designs. Lazy promotion simplifies cache-hit operations for higher throughput, while quick demotion effectively filters the low-popularity objects. However, the two designs either fail to identify burst objects or suffer from unstable demotion precision. In order to address the above problems, we propose FIFO-MEP, an efficient FIFO cache with Multiple Eviction Points. The key design of FIFO-MEP is to introduce multiple fixed-position eviction points near the head of a FIFO queue. These eviction points enable repeated inspections of objects, leading to effective identification of burst objects. Meanwhile, by fixing positions of these eviction points, FIFO-MEP delivers stable demotion precision. We implement FIFO-MEP using libCacheSim and evaluated it on 5439 production traces for three typical cache sizes, and further verify its efficiency based on Memcached. The evaluation results show that FIFO-MEP reduces the miss ratio by an average of 15.8 % across all experimental configurations. Compared to the state-of-the-art S3-FIFO, FIFO-MEP achieves cache efficiency improvement by up to 21.8 % for large cache sizes. Furthermore, FIFO-MEP yields the best performance under 51 % of all tested conditions. Ranhao Jia, Yunfei Gu, Chentao Wu, Jie Li 0002, Minyi Guo, Liqiang Zhang 0010 |
CLUSTER | 1 |
| 2025 | Gaze into the Pattern: Characterizing Spatial Patterns with Internal Temporal Correlations for Hardware PrefetchingabstractHardware prefetching is one of the most widely-used techniques for hiding long data access latency. To address the challenges faced by hardware prefetching, architects have proposed to detect and exploit the spatial locality at the granularity of spatial region. When a new region is activated, they try to find similar previously accessed regions for footprint prediction based on system-level environmental features such as the trigger instruction or data address. However, we find that such context-based prediction cannot capture the essential characteristics of access patterns, leading to limited flexibility, practicality and suboptimal prefetching performance. In this paper, inspired by the temporal property of memory accessing, we note that the temporal correlation exhibited within the spatial footprint is a key feature of spatial patterns. To this end, we propose Gaze, a simple and efficient hardware spatial prefetcher that skillfully utilizes footprint-internal temporal correlations to efficiently characterize spatial patterns. Meanwhile, we observe a unique unresolved challenge in utilizing spatial footprints generated by spatial streaming, which exhibit extremely high access density. Therefore, we further enhance Gaze with a dedicated two-stage approach that mitigates the over-prefetching problem commonly encountered in conventional schemes. Our comprehensive and diverse set of experiments show that Gaze can effectively enhance the performance across a wider range of scenarios. Specifically, Gaze improves performance by $\mathbf{5. 7 \%}$ and 5.4% at single-core, 11.4% and $\mathbf{8. 8 \%}$ at eight-core, compared to most recent low-cost solutions PMP and vBerti. Zixiao Chen, Chentao Wu, Yunfei Gu, Ranhao Jia, Jie Li 0002, Minyi Guo |
HPCA | 4 |
| 2024 | RL-Cache: An Efficient Reinforcement Learning Based Cache Partitioning Approach for Multi-Tenant CDN ServicesabstractContent Delivery Network (CDN) has been widely used to provide data transmission services to end users. The edge cache servers are important components in CDN, and their hit ratios significantly influence the quality of cache service. However, edge caches are shared by multiple tenants (i.e., Internet Content Providers or ICPs) and the resource contention among tenants presents a huge challenge to improve the cache performance. Cache partitioning is a common method to deal with this chal-lenge, and several approaches have been proposed but still have some drawbacks. Existing methods bring non-negligible temporal and spatial overheads while obtaining features. Although some learning based methods have reduced these costs, the learning model convergence is slow due to the large searching space. To address the above problems, we propose a lightweight Reinforcement Learning based Cache Partitioning Approach (RL-Cache), which increases overall hit ratios of edge cache servers in CDN. The core of RL-Cache is a new feature named Compulsory Miss Ratio (CMR). It can be obtained in linear complexity and reflect the tenants' demand of cache space. To demonstrate the effectiveness of our approach, we not only utilize open-source traces from industrial CDNs but also collect real-world workloads from Tencent Cloud CDN. We develop a simulator to conduct several experiments driven by various traces. The experimental results show that compared to the commonly used methods, RL-Cache reduces the upstream traffic by 12.6% on average and improves the hit ratio by up to 4%. Ranhao Jia, Zixiao Chen, Chentao Wu, Jie Li 0002, Minyi Guo, Hongwen Huang |
CLUSTER | 1 |
| 2022 | GRPU: An Efficient Graph-based Cross-Rack Parallel Update Scheme for Cloud Storage SystemsabstractErasure coding (EC) has been widely used in cloud storage systems to provide both high reliability and low storage cost. Previous literatures show that the cross-rack update operations are prevalent for many applications in erasure-coded cloud storage systems, which introduces significant I/O amplification, load imbalance and high latency. Several existing methods have been proposed to mitigate these problems. However, they ignore the correlations among chunks when performing data placement. Thus numerous stripes and racks participate in the update leading to extra I/Os and cross-rack traffic. Moreover, they don’t take into account the parallelism of network transmission which loses the potential update performance gains.To address the issues, we propose a novel Graph-based cross-Rack Parallel Update (GRPU) scheme to improve the update performance for erasure-coded cloud storage systems. The key idea of GRPU is to place the correlated chunks in the same stripe and rack, and transmit the chunks in parallel based on the network distance. The data placement and transmission paths selection are guided by two kinds of graphs. To demonstrate the effectiveness of GRPU, we conduct several experiments in a local cluster. The results show that, compared to the state-of-the-art methods, GRPU reduces the cross-rack traffic by up to 34.66% and the average response time by up to 61.69%, respectively. Ranhao Jia, Haiwei Deng, Yunfei Gu, Huangzhen Xue, Chentao Wu, Jie Li 0002, Guangtao Xue, Minyi Guo |
ICCD | 1 |
| 2022 | PRM: An Efficient Partial Recovery Method to Accelerate Training Data Reconstruction for Distributed Deep Learning Applications in Cloud Storage SystemsabstractDistributed deep learning is a typical machine learning method running in distributed environment such as cloud computing systems. The corresponding training, validation and test datasets are very large in general (e.g., several TBs), which need to be stored across multiple data nodes. Due to the high disk failure ratio in cloud storage systems, one of the critical issues for distributed deep learning is how to efficiently tolerate disk failures in the training procedures. These failures can lead to a large amount of data loss, which decreases the training accuracy and slows down the training process. Although several recovery methods are proposed to accelerate the data reconstruction, the related overhead is extremely high, such as high CPU/GPU utilization, a large number of I/Os, etc.To address the above problems, we propose a novel Partial-Recovery Method (called PRM) , which is an adaptive recovery method to accelerate data reconstruction for distributed deep learning applications in cloud storage systems. The key idea of PRM is combining the advantages of erasure coding’s ability to obtain global information on the data distribution with the AI’s ability to recover partial lost data, which can sharply reduce the overhead with acceptable training accuracy. To demonstrate the effectiveness of the PRM approach, we conduct several experiments. The results show that, compared to the state-of-the-art full or approximate recovery methods, PRM decreases the average network transmission time overhead by up to 64.50%, and reduces the recovery time by up to 55.90%, respectively. Piao Hu, Yunfei Gu, Ranhao Jia, Chentao Wu, Minyi Guo, Jie Li 0002 |
IWQoS | 3 |
| 2021 | A Graph-Assisted Out-of-Place Update Scheme for Erasure Coded Storage SystemsabstractErasure Codes (ECs) have widely been used in distributed storage systems to ensure data availability because of its low storage cost and high reliability. However, the update operations in erasure coded storage systems can bring extremely high I/O latency and load imbalance due to the complexity of relationships between data and parity blocks. Although several methods such as Parity Logging (PL) and Log-Structured Array (LSA) have been proposed to improve the performance of updates, they either bring extra I/O operations or decrease the performance of file access. Haiwei Deng, Ranhao Jia, Chentao Wu |
ICPP | 2 |