VLDB 2026 Research / reviewers in the wild / expert
Ji Zhang 0010
dblp:86/1953-10
· DBLP profile ↗
15ranked-venue papers
6as first author
10since 2021 · last 2024
0000-0002-3770-1463ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Speal: Achieving a More Accurate Model with Less Training Data in Performance Evaluation of Storage System through Sampling Optimization
Liang Bao, Hua Wang 0008, Ke Zhou 0001, Ji Zhang 0010, Xi Peng 0006, Renhai Chen, Gong Zhang 0001 |
DASFAA (2) | 5 |
| 2024 | HyperDB: a Novel Key Value Store for Reducing Background Traffic in Heterogeneous SSD StorageabstractLog-structured merge tree (LSM-tree) has been widely adopted by modern key-value stores. Deploying LSM-tree across heterogeneous SSD storage which combines the fast but expensive NVMe storage tier with the slow but economical SATA storage tier has emerged as the optimal choice for maximizing cost-effectiveness. However, existing studies typically focus on optimizing the performance of individual storage layers, thereby impeding the full utilization potential of both storage layers. We notice that they tend to over-rely on one storage layer and underutilize the other. In this paper, we present HyperDB, a novel hybrid key-value store designed to enhance the overall performance of both layers via deploying tailored data structures in different media. Especially, HyperDB devises a zone-based data layout for NVMe SSDs to reduce migration overhead, while also implementing a semi-sorted table on the SATA storage layer to minimize merge overhead. Furthermore, we propose a preemptive compaction method at the block-granularity level to further alleviate resource consumption caused by background compaction. Experimental results show that HyperDB achieves 2.25 × faster on average throughput and a 60.3% reduction in background task traffic, compared to the standard use of RocksDB in data centers today. Ruisong Zhou, Yuzhan Zhang, Chunhua Li 0002, Ke Zhou 0001, Peng Wang 0037, Gong Zhang 0001, Ji Zhang 0010 |
ICPP | 7 |
| 2024 | An optimized learning-based directory placement policy with two-rounds selection in distributed file systems
Yuanzhang Wang, Fengkui Yang, Ke Zhou 0001, Chunhua Li 0002, Ji Zhang 0010 |
Future Gener. Comput. Syst. | 6 |
| 2024 | SLAP: Segmented Reuse-Time-Label Based Admission Policy for Content Delivery Network Cachingabstract‘‘Learned” admission policies have shown promise in improving Content Delivery Network (CDN) cache performance and lowering operational costs. Unfortunately, existing learned policies are optimized with a few fixed cache sizes while in reality, cache sizes often vary over time in an unpredictable manner. As a result, existing solutions cannot provide consistent benefits in production settings. We present SLAP , a learned CDN cache admission approach based on segmented object reuse time prediction. SLAP predicts an object’s reuse time range using the Long-Short-Term-Memory model and admits objects that will be reused (before eviction) given the current cache size. SLAP decouples model training from cache size, allowing it to adapt to arbitrary sizes. The key to our solution is a novel segmented labeling scheme that makes SLAP without requiring precise prediction on object reuse time. To further make SLAP a practical and efficient solution, we propose aggressive reusing of computation and training on sampled traces to optimize model training, and a specialized predictor architecture that overlaps prediction computation with miss object fetching to optimize model inference. Our experiments using production CDN traces show that SLAP achieves significantly lower write traffic (38%-59%), longer SSDs lifetime (104%-178%), a consistently higher hit rate (3.2%-11.7%), and requires no effort to adapt to changing cache sizes, outperforming existing policies. Ke Liu 0014, Hua Wang 0008, Ke Zhou 0001, Peng Wang 0037, Ji Zhang 0010 |
ACM Trans. Archit. Code Optim. | 6 |
| 2023 | SLAP: An Adaptive, Learned Admission Policy for Content Delivery Network Cachingabstract"Learned" admission policies have shown promise in improving Content Delivery Network (CDN) cache performance and lowering operational costs. Unfortunately, existing learned policies are optimized with a few fixed cache sizes while in reality, cache sizes often vary over time in an unpredictable manner. As a result, existing solutions cannot provide consistent benefits in production settings.We present SLAP, a learned CDN cache admission approach based on segmented object reuse time prediction. SLAP predicts an object’s reuse time range using the Long-Short-Term-Memory model and admits objects that will be reused (before eviction) given the current cache size. SLAP separates model training from cache size, allowing it to adapt to arbitrary sizes. The key to our solution is a novel segmented labeling scheme that enables SLAP to precisely predict object reuse time. To further make SLAP a practical and efficient solution, we propose aggressive reusing of computation and training on sampled traces to optimize model training, and a specialized predictor architecture that overlaps prediction computation with miss object fetching to optimize model inference. Our experiments with production CDN traces show that SLAP achieves significantly lower write traffic (38%-59%), longer SSDs service life (104%-178%), a consistently higher hit rate (3.2%-11.7%), and requires no effort to adapt to changing cache sizes, outperforming existing policies. Ke Liu 0014, Hua Wang 0008, Ke Zhou 0001, Ji Zhang 0010 |
IPDPS | 5 |
| 2023 | SPAE: Lifelong disk failure prediction via end-to-end GAN-based anomaly detection with ensemble update
Yu Liu 0040, Yunchuan Guan, Tianming Jiang, Ke Zhou 0001, Hua Wang 0008, Guangxing Hu, Ji Zhang 0010, Ping Huang 0001 |
Future Gener. Comput. Syst. | 7 |
| 2022 | SS-LRU: a smart segmented LRU cachingabstractMany caching policies use machine learning to predict data reuse, but they ignore the impact of incorrect prediction on cache performance, especially for large-size objects. In this paper, we propose a smart segmented LRU (SS-LRU) replacement policy, which adopts a size-aware classifier designed for cache scenarios and considers the cache cost caused by misprediction. Besides, SS-LRU enhances the migration rules of segmented LRU (SLRU) and implements a smart caching with unequal priorities and segment sizes based on prediction and multiple access patterns. We conducted Extensive experiments under the real-world workloads to demonstrate the superiority of our approach over state-of-the-art caching policies. Chunhua Li 0002, Man Wu, Ke Zhou 0001, Ji Zhang 0010, Yunqing Sun |
DAC | 5 |
| 2022 | A Data-aware Learned Index Scheme for Efficient WritesabstractIndex structure is very important for efficient data access and system performance in the storage system. Learned index utilizes recursive index models to replace range index structure (such as B+ Tree) so as to predict the position of a lookup key in a dataset. This new paradigm greatly reduces query time and index size, however it only supports read-only workloads. Although some studies reserve gaps between keys for new data to support update, they incur high memory space and shift cost when a large number of data are inserted. Li Liu 0047, Chunhua Li 0002, Ke Zhou 0001, Ji Zhang 0010 |
ICPP | 6 |
| 2022 | LDPP: A Learned Directory Placement Policy in Distributed File SystemsabstractLoad balance is a critical problem in distributed file systems. Previous works focus on how to distribute data evenly on different nodes or storage devices from the perspective of file level, but neglect to effectively take advantage of the directory’s locality and the long duration of the directory’s hotness, which may affect the degree of balance and cause performance degradation. To overcome this shortcoming, in this paper, we propose a learning-based directory placement policy, called LDPP, which determines the data layout by predicting the load. We first establish a relationship between directory request characteristics and state information to predict the state information of the directory (storage capacity, bandwidth, and IOPS). Then, the new directory is placed on different nodes in a multi-dimensional manner based on the Manhattan distance according to the predicted multidimensional state information. In addition, we also take into account the trade-off between the same category directory classified by the load prediction module and the peer directories and explore their influence on the balance. Extensive experiments demonstrate that LDPP not only efficiently alleviates load imbalance and increases the utilization of the resources but also improves DFS performance in practice, which can reduce service latency by up to 36 and increase IOPS and bandwidth by 8 and 9, respectively. Yuanzhang Wang, Fengkui Yang, Ji Zhang 0010, Chunhua Li 0002, Ke Zhou 0001, Jinhu Liu |
ICPP | 3 |
| 2021 | $\hbox {CDBTune}^{+}$: An efficient deep reinforcement learning-based automatic cloud database tuning systemabstractAbstract Configuration tuning is vital to optimize the performance of a database management system (DBMS). It becomes more tedious and urgent for cloud databases (CDB) due to diverse database instances and query workloads, which make the job of a database administrator (DBA) very difficult. Existing solutions for automatic DBMS configuration tuning have several limitations. Firstly, they adopt a pipelined learning model but cannot optimize the overall performance in an end-to-end manner. Secondly, they rely on large-scale high-quality training samples which are hard to obtain. Thirdly, existing approaches cannot recommend reasonable configurations for a large number of knobs to tune whose potential values live in such high-dimensional continuous space. Lastly, in cloud environments, existing approaches can hardly cope with the changes of hardware configurations and workloads, and have poor adaptability. To address these challenges, we design an end-to-end automatic CDB tuning system, $${\texttt {CDBTune}}^{+}$$ CDBTune + , using deep reinforcement learning (RL). $${\texttt {CDBTune}}^{+}$$ CDBTune + utilizes the deep deterministic policy gradient method to find the optimal configurations in a high-dimensional continuous space. $${\texttt {CDBTune}}^{+}$$ CDBTune + adopts a trial-and-error strategy to learn knob settings with a limited number of samples to accomplish the initial training, which alleviates the necessity of collecting a massive amount of high-quality samples. $${\texttt {CDBTune}}^{+}$$ CDBTune + adopts the reward-feedback mechanism in RL instead of traditional regression, which enables end-to-end learning and accelerates the convergence speed of our model and improves the efficiency of online tuning. Besides, we propose effective techniques to improve the training and tuning efficiency of $${\texttt {CDBTune}}^{+}$$ CDBTune + for practical usage in a cloud environment. We conducted extensive experiments under 7 different workloads on real cloud databases to evaluate $${\texttt {CDBTune}}^{+}$$ CDBTune + . Experimental results showed that $${\texttt {CDBTune}}^{+}$$ CDBTune + adapts well to a new hardware environment or workload, and significantly outperformed the state-of-the-art tuning tools and DBA experts. Ji Zhang 0010, Ke Zhou 0001, Guoliang Li 0001, Yu Liu 0040, Jiashu Xing |
VLDB J. | 1 |
| 2020 | Tier-Scrubbing: An Adaptive and Tiered Disk Scrubbing Scheme with Improved MTTD and Reduced CostabstractSector errors are a common type of error in modern disks. A sector error that occurs during I/O operations might cause inaccessibility of an application. Even worse, it could result in permanent data loss if the data is being reconstructed, and thereby severely affects the reliability of a storage system. Many disk scrubbing schemes have been proposed to solve this problem. However, existing approaches have several limitations. First, schemes use machine learning (ML) to predict latent sector errors (LSEs), but only leverage a single snapshot of training data to make a prediction, and thereby ignore sequential dependencies between different statuses of a hard disk over time. Second, they accelerate the scrubbing at a fixed rate based on the results of a binary classification model, which may result in unnecessary increases in scrubbing cost. Third, they naively accelerate the scrubbing of the full disk which has LSEs based on the predictive results, but neglect partial high-risk areas (the areas that have a higher probability of encountering LSEs). Lastly, they do not employ strategies to scrub these high-risk areas in advance based on I/O accesses patterns, in order to further increase the efficiency of scrubbing.We address these challenges by designing a Tier-Scrubbing (TS) scheme that combines a Long Short-Term Memory (LSTM) based Adaptive Scrubbing Rate Controller (ASRC), a module focusing on sector error locality to locate high-risk areas in a disk, and a piggyback scrubbing strategy to improve the reliability of a storage system. Our evaluation results on realistic datasets and workloads from two real world data centers demonstrate that TS can simultaneously decrease the Mean-Time-To-Detection (MTTD) by about 80% and the scrubbing cost by 20%, compared to a state-of-the-art scrubbing scheme. Ji Zhang 0010, Yuanzhang Wang, Yangtao Wang, Ke Zhou 0001, Sebastian Schelter, Ping Huang 0001, Yong-guang Ji |
DAC | 1 |
| 2020 | HDDse: Enabling High-Dimensional Disk State Embedding for Generic Failure Detection System of Heterogeneous Disks in Large Data Centers
Ji Zhang 0010, Ping Huang 0001, Ke Zhou 0001, Sebastian Schelter |
USENIX ATC | 1 |
| 2020 | Minority Disk Failure Prediction Based on Transfer Learning in Large Data Centers of Heterogeneous Disk SystemsabstractThe storage system in large scale data centers is typically built upon thousands or even millions of disks, where disk failures constantly happen. A disk failure could lead to serious data loss and thus system unavailability or even catastrophic consequences if the lost data cannot be recovered. While replication and erasure coding techniques have been widely deployed to guarantee storage availability and reliability, disk failure prediction is gaining popularity as it has the potential to prevent disk failures from occurring in the first place. Recent trends have turned toward applying machine learning approaches based on disk SMART attributes for disk failure predictions. However, traditional machine learning (ML) approaches require a large set of training data in order to deliver good predictive performance. In large-scale storage systems, new disks enter gradually to augment the storage capacity or to replace failed disks, leading storage systems to consist of small amounts of new disks from different vendors and/or different models from the same vendor as time goes on. We refer to this relatively small amount of disks as minority disks. Due to the lack of sufficient training data, traditional ML approaches fail to deliver satisfactory predictive performance in evolving storage systems which consist of heterogeneous minority disks. To address this challenge and improve the predictive performance for minority disks in large data centers, we propose a minority disk failure prediction model named TLDFP based on a transfer learning approach. Our evaluation results in two realistic datasets have demonstrated that TLDFP can deliver much more precise results and lower additional maintenance cost, compared to four popular prediction models based on traditional ML algorithms and two state-of-the-art transfer learning methods. Ji Zhang 0010, Ke Zhou 0001, Ping Huang 0001, Xubin He, Yong-guang Ji, Yinhu Wang |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2019 | Transfer Learning based Failure Prediction for Minority Disks in Large Data Centers of Heterogeneous Disk SystemsabstractThe storage system in large scale data centers is typically built upon thousands or even millions of disks, where disk failures constantly happen. A disk failure could lead to serious data loss and thus system unavailability or even catastrophic consequences if the lost data cannot be recovered. While replication and erasure coding techniques have been widely deployed to guarantee storage availability and reliability, disk failure prediction is gaining popularity as it has the potential to prevent disk failures from occurring in the first place. Recent trends have turned toward applying machine learning approaches based on disk SMART attributes for disk failure predictions. However, traditional machine learning (ML) approaches require a large set of training data in order to deliver good predictive performance. In large-scale storage systems, new disks enter gradually to augment the storage capacity or to replace failed disks, leading storage systems to consist of small amounts of new disks from different vendors and/or different models from the same vendor as time goes on. We refer to this relatively small amount of disks as minority disks. Due to the lack of sufficient training data, traditional ML approaches fail to deliver satisfactory predictive performance in evolving storage systems which consist of heterogeneous minority disks. To address this challenge and improve the predictive performance for minority disks in large data centers, we propose a minority disk failure prediction model named TLDFP based on a transfer learning approach. Our evaluation results on two realistic datasets have demonstrated that TLDFP can deliver much more precise results, compared to four popular prediction models based on traditional ML algorithms and two state-of-the-art transfer learning methods. Ji Zhang 0010, Ke Zhou 0001, Ping Huang 0001, Xubin He, Zhili Xiao, Yong-guang Ji, Yinhu Wang |
ICPP | 1 |
| 2019 | An End-to-End Automatic Cloud Database Tuning System Using Deep Reinforcement LearningabstractConfiguration tuning is vital to optimize the performance of database management system (DBMS). It becomes more tedious and urgent for cloud databases (CDB) due to the diverse database instances and query workloads, which make the database administrator (DBA) incompetent. Although there are some studies on automatic DBMS configuration tuning, they have several limitations. Firstly, they adopt a pipelined learning model but cannot optimize the overall performance in an end-to-end manner. Secondly, they rely on large-scale high-quality training samples which are hard to obtain. Thirdly, there are a large number of knobs that are in continuous space and have unseen dependencies, and they cannot recommend reasonable configurations in such high-dimensional continuous space. Lastly, in cloud environment, they can hardly cope with the changes of hardware configurations and workloads, and have poor adaptability. To address these challenges, we design an end-to-end automatic CDB tuning system, CDBTune, using deep reinforcement learning (RL). CDBTune utilizes the deep deterministic policy gradient method to find the optimal configurations in high-dimensional continuous space. CDBTune adopts a try-and-error strategy to learn knob settings with a limited number of samples to accomplish the initial training, which alleviates the difficulty of collecting massive high-quality samples. CDBTune adopts the reward-feedback mechanism in RL instead of traditional regression, which enables end-to-end learning and accelerates the convergence speed of our model and improves efficiency of online tuning. We conducted extensive experiments under 6 different workloads on real cloud databases to demonstrate the superiority of CDBTune. Experimental results showed that CDBTune had a good adaptability and significantly outperformed the state-of-the-art tuning tools and DBA experts. Ji Zhang 0010, Yu Liu 0040, Ke Zhou 0001, Guoliang Li 0001, Zhili Xiao, Jiashu Xing, Yangtao Wang, Tianheng Cheng, Li Liu 0047, Minwei Ran, Zekang Li |
SIGMOD Conference | 1 |