Yong-guang Ji

dblp:27/4062 · also Yongguang Ji · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 since 2021Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Storage systems · 72% Memory systems · 17% Hardware accelerators and domain-specific architectures · 8%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems
storage reliability
1.032020
Minority Disk Failure Prediction Based on Transfer Learning in Large Data Centers of Heterogeneous Disk Systems · IEEE Trans. Parallel Distributed Syst. 2020
Tier-Scrubbing: An Adaptive and Tiered Disk Scrubbing Scheme with Improved MTTD and Reduced Cost · DAC 2020
OSCA: An Online-Model Based Cache Allocation Scheme in Cloud Block Storage Systems · USENIX ATC 2020
Storage systems › storage architecture › block storage
cloud block storage
0.622020
OSCA: An Online-Model Based Cache Allocation Scheme in Cloud Block Storage Systems · USENIX ATC 2020
Efficient SSD Cache for Cloud Block Storage via Leveraging Block Reuse Distances · IEEE Trans. Parallel Distributed Syst. 2020
Memory systems › cache management
cache allocation
0.412020
OSCA: An Online-Model Based Cache Allocation Scheme in Cloud Block Storage Systems · USENIX ATC 2020
Memory systems › cache management
cache replacement
0.412020
Efficient SSD Cache for Cloud Block Storage via Leveraging Block Reuse Distances · IEEE Trans. Parallel Distributed Syst. 2020
Storage systems › storage reliability
disk failure prediction
0.412020
Minority Disk Failure Prediction Based on Transfer Learning in Large Data Centers of Heterogeneous Disk Systems · IEEE Trans. Parallel Distributed Syst. 2020
Storage systems
flash and SSD
0.412020
Efficient SSD Cache for Cloud Block Storage via Leveraging Block Reuse Distances · IEEE Trans. Parallel Distributed Syst. 2020
Storage systems › storage reliability › disk reliability
latent sector errors
0.412020
Tier-Scrubbing: An Adaptive and Tiered Disk Scrubbing Scheme with Improved MTTD and Reduced Cost · DAC 2020
Storage systems › storage reliability
scrubbing
0.412020
Tier-Scrubbing: An Adaptive and Tiered Disk Scrubbing Scheme with Improved MTTD and Reduced Cost · DAC 2020
Storage systems › flash and SSD
SSD cache
0.412020
Efficient SSD Cache for Cloud Block Storage via Leveraging Block Reuse Distances · IEEE Trans. Parallel Distributed Syst. 2020
Hardware accelerators and domain-specific architectures
transfer learning
0.412020
Minority Disk Failure Prediction Based on Transfer Learning in Large Data Centers of Heterogeneous Disk Systems · IEEE Trans. Parallel Distributed Syst. 2020
Machine learning and data management
transfer learning
0.112020
Minority Disk Failure Prediction Based on Transfer Learning in Large Data Centers of Heterogeneous Disk Systems · IEEE Trans. Parallel Distributed Syst. 2020

Methods — techniques the papers use, named apart from their topics

transfer learning · 0.9SMART attributes · 0.9reuse distance analysis · 0.4piggyback scrubbing · 0.4lazy eviction · 0.4adaptive scrubbing rate control · 0.4LSTM · 0.4
YearPublicationVenuePosition
2024 CAMS: A Cost-Aware Migration Scheme for Cloud Object Storage Systems
abstract
Cloud object storage systems provide massive data storage capabilities where data is stored in different storage clusters. Storing data according to access characteristics efficiently in different clusters is a challenging task. Methods considering past data access frequency bring the problem of low storage utilization and load imbalance. We propose a Cost-Aware Migration Scheme for cloud object storage systems(CAMS) based on object hotness and life cycle to improve the utilization of cloud object storage systems and reduce the Total Cost of Ownership(TCO). CAMS establishes an accurate object hotness standard, it uses object hotness and lifecycle prediction to guide data migration. CAMS was tested using real-world datasets from production cloud object storage system, the results show that CAMS strategies outperform Cold, CoinFlip and RejectX strategies with gains of up to 19.79% on the estimated TCO.
Ke Zhou 0001, Hua Wang 0008, Han Kong, Yong-guang Ji
NAS6
2020 Tier-Scrubbing: An Adaptive and Tiered Disk Scrubbing Scheme with Improved MTTD and Reduced Cost
abstract
Sector errors are a common type of error in modern disks. A sector error that occurs during I/O operations might cause inaccessibility of an application. Even worse, it could result in permanent data loss if the data is being reconstructed, and thereby severely affects the reliability of a storage system. Many disk scrubbing schemes have been proposed to solve this problem. However, existing approaches have several limitations. First, schemes use machine learning (ML) to predict latent sector errors (LSEs), but only leverage a single snapshot of training data to make a prediction, and thereby ignore sequential dependencies between different statuses of a hard disk over time. Second, they accelerate the scrubbing at a fixed rate based on the results of a binary classification model, which may result in unnecessary increases in scrubbing cost. Third, they naively accelerate the scrubbing of the full disk which has LSEs based on the predictive results, but neglect partial high-risk areas (the areas that have a higher probability of encountering LSEs). Lastly, they do not employ strategies to scrub these high-risk areas in advance based on I/O accesses patterns, in order to further increase the efficiency of scrubbing.We address these challenges by designing a Tier-Scrubbing (TS) scheme that combines a Long Short-Term Memory (LSTM) based Adaptive Scrubbing Rate Controller (ASRC), a module focusing on sector error locality to locate high-risk areas in a disk, and a piggyback scrubbing strategy to improve the reliability of a storage system. Our evaluation results on realistic datasets and workloads from two real world data centers demonstrate that TS can simultaneously decrease the Mean-Time-To-Detection (MTTD) by about 80% and the scrubbing cost by 20%, compared to a state-of-the-art scrubbing scheme.
Ji Zhang 0010, Yuanzhang Wang, Yangtao Wang, Ke Zhou 0001, Sebastian Schelter, Ping Huang 0001, Yong-guang Ji
DAC8
2020 A Machine Learning Based Write Policy for SSD Cache in Cloud Block Storage
abstract
Nowadays, SSD cache plays an important role in cloud storage systems. The associated write policy, which enforces an admission control policy regarding filling data into the cache, has a significant impact on the performance of the cache system and the amount of write traffic to SSD caches. Based on our analysis on a typical cloud block storage system, approximately 47.09% writes are write-only, i.e., writes to the blocks which have not been read during a certain time window. Naively writing the write-only data to the SSD cache unnecessarily introduces a large number of harmful writes to the SSD cache without any contribution to cache performance. On the other hand, it is a challenging task to identify and filter out those write-only data in a real-time manner, especially in a cloud environment running changing and diverse workloads.In this paper, to alleviate the above cache problem, we propose an ML-WP, Machine Learning Based Write Policy, which reduces write traffic to SSDs by avoiding writing write-only data. The main challenge in this approach is to identify write-only data in a real-time manner. To realize ML-WP and achieve accurate write-only data identification, we use machine learning methods to classify data into two groups (i.e., write-only and normal data). Based on this classification, the write-only data is directly written to backend storage without being cached. Experimental results show that, compared with the industry widely deployed write-back policy, ML-WP decreases write traffic to SSD cache by 41.52%, while improving the hit ratio by 2.61% and reducing the average read latency by 37.52%.
Yu Zhang 0101, Ke Zhou 0001, Ping Huang 0001, Hua Wang 0008, Jianying Hu, Yangtao Wang, Yong-guang Ji
DATE7
2020 OSCA: An Online-Model Based Cache Allocation Scheme in Cloud Block Storage Systems
Yu Zhang 0101, Ping Huang 0001, Ke Zhou 0001, Hua Wang 0008, Jianying Hu, Yong-guang Ji
USENIX ATC6
2020 Minority Disk Failure Prediction Based on Transfer Learning in Large Data Centers of Heterogeneous Disk Systems
abstract
The storage system in large scale data centers is typically built upon thousands or even millions of disks, where disk failures constantly happen. A disk failure could lead to serious data loss and thus system unavailability or even catastrophic consequences if the lost data cannot be recovered. While replication and erasure coding techniques have been widely deployed to guarantee storage availability and reliability, disk failure prediction is gaining popularity as it has the potential to prevent disk failures from occurring in the first place. Recent trends have turned toward applying machine learning approaches based on disk SMART attributes for disk failure predictions. However, traditional machine learning (ML) approaches require a large set of training data in order to deliver good predictive performance. In large-scale storage systems, new disks enter gradually to augment the storage capacity or to replace failed disks, leading storage systems to consist of small amounts of new disks from different vendors and/or different models from the same vendor as time goes on. We refer to this relatively small amount of disks as minority disks. Due to the lack of sufficient training data, traditional ML approaches fail to deliver satisfactory predictive performance in evolving storage systems which consist of heterogeneous minority disks. To address this challenge and improve the predictive performance for minority disks in large data centers, we propose a minority disk failure prediction model named TLDFP based on a transfer learning approach. Our evaluation results in two realistic datasets have demonstrated that TLDFP can deliver much more precise results and lower additional maintenance cost, compared to four popular prediction models based on traditional ML algorithms and two state-of-the-art transfer learning methods.
Ji Zhang 0010, Ke Zhou 0001, Ping Huang 0001, Xubin He, Yong-guang Ji, Yinhu Wang
IEEE Trans. Parallel Distributed Syst.7
2020 Efficient SSD Cache for Cloud Block Storage via Leveraging Block Reuse Distances
abstract
Solid State Drives (SSDs) are popularly used for caching in large scale cloud storage systems nowadays. Traditionally, most cache algorithms make replacement upon each miss when cache space is full. However, we observe that in a typical Cloud Block Storage (CBS) system, there is a great percentage of blocks with large reuse distances, which would result in large number of blocks being evicted out of the cache before they ever have a chance to be referenced while they are cached, significantly jeopardizing the cache efficiency. In this article, we propose LEA, Lazy Eviction cache Algorithm, for cloud block storage to efficiently remedy the cache inefficiencies caused by cache blocks with large reuse distances. LEA mainly employs two lists, Lazy Eviction List (LEL) and Block Identity List (BIL), which keep track of two types of victim blocks respectively based on their cache duration when replacements occur, to improve cache efficiency. When a cache miss happens, if the victim block has not resided in cache for longer than its reuse distance, LEA inserts the missed block identity into BIL. Otherwise, it inserts the missed block entry into LEL. We have evaluated LEA by using IO traces collected from Tencent, one of the largest network service providers in the world, and several open source traces. Experimental results show that LEA not only outperforms most of the state-of-the-art cache algorithms in hit ratio, but also greatly reduces the number of SSD writes.
Ke Zhou 0001, Yu Zhang 0101, Ping Huang 0001, Hua Wang 0008, Yong-guang Ji
IEEE Trans. Parallel Distributed Syst.5
2019 Transfer Learning based Failure Prediction for Minority Disks in Large Data Centers of Heterogeneous Disk Systems
abstract
The storage system in large scale data centers is typically built upon thousands or even millions of disks, where disk failures constantly happen. A disk failure could lead to serious data loss and thus system unavailability or even catastrophic consequences if the lost data cannot be recovered. While replication and erasure coding techniques have been widely deployed to guarantee storage availability and reliability, disk failure prediction is gaining popularity as it has the potential to prevent disk failures from occurring in the first place. Recent trends have turned toward applying machine learning approaches based on disk SMART attributes for disk failure predictions. However, traditional machine learning (ML) approaches require a large set of training data in order to deliver good predictive performance. In large-scale storage systems, new disks enter gradually to augment the storage capacity or to replace failed disks, leading storage systems to consist of small amounts of new disks from different vendors and/or different models from the same vendor as time goes on. We refer to this relatively small amount of disks as minority disks. Due to the lack of sufficient training data, traditional ML approaches fail to deliver satisfactory predictive performance in evolving storage systems which consist of heterogeneous minority disks. To address this challenge and improve the predictive performance for minority disks in large data centers, we propose a minority disk failure prediction model named TLDFP based on a transfer learning approach. Our evaluation results on two realistic datasets have demonstrated that TLDFP can deliver much more precise results, compared to four popular prediction models based on traditional ML algorithms and two state-of-the-art transfer learning methods.
Ji Zhang 0010, Ke Zhou 0001, Ping Huang 0001, Xubin He, Zhili Xiao, Yong-guang Ji, Yinhu Wang
ICPP7
2018 LEA: A Lazy Eviction Algorithm for SSD Cache in Cloud Block Storage
abstract
Solid State Drives (SSDs) are popularly used for caching in large scale cloud storage systems nowadays. Traditionally, most cache algorithms make replacement at each miss when cache space is full. However, we observe that in a typical Cloud Block Storage (CBS), there is a great percentage of blocks with large reuse distances, which would result in large number of blocks being evicted out of the cache before they ever have a chance to be referenced while they are cached, significantly jeopardizing the cache efficiency. In this paper, we propose LEA, Lazy Eviction cache Algorithm, for cloud block storage to efficiently remedy the cache inefficiencies caused by cache blocks with large reuse distances. Specifically, LEA uses two lists, Lazy Eviction List (LEL) and Block Identity List (BIL). When a cache miss happens, if the candidate evicted-block has not resided in cache for longer than its reuse distance, LEA inserts the missed block identity into BIL. Otherwise, it inserts the missed block entry into LEL. We have evaluated LEA by using IO traces collected from Tencent, one of the largest network service providers in the world, and several open source traces. Experimental results show that LEA not only outperforms most of the state-of-the-art cache algorithms in hit ratio, but also reduces the number of SSD writes greatly.
Ke Zhou 0001, Yu Zhang 0101, Ping Huang 0001, Hua Wang 0008, Yong-guang Ji
ICCD5