Jiongzhou Liu

dblp:291/3638 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2022 An In-Depth Correlative Study Between DRAM Errors and Server Failures in Production Data Centers
abstract
Dynamic Random Access Memory (DRAM) errors are prevalent and lead to server failures in production data centers. However, little is known about the correlation between DRAM errors and server failures in state-of-the-art field studies on DRAM error measurement. To fill this void, we present an in-depth data-driven correlative analysis between DRAM errors and server failures, with the primary goal of predicting server failures based on DRAM error characterization and hence enabling proactive reliability maintenance for production data centers. Our analysis is based on an eight-month dataset collected from over three million memory modules in the production data centers at Alibaba. We find that the correctable DRAM errors of most server failures only manifest within a short time before the failures happen, implying that server failure prediction should be conducted regularly at short time intervals for accurate prediction. We also study various impacting factors (including component failures in the memory subsystem, DRAM configurations, types of correctable DRAM errors) on server failures. Furthermore, we design a machine-learning-based server failure prediction workflow and demonstrate the feasibility of server failure prediction based on DRAM error characterization. To this end, we report 14 findings from our measurement and prediction studies.
Zhinan Cheng, Shujie Han 0001, Patrick P. C. Lee, Jiongzhou Liu
SRDS5
2021 General Feature Selection for Failure Prediction in Large-scale SSD Deployment
abstract
Solid-state drive (SSD) failures are likely to cause system-level failures leading to downtime, enabling SSD failure prediction to be critical to large-scale SSD deployment. Existing SSD failure prediction studies are mostly based on customized SSDs with proprietary monitoring metrics, which are difficult to reproduce. To support general SSD failure prediction of different drive models and vendors, this paper proposes Wear-out-updating Ensemble Feature Ranking (WEFR) to select the SMART attributes as learning features in an automated and robust manner. WEFR combines different feature ranking results and automatically generates the final feature selection based on the complexity measures and the change point detection of wear-out degrees. We evaluate our approach using a dataset of nearly 500K working SSDs at Alibaba. Our results show that the proposed approach is effective and outperforms related approaches. We have successfully applied the proposed approach to improve the reliability of cloud storage systems in production SSD-based data centers. We release our dataset for public use.
Shujie Han 0001, Patrick P. C. Lee, Jiongzhou Liu
DSN6
2021 An In-Depth Study of Correlated Failures in Production SSD-Based Data Centers
Shujie Han 0001, Patrick P. C. Lee, Jiongzhou Liu
FAST6
2021 Adversarial Domain Adaptation with Correlation-Based Association Networks for Longitudinal Disk Fault Prediction
abstract
Disk fault is known to be the key cause of data loss in the modern large-scale data center, which affects the reliability and stability of the server and even the whole IT infrastructure, resulting in high financial cost. Recent works on disk fault prediction demonstrate the ability of machine learning techniques in the early prediction of disk failure. However, two limitations hinder the real-world application of current methods. First, they ignore the data heterogeneity in the data center, where distribution shifts commonly exist across different disk model types that decrease the model's performance. Second, the number of disks of different disk model types varies greatly in the data center, and current methods failed to deliver an acceptable performance over disk model types with few training samples. To address these limitations, we propose adversarial domain adaptation with correlation-based association networks (ADA-CBAN) to both mitigate the distribution shift problem and also to boost the performance on small-scaled data. Extensive experiments prove that our proposed model is effective and achieves new state-of-the-art. In addition, our post model analysis can also reveal important feature interactions and highlight the crucial period before disk faults, where both are useful information in real-world applications.
Xiang Lan 0004, Dianwen Ng, Jiongzhou Liu, Mengling Feng
IJCNN4