Qiao Yu 0003

dblp:162/6793-3 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0002-6542-7270ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 M2-MFP: A Multi-Scale and Multi-Level Memory Failure Prediction Framework for Reliable Cloud Infrastructure
abstract
As cloud services become increasingly integral to modern IT infrastructure, ensuring hardware reliability is essential to sustain high-quality service. Memory failures pose a significant threat to overall system stability, making accurate failure prediction through the analysis of memory error logs (i.e., Correctable Errors) imperative. Existing memory failure prediction approaches have notable limitations: rule-based expert models suffer from limited generalizability and low recall rates, while automated feature extraction methods exhibit suboptimal performance. To address these limitations, we propose M2-MFP: a Multi-scale and Multi-Level Memory Failure Prediction framework designed to enhance the reliability and availability of cloud infrastructure. M2-MFP converts correctable errors (CEs) into multi-level binary matrix representations and introduces a Binary Spatial Feature Extractor (BSFE) to automatically extract high-order features at both DIMM-level and bit-level. Building upon the BSFE outputs, we develop a dual-path temporal modeling architecture: 1) a time-patch module that aggregates multi-level features within observation windows, and 2) a time-point module that employs interpretable rule-generation trees trained on bit-level patterns. Experiments on both benchmark datasets and real-world deployment show the superiority of M2-MFP as it outperforms existing state-of-the-art methods by significant margins. Code and data are available at this repository: https://github.com/hwcloud-RAS/M2-MFP.
Hongyi Xie, Min Zhou 0006, Qiao Yu 0003, Jialiang Yu, Zhenli Sheng, Hong Xie 0004, Defu Lian
KDD (2)3
2024 Unveiling DRAM Failures Across Different CPU Architectures in Large-Scale Datacenters
abstract
Memory failures frequently lead to server break-downs in large-scale datacenters, with uncorrectable errors (UEs) serving as primary indicators of defects in Dual Inline Memory Modules (DIMMs). Existing approaches mainly focus on predicting UEs using correctable errors (CEs), but often neglect the relationships of these errors across various CPU architectures, especially in the context of Error Correction Code (ECC). In this paper, we explore the correlation between CEs and UEs across different CPU Architectures, such as x86 and ARM. Our analysis reveals distinctive failure patterns in memory specific to each processor platform. Leveraging Machine Learning (ML) techniques on production datasets, we demonstrate that our approach substantially enhances prediction performance up to 15% in F1-score compared to the existing alaorithm.
Qiao Yu 0003, Jorge Cardoso 0001, Odej Kao
ICDCS1
2023 An Optical Transceiver Reliability Study based on SFP Monitoring and OS-level Metric Data
abstract
The increasing demand for cloud computing drives the expansion in scale of datacenters and their internal optical network, in a strive for increasing bandwidth, high reliability, and lower latency. Optical transceivers are essential elements of optical networks, whose reliability has not been well-studied compared to other hardware components. In this paper, we leverage high quantities of monitoring data from optical transceivers and OS-level metrics to provide statistical insights about the occurrence of optical transceiver failures. We estimate transceiver failure rates and normal operating ranges for monitored attributes, correlate early-observable patterns to known failure symptoms, and finally develop failure prediction models based on our analyses. Our results enable network administrators to deploy early-warning systems and enact predictive maintenance strategies, such as replacement or traffic re-routing, reducing the number of incidents and their associated costs.
Paolo Notaro, Qiao Yu 0003, Soroush Haeri, Jorge Cardoso 0001, Michael Gerndt
CCGrid2
2023 HiMFP: Hierarchical Intelligent Memory Failure Prediction for Cloud Service Reliability
abstract
In large-scale datacenters, memory failure is one of the leading causes of server crashes, and uncorrectable error (UCE) is the major fault type indicating defects of memory modules. Existing approaches tend to predict UCEs using Correctable Errors (CE). However, bit-level CE information has not been completely discussed in previous works and CEs with error bit patterns are strongly correlated with UCE occurrences. In this paper, we present a novel Hierarchical Intelligent Memory Failure Prediction (HiMFP) framework which can predict UCEs on multiple levels of the memory system and associate with memory recovery techniques. Particularly, we leverage CE addresses on multiple levels of memory, especially bit-level, and construct machine learning models based on spatial and temporal CE information. Results of algorithm evaluation using real-world datasets indicate that HiMFP significantly enhances the prediction performance compared with the baseline algorithm. Overall, Virtual Machines (VM) interruptions caused by UCEs can be reduced by around 45% using HiMFP.
Qiao Yu 0003, Wengui Zhang, Paolo Notaro, Soroush Haeri, Jorge Cardoso 0001, Odej Kao
DSN1
2023 Exploring Error Bits for Memory Failure Prediction: An In-Depth Correlative Study
abstract
In large-scale datacenters, memory failure is a common cause of server crashes, with uncorrectable errors (UEs) being a major indicator of Dual Inline Memory Module (DIMM) defects. Existing approaches primarily focus on predicting UEs using correctable errors (CEs), without fully considering the information provided by error bits. However, error bit patterns have a strong correlation with the occurrence of uncorrectable errors (UEs). In this paper, we present a comprehensive study on the correlation between CEs and UEs, specifically emphasizing the importance of spatio-temporal error bit information. Our analysis reveals a strong correlation between spatio-temporal error bits and UE occurrence. Through evaluations using real-world datasets, we demonstrate that our approach significantly improves prediction performance by 15% in F1-score compared to the state-of-the-art algorithms. Overall, our approach effectively reduces the number of virtual machine interruptions caused by UEs by approximately 59%.
Qiao Yu 0003, Wengui Zhang, Jorge Cardoso 0001, Odej Kao
ICCAD1
2022 First CE Matters: On the Importance of Long Term Properties on Memory Failure Prediction
abstract
Dynamic random access memory failures are a threat to the reliability of data centres as they lead to data loss and system crashes. Timely predictions of memory failures allow for taking preventive measures such as server migration and memory replacement. Thereby, memory failure prediction prevents failures from externalizing, and it is a vital task to improve system reliability. In this paper, we revisited the problem of memory failure prediction. We analyzed the correctable errors (CEs) from hardware logs as indicators for a degraded memory state. As memories do not always work with full occupancy, access to faulty memory parts is time distributed. Following this intuition, we observed that important properties for memory failure prediction are distributed through long time intervals. In contrast, related studies, to fit practical constraints, frequently only analyze the CEs from the last fixed-size time interval while ignoring the predating information. Motivated by the observed discrepancy, we study the impact of including the overall (long-range) CE evolution and propose novel features that are calculated incrementally to preserve long-range properties. By coupling the extracted features with machine learning methods, we learn a predictive model to anticipate upcoming failures three hours in advance while improving the average relative precision and recall for 21% and 19% accordingly. We evaluated our methodology on real-world memory failures from the server fleet of a large cloud provider, justifying its validity and practicality.
Jasmin Bogatinovski, Odej Kao, Qiao Yu 0003, Jorge Cardoso 0001
IEEE Big Data3
2020 Using a Set of Triangle Inequalities to Accelerate K-means Clustering
Qiao Yu 0003, Kuan-Hsun Chen, Jian-Jia Chen
SISAP1
2017 Accelerating K-Means by Grouping Points Automatically
Qiao Yu 0003, Bi-Ru Dai
DaWaK1