EDBT 2026 Demo / reviewers in the wild / expert
Xuhua Ma
dblp:325/4229
· DBLP profile ↗
7ranked-venue papers
0as first author
7since 2021 · last 2026
0009-0008-9852-2427ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LMID: A Comprehensive Multimodal Dataset for Failure Prediction in Cloud Computing SystemsabstractFailure prediction is crucial for ensuring the stability of cloud computing systems and has garnered extensive attention from both academia and industry. Generally, data used for prediction includes two modalities: 1) Text data, such as logs; and, 2) Numerical data, such as error counts and monitoring metrics. However, most existing failure prediction algorithms for cloud computing only focus on a single modality. The lack of high-quality multimodal datasets from real-world production environments constrains academic research on multimodal failure prediction. To fill this gap, this paper releases a large multimodal dataset of operation data from the Alibaba cloud computing platform, namely, Logs and Metrics Integration Dataset (LMID). It consists of 100 million pieces of logs (textual data) and 37 dimensions of monitoring metrics (numerical data) from 220,000 physical machines. To our knowledge, it is the first multimodal dataset for cloud computing system failure prediction, and is expected to greatly benefit the community. This paper provides a detailed introduction to the construction of LMID, its contents, and the performance of state-of-the-art algorithms on it. It also conducts extensive experiments to reveal a new insight that cross-modality connections are effective for failure prediction. LMID is now available at https://huggingface.co/datasets/AliyunECSAlgos/LMID. Lingfei Deng, Ruqiao Xu, Yunong Wang, Xuhua Ma, Dongrui Wu |
KDD (1) | 5 |
| 2026 | Collaborative Prediction of Cloud DRAM Failures With Rules and Machine LearningabstractDRAM faults are the main hardware cause of node unavailability in clouds. To enable early preventive actions and mitigate DRAM fault impacts, prior studies focus on predicting DRAM uncorrectable errors (UEs) that typically cause immediate node unavailability. However, in Alibaba Cloud, we observe that correctable error(CE) storm (numerous CEs occur in a short period) dominates 41% DRAM-caused node unavailability (DCNU). In this paper, we propose to predict DCNU by taking into account both UEs and CE storms. Specifically, observing that DCNUs have strong relevance to the temporal statistics and spatial patterns of CEs, we design novel spatio-temporal features and use soft labels to build a DCNU predictor. We propose novel rule mining algorithms that can generate accurate and interpretable rules to improve the prediction performance. Considering the predictor’s real effects cannot be evaluated by traditional metrics like F1-score, we propose a new metric, NURR, to quantify the node unavailability reduction rate and tune model hyperparameters with NURR. The comparative study shows that our approach achieves over 40% better NURR than existing methods and runs stably in the production environment. Yaoguang Yong, Yunong Wang, Xuhua Ma, Bin Yao 0002, Linquan Jiang |
IEEE Trans. Computers | 5 |
| 2025 | Predicting DRAM-Caused Risky VMs in Large-Scale CloudsabstractDRAM failures are one of the leading failures in large-scale clouds. Previous studies focus on predicting DRAM uncorrectable errors (UEs) and mitigating the impact of DRAM failures through node-level workload migration and proactive dual in-line memory module (DIMM) replacement. For cloud systems, node migration means migrating all virtual machines (VMs) to other nodes. Such a coarse-grained migration consumes considerable resources. Inspired by the observation that DRAM errors tend to cluster in space, this paper proposes Pegasus, the first VM-level solution to mitigate DRAM failures. The key idea of Pegasus is to use the VM as the basic unit for predicting DRAM failures instead of the whole node. Specifically, we introduce a new concept of DRAM-caused risky VMs, which causes node unavailability when accessed. We design a novel Error-VM mapping framework deployed on a large-scale cloud. Statistical results confirm that DRAM errors are concentrated in the address space managed by a small number of VMs. By combining ECC and spatio-temporal features, our predictor achieves decent performance. Pegasus has been deployed online on over $\mathbf{3 0 0, 0 0 0}$ nodes in the production cloud. The comparative study shows that the prediction performance of Pegasus is comparable to node-level solutions. Meanwhile, our approach offers over $\mathbf{7 0 \%}$ lower costs and avoids more than $\mathbf{1 0 \%}$ of VM crashes compared to node-level mitigation. Yaoguang Yong, Xuhua Ma, Bin Yao 0002, Huite Yi |
HPCA | 3 |
| 2024 | Time-Aware Attention-Based Transformer (TAAT) for Cloud Computing System Failure PredictionabstractLog-based failure prediction helps identify and mitigate system failures ahead of time, increasing the reliability of cloud elastic computing systems.However, most existing log-based failure prediction approaches only focus on semantic information, and do not make full use of the information contained in the timestamps of log messages.This paper proposes time-aware attention-based transformer (TAAT), a failure prediction approach that extracts semantic and temporal information simultaneously from log messages and their timestamps.TAAT first tokenizes raw log messages into specific exceptions, and then performs: 1) exception sequence embedding that reorganizes the exceptions of each node as an ordered sequence and converts them to vectors; 2) time relation estimation that computes time relation matrices from the timestamps; and, 3) time-aware attention that computes semantic correlation matrices from the exception sequences and then combines them with time relation matrices.Experiments on Alibaba Cloud demonstrated that TAAT achieves an approximately 10% performance improvement compared with the state-of-the-art approaches.TAAT is now used in the daily operation of Alibaba Cloud.Moreover, this paper also releases the real-world cloud computing failure prediction dataset used in our study, which consists of about 2.7 billion syslogs from about 300,000 node controllers during a 4-month period.To our * Both authors contributed equally to this research. Lingfei Deng, Yunong Wang, Xuhua Ma, Dongrui Wu |
KDD | 4 |
| 2024 | MISP: A Multimodal-based Intelligent Server Failure Prediction Model for Cloud Computing SystemsabstractTraditional server failure prediction methods predominantly rely on single-modality data such as system logs or system status curves. This reliance may lead to an incomplete understanding of system health and impending issues, proving inadequate for the complex and dynamic landscape of contemporary cloud computing environments. The potential of multimodal data to provide comprehensive insights is widely acknowledged, yet the lack of a holistic dataset and the challenges inherent in integrating features from both structured and unstructured data have impeded the exploration of multimodal-based server failure prediction. Addressing these challenges, this paper presents an industrial-scale, comprehensive dataset for server failure prediction, comprising nearly 80 types of structured and unstructured data sourced from real-world industrial cloud systems 1. Building on this resource, we introduce MISP, a model that leverages multimodal fusion techniques for server failure prediction. MISP transforms multimodal data into multi-dimensional sequences, extracts and encodes features both within and across the modalities, and ultimately computes the failure probability from the synthesized features. Experiments demonstrate that MISP significantly outperforms existing methods, enhancing prediction accuracy by approximately 25% over previous state-of-the-art approaches. Xianting Lu, Yunong Wang, Yu Fu 0008, Qi Sun 0002, Xuhua Ma, Cheng Zhuo |
KDD | 5 |
| 2023 | DARC: High-dimensional Diffusing Anomaly Detection and Root Cause Location in Cloud Computing SystemsabstractThe modern cloud computing system has evolved into a highly dynamic and complex ecosystem with thousands of modules. Modifications and updates to these modules occur every day to accommodate customer needs. Unfortunately, these frequent changes may introduce anomalies to the system, whose diffusion can undermine the system performance and even cause system outage. Though very important, it is challenging to detect anomalies at the early stages and locate their root causes, due to the complexity of the cloud ecosystem and the huge number of attribute combinations. This paper proposes DARC for high-dimensional diffusing anomaly detection and root cause location in cloud computing systems. DARC uses first two-stage percentile analysis and Mann-Kendall score thresholding to detect rare anomalies, and then a bottom-up search strategy with three computational complexity reduction techniques to efficiently locate the root causes. Extensive experiments showed that DARC is able to accurately and efficiently locate root causes of diffusing anomalies. It has been successfully used in the daily practice of Alibaba Cloud, one of the world’s largest cloud computing service providers. Wanning Sun, Xuhua Ma, Ruimin Peng, Yifan Xu 0015, Dongrui Wu |
IEEE Big Data | 2 |
| 2022 | Predicting DRAM-Caused Node Unavailability in Hyper-Scale CloudsabstractDRAM faults are major hardware sources of cloud node unavailability. To enable early preventive actions and mitigate DRAM fault impacts, prior studies focus on predicting DRAM uncorrectable errors (UEs) that typically cause immediate node unavailability. In our cloud with over half a million nodes, we firstly observe that the correctable error storm (numerous CEs occur in a short period) dominates 56% DRAM-caused node unavailability (DCNU). Therefore, we propose to predict DCNU that takes account into both UEs and CE storms. Observing that DCNUs have strong relevance to temporal statistics and spatial patterns of CEs, we design novel spatio-temporal features to train the prediction model. Considering the model’s real effects cannot be evaluated by traditional metrics like F1-score, we propose a new metric NURR to quantify the node unavailability reduction and tune model hyperparameters with NURR. Our approach achieves over 40% better NURR than existing methods on historical data and runs stably in the production environment. Yunong Wang, Xuhua Ma, Yaoheng Xu, Bin Yao 0002, Linquan Jiang |
DSN | 3 |