Yunong Wang

dblp:166/2625 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 1 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LMID: A Comprehensive Multimodal Dataset for Failure Prediction in Cloud Computing Systems
abstract
Failure prediction is crucial for ensuring the stability of cloud computing systems and has garnered extensive attention from both academia and industry. Generally, data used for prediction includes two modalities: 1) Text data, such as logs; and, 2) Numerical data, such as error counts and monitoring metrics. However, most existing failure prediction algorithms for cloud computing only focus on a single modality. The lack of high-quality multimodal datasets from real-world production environments constrains academic research on multimodal failure prediction. To fill this gap, this paper releases a large multimodal dataset of operation data from the Alibaba cloud computing platform, namely, Logs and Metrics Integration Dataset (LMID). It consists of 100 million pieces of logs (textual data) and 37 dimensions of monitoring metrics (numerical data) from 220,000 physical machines. To our knowledge, it is the first multimodal dataset for cloud computing system failure prediction, and is expected to greatly benefit the community. This paper provides a detailed introduction to the construction of LMID, its contents, and the performance of state-of-the-art algorithms on it. It also conducts extensive experiments to reveal a new insight that cross-modality connections are effective for failure prediction. LMID is now available at https://huggingface.co/datasets/AliyunECSAlgos/LMID.
Lingfei Deng, Ruqiao Xu, Yunong Wang, Xuhua Ma, Dongrui Wu
KDD (1)3
2026 TSUIE: Efficient two-stage underwater image enhancement framework
Lingfeng Chen, Tianheng Ma, Yuanxin Xu, Zhihan Xu, Shibo Lu, Yunong Wang
Neurocomputing6
2026 Collaborative Prediction of Cloud DRAM Failures With Rules and Machine Learning
abstract
DRAM faults are the main hardware cause of node unavailability in clouds. To enable early preventive actions and mitigate DRAM fault impacts, prior studies focus on predicting DRAM uncorrectable errors (UEs) that typically cause immediate node unavailability. However, in Alibaba Cloud, we observe that correctable error(CE) storm (numerous CEs occur in a short period) dominates 41% DRAM-caused node unavailability (DCNU). In this paper, we propose to predict DCNU by taking into account both UEs and CE storms. Specifically, observing that DCNUs have strong relevance to the temporal statistics and spatial patterns of CEs, we design novel spatio-temporal features and use soft labels to build a DCNU predictor. We propose novel rule mining algorithms that can generate accurate and interpretable rules to improve the prediction performance. Considering the predictor’s real effects cannot be evaluated by traditional metrics like F1-score, we propose a new metric, NURR, to quantify the node unavailability reduction rate and tune model hyperparameters with NURR. The comparative study shows that our approach achieves over 40% better NURR than existing methods and runs stably in the production environment.
Yaoguang Yong, Yunong Wang, Xuhua Ma, Bin Yao 0002, Linquan Jiang
IEEE Trans. Computers3
2025 Inter-graph and Intra-graph: Utilizing global financial markets and constituent stocks for stock index prediction
Yong Shi 0001, Yunong Wang
Eng. Appl. Artif. Intell.2
2024 Time-Aware Attention-Based Transformer (TAAT) for Cloud Computing System Failure Prediction
abstract
Log-based failure prediction helps identify and mitigate system failures ahead of time, increasing the reliability of cloud elastic computing systems.However, most existing log-based failure prediction approaches only focus on semantic information, and do not make full use of the information contained in the timestamps of log messages.This paper proposes time-aware attention-based transformer (TAAT), a failure prediction approach that extracts semantic and temporal information simultaneously from log messages and their timestamps.TAAT first tokenizes raw log messages into specific exceptions, and then performs: 1) exception sequence embedding that reorganizes the exceptions of each node as an ordered sequence and converts them to vectors; 2) time relation estimation that computes time relation matrices from the timestamps; and, 3) time-aware attention that computes semantic correlation matrices from the exception sequences and then combines them with time relation matrices.Experiments on Alibaba Cloud demonstrated that TAAT achieves an approximately 10% performance improvement compared with the state-of-the-art approaches.TAAT is now used in the daily operation of Alibaba Cloud.Moreover, this paper also releases the real-world cloud computing failure prediction dataset used in our study, which consists of about 2.7 billion syslogs from about 300,000 node controllers during a 4-month period.To our * Both authors contributed equally to this research.
Lingfei Deng, Yunong Wang, Xuhua Ma, Dongrui Wu
KDD2
2024 MISP: A Multimodal-based Intelligent Server Failure Prediction Model for Cloud Computing Systems
abstract
Traditional server failure prediction methods predominantly rely on single-modality data such as system logs or system status curves. This reliance may lead to an incomplete understanding of system health and impending issues, proving inadequate for the complex and dynamic landscape of contemporary cloud computing environments. The potential of multimodal data to provide comprehensive insights is widely acknowledged, yet the lack of a holistic dataset and the challenges inherent in integrating features from both structured and unstructured data have impeded the exploration of multimodal-based server failure prediction. Addressing these challenges, this paper presents an industrial-scale, comprehensive dataset for server failure prediction, comprising nearly 80 types of structured and unstructured data sourced from real-world industrial cloud systems 1. Building on this resource, we introduce MISP, a model that leverages multimodal fusion techniques for server failure prediction. MISP transforms multimodal data into multi-dimensional sequences, extracts and encodes features both within and across the modalities, and ultimately computes the failure probability from the synthesized features. Experiments demonstrate that MISP significantly outperforms existing methods, enhancing prediction accuracy by approximately 25% over previous state-of-the-art approaches.
Xianting Lu, Yunong Wang, Yu Fu 0008, Qi Sun 0002, Xuhua Ma, Cheng Zhuo
KDD2
2023 FFT-based efficient Poisson solver in nonrectangular domain
abstract
Abstract Poisson's equation is one of the most popular partial differential equation (PDE), which is widely used in image processing, computer graphics and other fields. However, solving a large‐scale Poisson's equation always costs huge computational resources. Fast Fourier transform (FFT) is an efficient Poisson solver but it only works in rectangular domain. In this paper, we propose a FFT‐based Poisson solver in nonrectangular domain on regular grids combined with algebraic multigrid (AMG). We extend the original Poisson's equation to a rectangular domain to construct an equivalent equation, so that it can use FFT algorithm to accelerate the solving to Poisson's equation. Experiments show that the FFT‐based Poisson solver can improve the solving speed of large‐scale Poisson's equations in nonrectangular domain. We demonstrate the solver in applications of image processing and fluid simulation.
Yunong Wang, Fei Hou 0001, Wencheng Wang 0001
Comput. Animat. Virtual Worlds1
2022 Predicting DRAM-Caused Node Unavailability in Hyper-Scale Clouds
abstract
DRAM faults are major hardware sources of cloud node unavailability. To enable early preventive actions and mitigate DRAM fault impacts, prior studies focus on predicting DRAM uncorrectable errors (UEs) that typically cause immediate node unavailability. In our cloud with over half a million nodes, we firstly observe that the correctable error storm (numerous CEs occur in a short period) dominates 56% DRAM-caused node unavailability (DCNU). Therefore, we propose to predict DCNU that takes account into both UEs and CE storms. Observing that DCNUs have strong relevance to temporal statistics and spatial patterns of CEs, we design novel spatio-temporal features to train the prediction model. Considering the model’s real effects cannot be evaluated by traditional metrics like F1-score, we propose a new metric NURR to quantify the node unavailability reduction and tune model hyperparameters with NURR. Our approach achieves over 40% better NURR than existing methods on historical data and runs stably in the production environment.
Yunong Wang, Xuhua Ma, Yaoheng Xu, Bin Yao 0002, Linquan Jiang
DSN2
2022 A resource-constrained distributed task allocation method based on a two-stage coalition formation methodology for multi-UAVs
Mi Yang 0002, An Zhang 0002, Wenhao Bi, Yunong Wang
J. Supercomput.4
2017 Neural network with saliency based feature selection ability
abstract
Convolutional neural network (CNN) is inspired by the biological structure of human visual system (HVS). And there are still mechanisms in HVS that are worthy learn from. We were inspired by the function of feature selection in HVS which is named as visual saliency and proposed Sal-Mask connection and Ada-Sal Network to implement similar function in neural networks. In this paper, we did further research and tried three different improvement schemes on Ada-Sal Network. By visualizing features we illustrated the feature selection ability of Sal-Mask connection helps the neural network to extract features from more subtle details of input images. The experiment results proved that neural network works better with this saliency based feature selection function. We also found that Sal-Mask connection works best on features from the first convolutional connection.
Yunong Wang, Huanyu Bian, Nenghai Yu
ICIP1
2016 Ada-Sal Network: emulate the Human Visual System
Yunong Wang, Nenghai Yu, Taifeng Wang
Signal Process. Image Commun.1
2015 Improve Neural Network Using Saliency
Yunong Wang, Nenghai Yu, Taifeng Wang
ICIG (2)1