Zhaoyang Yu 0002

dblp:07/6449-2 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0003-3179-3894ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Triangle: Empowering Incident Triage with Multi-Agent
abstract
As cloud service systems grow in scale and complexity, incidents that indicate unplanned interruptions and outages become unavoidable. Rapid and accurate triage of these incidents to the appropriate responsible teams is crucial to maintain service reliability and prevent significant financial losses. However, existing incident triage methods relying on manual operations and predefined rules often struggle with efficiency and accuracy due to the heterogeneity of incident data and the dynamic nature of domain knowledge across multiple teams.To solve these issues, we propose Triangle, an end-to-end incident triage system based on a Multi-Agent framework. Triangle leverages a semantic distillation mechanism to tackle the issue of semantic heterogeneity in incident data, enhancing the accuracy of incident triage. Additionally, we introduce multi-role agents and a negotiation mechanism to emulate human engineers’ workflows, effectively handling decentralized and dynamic domain knowledge from multiple teams. Furthermore, our system incorporates an automated troubleshooting information collection and mitigation mechanism, reducing the reliance on human labor and enabling fully automated end-to-end incident triage. Extensive experiments conducted on a real-world cloud production environment demonstrate that Triangle significantly improved incident triage accuracy (up to 97%) and reduced Time to Engage (TTE) by as much as 91%, demonstrating substantial operational impact across diverse cloud services.
Zhaoyang Yu 0002, Aoyang Fang, Minghua Ma, Jaskaran Singh Walia, Chaoyun Zhang, Shu Chi, Ze Li 0005, Murali Chintalapati, Xuchao Zhang, Rujia Wang, Chetan Bansal, Saravan Rajmohan, Qingwei Lin, Shenglin Zhang, Dan Pei, Pinjia He
ASE1
2024 Causality Enhanced Graph Representation Learning for Alert-Based Root Cause Analysis
abstract
Accurate and efficient root cause identification in online service systems is critical for service stability and user experience. When a system failure occurs, numerous alerts are generated, but existing methods fail to effectively integrate all these multi-modal data to pinpoint the root causes. Moreover, most existing approaches are inefficient for large-scale online services due to their high reliance on handcrafted rules and domain expertise. This paper introduces AlertRCA, an algorithm for Root Cause Analysis (RCA) based on Alert events. It utilizes a pre-trained Alert2Vec module to encode multi-modal alert information into vectors, and implements an RCA-oriented causality prediction graph attention network (CPGAT) to automatically gauge causal relationships between alerts. Further, we devise a novel dispersing and aggregating graph neural network (DAGNN) to identify root causes. Experiments on a real-world dataset collected from a top-tier e-commerce company reveal AlertRCA’s superior performance, achieving 83.9% top-1 and 96.8% top-3 accuracy on average. Our codes are available at https://github.com/NetManAIOps/AlertRCA.
Zhaoyang Yu 0002, Qianyu Ouyang, Changhua Pei, Xin Wang 0001, Wenxiao Chen, Liangfei Su, Huai Jiang, Xuanrun Wang, Dan Pei
CCGrid1
2024 Guardian of the Resiliency: Detecting Erroneous Software Changes Before They Make Your Microservice System Less Fault-Resilient
abstract
The microservice system’s resilience is crucial for ensuring the quality of service. Nowadays, software changes are frequent and error-prone, and erroneous software changes could reduce microservice systems’ resilience to handle faults, leading to service failures and negatively impacting user experience. To better understand erroneous software changes, we conducted an empirical study on 256 real-world incidents from four famous microservice systems. Our quantitative results indicate that 37.87% of erroneous software changes make the microservice systems less fault-resilient; that is, when a fault (e.g., network fluctuation, high CPU usage, etc.) happens in the system after the software change, the services are more likely to experience failures. We refer to these software changes as Erroneous Software Changes that Reduce fault Resilience(ESCR). Traditional methods struggle to detect ESCRs effectively because the occurrence of faults is unpredictable and can hardly be in their post-change monitoring windows. In this paper, we propose a novel framework named ResilienceGuardian, aiming to detect ESCRs before they make microservice systems less fault-resilient. The key idea is utilizing fault injection techniques to evaluate systems’ fault resilience in the staging environment and then training lightweight classifiers of KPI segment pairs to detect ESCRs. The performance of ResilienceGuardian is systematically evaluated on three datasets with various faults and erroneous software changes. The results show that ResilienceGuardian significantly outperforms all the baselines with a 0.9 F1-score in identifying ESCRs and reduces the training time by 56.23% to 97.53%. Besides, ResilienceGuardian can achieve minute-level ESCR detection in large-scale microservice systems.
Guanglei He, Xiaohui Nie, Ruming Tang, Zhaoyang Yu 0002, Xidao Wen, Kanglin Yin, Dan Pei
IWQoS5
2024 Pre-trained KPI Anomaly Detection Model Through Disentangled Transformer
abstract
In large-scale online service systems, numerous Key Performance Indicators (KPIs), such as service response time and error rate, are gathered in a time-series format. KPI Anomaly Detection (KAD) is a critical data mining problem due to its widespread applications in real-world scenarios. However, KAD faces the challenges of dealing with KPI heterogeneity and noisy data. We propose KAD-Disformer, a KPI Anomaly Detection approach through Disentangled Transformer. KAD-Disformer pre-trains a model on existing accessible KPIs, and the pre-trained model can be effectively "fine-tuned" to unseen KPI using only a handful of samples from the unseen KPI. We propose a series of innovative designs, including disentangled projection for transformer, unsupervised few-shot fine-tuning (uTune), and denoising modules, each of which significantly contributes to the overall performance. Our extensive experiments demonstrate that KAD-Disformer surpasses the state-of-the-art universal anomaly detection model by 13% in F1-score and achieves comparable performance using only 1/8 of the finetuning samples saving about 25 hours. KAD-Disformer has been successfully deployed in the real-world cloud system serving millions of users, attesting to its feasibility and robustness. Our code is available at https://github.com/NetManAIOps/KAD-Disformer.
Zhaoyang Yu 0002, Changhua Pei, Xin Wang 0001, Minghua Ma, Chetan Bansal, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001, Xidao Wen, Gaogang Xie, Dan Pei
KDD1
2024 Supervised Fine-Tuning for Unsupervised KPI Anomaly Detection for Mobile Web Systems
abstract
With the rapid development of cellular networks, wireless base stations (WBSes) have become crucial infrastructure for mobile web systems. To ensure service quality, operators constantly monitor the operation status of WBSes and deploy anomaly detection methods to identify anomalies promptly. After the deployment of anomaly detection methods, operators periodically collect feedback, which holds significant value in improving anomaly detection performance. In real-world industrial environments, the frequency of false negative feedback is usually very low, and the newly generated data's distribution can differ significantly from that of the original training data. Therefore, the feedback-based performance improvement of the previously proposed methods is limited. In this paper, we propose AnoTuner, which incorporates a false negative augmentation mechanism to generate similar false negative feedback cases, effectively compensating for the low feedback frequency. Additionally, we introduce a Two-Stage Active Learning (TSAL) mechanism that minimizes data contamination issues caused by the difference between the distribution of feedback data and that of the training data. Experiments conducted on the real-world data collected from a top-tier global Internet Service Provider (ISP) demonstrate that the performance improvement of AnoTuner after feedback-based fine-tuning is significantly higher than that of the best baseline method.
Zhaoyang Yu 0002, Shenglin Zhang, Yingke Li, Yankai Zhao, Xiaolei Hua, Xidao Wen, Dan Pei
WWW1
2023 AutoKAD: Empowering KPI Anomaly Detection with Label-Free Deployment
abstract
Monitoring Key Performance Indicators (KPIs) and detecting anomalies in online service systems is critical. However, choosing the right KPI anomaly detection algorithm and appropriate hyperparameters presents a challenge. Conventional Automated Machine Learning (AutoML) struggles to address this because the hold-out dataset lacks labels and its loss doesn’t reliably reflect anomaly detection accuracy. To address the above challenges, this paper introduces AutoKAD, an AutoML framework designed to solve the combined algorithm selection and hyperparameter optimization problem for unsupervised KPI Anomaly Detection. We propose a label-free universal objective function, inspired by the Local Outlier Factor (LOF), for evaluating AutoML trials. Additionally, we improve the acquisition function and designs a cluster-based warm start strategy to enhance exploration effectiveness and efficiency. The experimental results on three real-world datasets show that our approach outperforms the SOTA model selection algorithm by 11% in F1-score and achieves comparable performance (99%) with theoretically optimal results. We believe that AutoKAD can greatly improve the deployment feasibility of existing anomaly detection algorithms in real-world systems. Our code is anonymously released at https://github.com/NetManAIOps/AutoKAD.
Zhaoyang Yu 0002, Changhua Pei, Shenglin Zhang, Xidao Wen, Gaogang Xie, Dan Pei
ISSRE1
2021 Identifying bad software changes via multimodal anomaly detection for online service systems
abstract
In large-scale online service systems, software changes are inevitable and frequent. Due to importing new code or configurations, changes are likely to incur incidents and destroy user experience. Thus it is essential for engineers to identify bad software changes, so as to reduce the influence of incidents and improve system re- liability. To better understand bad software changes, we perform the first empirical study based on large-scale real-world data from a large commercial bank. Our quantitative analyses indicate that about 50.4% of incidents are caused by bad changes, mainly be- cause of code defect, configuration error, resource contention, and software version. Besides, our qualitative analyses show that the current practice of detecting bad software changes performs not well to handle heterogeneous multi-source data involved in soft- ware changes. Based on the findings and motivation obtained from the empirical study, we propose a novel approach named SCWarn aiming to identify bad changes and produce interpretable alerts accurately and timely. The key idea of SCWarn is drawing support from multimodal learning to identify anomalies from heterogeneous multi-source data. An extensive study on two datasets with various bad software changes demonstrates our approach significantly outperforms all the compared approaches, achieving 0.95 F1-score on average and reducing MTTD (mean time to detect) by 20.4%∼60.7%. In particular, we shared some success stories and lessons learned from the practical usage.
Nengwen Zhao, Junjie Chen 0003, Zhaoyang Yu 0002, Honglin Wang, Jiesong Li, Bin Qiu, Hongyu Xu, Wenchi Zhang, Kaixin Sui, Dan Pei
ESEC/SIGSOFT FSE3
2021 LogClass: Anomalous Log Identification and Classification With Partial Labels
abstract
Logs are imperative in the management process of networks and services. However, manually identifying and classifying anomalous logs is time-consuming, error-prone, and labor-intensive. Additionally, rule-based approaches cannot tackle the challenges underlying anomalous log identification and classification resulting from new types of logs and partial labels. We propose LogClass, a framework to automatically and robustly identify and classify anomalous logs for network and service based onpartial labels. LogClass combines a word representation method, a positive and unlabeled learning (PU learning) model, and a machine learning classifier. Besides, we propose a novel Inverse Location Frequency (ILF) method to weight the words of logs in feature construction properly. We evaluate the performance of LogClass based on 18 million+ real-world switch logs and six public log datasets. It achieves 99.56% and 98% F1 scores in anomalous log identification on switch logs and publicly available supercomputer logs, respectively, and very-close-to-one F1 score in anomalous log classification. Moreover, we have conducted extensive experiments to demonstrate LogClass’ superior performance in addressing partial labels and new types of logs.
Weibin Meng, Ying Liu 0024, Shenglin Zhang, Federico Zaiter, Zhaoyang Yu 0002, Dan Pei
IEEE Trans. Netw. Serv. Manag.7