Wenchi Zhang

dblp:223/0795 · DBLP profile ↗
← Back
16ranked-venue papers
0as first author
9since 2021 · last 2024
0000-0002-5599-030XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 5 since 2021Computer networks · 5 · 2 since 2021Systems, architecture and hardware · 2Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2024 A survey on intelligent management of alerts and incidents in IT services
Qingyang Yu, Nengwen Zhao, Mingjie Li 0005, Zeyan Li 0001, Honglin Wang, Wenchi Zhang, Kaixin Sui, Dan Pei
J. Netw. Comput. Appl.6
2023 Multi-stage Location for Root-Cause Metrics in Online Service Systems
abstract
The failure of the online service system will seriously affect the user experience and bring huge economic losses. Therefore, the operators usually monitor service-level metrics and machine-level metrics to help quickly find failures, locate root-cause metrics, and reduce MTTR(mean time to repair). Many methods have emerged in recent years to automatically locate root-cause metrics. However, the existing methods cannot meet the requirements of efficiency, accuracy, and ease of deployment at the same time, and are difficult to use in practice. To overcome their limitations, we propose MetricMiner- a multi-stage location method for root-cause metrics in online service systems. Our approach is based on a key observation from numerous real-world cases: root-cause metrics tend to be unique in both the time dimension and the machine dimension. Therefore, we divide the root-cause metrics localization into three stages: first, quickly filter out normal metrics with limited historical data; second, obtain sufficient historical data to eliminate abnormal metrics; finally, according to the clustering of abnormal metrics between machines to sort and locate root-cause metrics. Experimental results on two real-world datasets with 194 cases show that our method can significantly outperform the state-of-the-art methods. Moreover, MetricMiner has been deployed to multiple banking services for more than six months, and we also shared some lessons learned from real deployment.
Wenchi Zhang, Shize Zhang, Kaixin Sui, Enhuan Dong, Jiahai Yang 0001
NOMS2
2022 Identifying Erroneous Software Changes through Self-Supervised Contrastive Learning on Time Series Data
abstract
Software changes are frequent and inevitable. How-ever, erroneous software changes may cause failures and incidents, degrading user experience and system stability. Thus, it is critical to distinguish erroneous software changes from normal ones. Our empirical study from a global data center reveals that erroneous software changes have caused nearly one-third of the critical incidents in the last two years. Some quantitative results also imply that the number of software changes and that of the Key Performance Indicator (KPI) time series related to a software change are relatively large. Based on the observations, we propose Kontrast, a self-supervised, generic and adaptive approach using contrastive learning, aiming to identify erroneous software changes on time. Its key idea is to compare pre-change and post-change KPI time series related to the software change, assuring the time series is still in a normal state after the software change. Since contrastive learning approaches need a fully-labeled dataset, we propose a novel data augmentation technique inspired by self-supervised learning to generate data with pseudo labels. Our model significantly outperforms all the compared approaches on two datasets with a millisecond-level speed for each KPI and is proven to obtain cross-dataset adaptability. To better certify our contribution, we also exhibit some success cases of Kontrast from its deployment.
Xuanrun Wang, Kanglin Yin, Qianyu Ouyang, Xidao Wen, Shenglin Zhang, Wenchi Zhang, Jiuxue Han, Dan Pei
ISSRE6
2022 Causal Inference-Based Root Cause Analysis for Online Service Systems with Intervention Recognition
abstract
Fault diagnosis is critical in many domains, as faults may lead to safety threats or economic losses. In the field of online service systems, operators rely on enormous monitoring data to detect and mitigate failures. Quickly recognizing a small set of root cause indicators for the underlying fault can save much time for failure mitigation. In this paper, we formulate the root cause analysis problem as a new causal inference task namedintervention recognition. We proposed a novel unsupervised causal inference-based method namedCausal Inference-based Root Cause Analysis (CIRCA). The core idea is a sufficient condition for a monitoring variable to be a root cause indicator,i.e., the change of probability distribution conditioned on the parents in the Causal Bayesian Network (CBN). Towards the application in online service systems, CIRCA constructs a graph among monitoring metrics based on the knowledge of system architecture and a set of causal assumptions. The simulation study illustrates the theoretical reliability of CIRCA. The performance on a real-world dataset further shows that CIRCA can improve the recall of the top-1 recommendation by 25% over the best baseline method.
Mingjie Li 0005, Zeyan Li 0001, Kanglin Yin, Xiaohui Nie, Wenchi Zhang, Kaixin Sui, Dan Pei
KDD5
2022 Actionable and interpretable fault localization for recurring failures in online service systems
abstract
Fault localization is challenging in an online service system due to its monitoring data's large volume and variety and complex dependencies across/within its components (e.g., services or databases). Furthermore, engineers require fault localization solutions to be actionable and interpretable, which existing research approaches cannot satisfy. Therefore, the common industry practice is that, for a specific online service system, its experienced engineers focus on localization for recurring failures based on the knowledge accumulated about the system and historical failures. More specifically, 1) they can identify the underlying root causes and take mitigation actions when pinpointing a group of indicative metrics on the faulty component; 2) their diagnosis knowledge is roughly based on how one failure might affect the components in the whole system.
Zeyan Li 0001, Nengwen Zhao, Mingjie Li 0005, Xianglin Lu, Dongdong Chang, Xiaohui Nie, Wenchi Zhang, Kaixin Sui, Guoqiang Duan, Dan Pei
ESEC/SIGSOFT FSE9
2021 Identifying Root-Cause Metrics for Incident Diagnosis in Online Service Systems
abstract
Incidents in online service systems could incur poor user experience and tremendous economic loss. To reduce the influence of incidents and guarantee service reliability, it is critical to identify root-cause metrics for engineers with clues to assist incident diagnosis. However, it is a challenging task due to the complicated dependencies and huge volume of various metrics in large-scale systems. Existing approaches are based on either anomaly detection or correlation analysis, performing not well in terms of accuracy or efficiency. To better understand the problem of root-cause metric identification, we conduct a preliminary study based on real-world data analysis and interactions with engineers. The key observation is that root-cause metrics should satisfy two requirements. One is that the metric is expected to behave abnormally during the incident; the other is that the anomaly pattern should meet physical meaning and engineers' demand. Motivated by the findings obtained from the study, we propose an effective approach named PatternMatcher to identifying root-cause metrics accurately. Specifically, PatternMatcher contains three steps, where coarse-grained anomaly detection aiming to filter out normal metrics, anomaly pattern classification aiming to filter out unimportant anomaly patterns, and root-cause metric ranking. An extensive study on four real-world datasets including 113 incident cases from a large commercial bank demonstrates that PatternMatcher outperforms all baseline approaches, achieving top-3 average accuracy of 0.91. Moreover, we have deployed PatternMatcher in practice and shared some successful cases from real deployment.
Canhua Wu, Nengwen Zhao, Xiaoqin Yang, ShiNing Li, Xidao Wen, Xiaohui Nie, Wenchi Zhang, Kaixin Sui, Dan Pei
ISSRE10
2021 Practical Root Cause Localization for Microservice Systems via Trace Analysis
abstract
Microservice architecture is applied by an increasing number of systems because of its benefits on delivery, scalability, and autonomy. It is essential but challenging to localize root-cause microservices promptly when a fault occurs. Traces are helpful for root-cause microservice localization, and thus many recent approaches utilize them. However, these approaches are less practical due to relying on supervision or other unrealistic assumptions. To overcome their limitations, we propose a more practical root-cause microservice localization approach named TraceRCA. The key insight of TraceRCA is that a microservice with more abnormal and less normal traces passing through it is more likely to be the root cause. Based on it, TraceRCA is composed of trace anomaly detection, suspicious microservice set mining and microservice ranking. We conducted experiments on hundreds of injected faults in a widely-used open-source microservice benchmark and a production system. The results show that TraceRCA is effective in various situations. The top-1 accuracy of TraceRCA outperforms the state-of-the-art unsupervised approaches by 44.8%. Besides, TraceRCA is applied in a large commercial bank, and it helps operators localize root causes for real-world faults accurately and efficiently. We also share some lessons learned from our real-world deployment.
Zeyan Li 0001, Junjie Chen 0003, Nengwen Zhao, Shuwei Zhang, Long Jiang, Leiqin Yan, Zikai Wang 0010, Zhekang Chen, Wenchi Zhang, Xiaohui Nie, Kaixin Sui, Dan Pei
IWQoS12
2021 Identifying bad software changes via multimodal anomaly detection for online service systems
abstract
In large-scale online service systems, software changes are inevitable and frequent. Due to importing new code or configurations, changes are likely to incur incidents and destroy user experience. Thus it is essential for engineers to identify bad software changes, so as to reduce the influence of incidents and improve system re- liability. To better understand bad software changes, we perform the first empirical study based on large-scale real-world data from a large commercial bank. Our quantitative analyses indicate that about 50.4% of incidents are caused by bad changes, mainly be- cause of code defect, configuration error, resource contention, and software version. Besides, our qualitative analyses show that the current practice of detecting bad software changes performs not well to handle heterogeneous multi-source data involved in soft- ware changes. Based on the findings and motivation obtained from the empirical study, we propose a novel approach named SCWarn aiming to identify bad changes and produce interpretable alerts accurately and timely. The key idea of SCWarn is drawing support from multimodal learning to identify anomalies from heterogeneous multi-source data. An extensive study on two datasets with various bad software changes demonstrates our approach significantly outperforms all the compared approaches, achieving 0.95 F1-score on average and reducing MTTD (mean time to detect) by 20.4%∼60.7%. In particular, we shared some success stories and lessons learned from the practical usage.
Nengwen Zhao, Junjie Chen 0003, Zhaoyang Yu 0002, Honglin Wang, Jiesong Li, Bin Qiu, Hongyu Xu, Wenchi Zhang, Kaixin Sui, Dan Pei
ESEC/SIGSOFT FSE8
2021 An empirical investigation of practical log anomaly detection for online service systems
abstract
Log data is an essential and valuable resource of online service systems, which records detailed information of system running status and user behavior. Log anomaly detection is vital for service reliability engineering, which has been extensively studied. However, we find that existing approaches suffer from several limitations when deploying them into practice, including 1) inability to deal with various logs and complex log abnormal patterns; 2) poor interpretability; 3) lack of domain knowledge. To help understand these practical challenges and investigate the practical performance of existing work quantitatively, we conduct the first empirical study and an experimental study based on large-scale real-world data. We find that logs with rich information indeed exhibit diverse abnormal patterns (e.g., keywords, template count, template sequence, variable value, and variable distribution). However, existing approaches fail to tackle such complex abnormal patterns, producing unsatisfactory performance. Motivated by obtained findings, we propose a generic log anomaly detection system named LogAD based on ensemble learning, which integrates multiple anomaly detection approaches and domain knowledge, so as to handle complex situations in practice. About the effectiveness of LogAD, the average F1-score achieves 0.83, outperforming all baselines. Besides, we also share some success cases and lessons learned during our study. To our best knowledge, we are the first to investigate practical log anomaly detection in the real world deeply. Our work is helpful for practitioners and researchers to apply log anomaly detection to practice to enhance service reliability.
Nengwen Zhao, Honglin Wang, Zeyan Li 0001, Zhu Pan, Xidao Wen, Wenchi Zhang, Kaixin Sui, Dan Pei
ESEC/SIGSOFT FSE10
2020 Practical and White-Box Anomaly Detection through Unsupervised and Active Learning
abstract
To ensure quality of service and user experience, large Internet companies often monitor various Key Performance Indicators (KPIs) of their systems so that they can detect anomalies and identify failure in real time. However, due to a large number of various KPIs and the lack of high-quality labels, existing KPI anomaly detection approaches either perform well only on certain types of KPIs or consume excessive resources. Therefore, to realize generic and practical KPI anomaly detection in the real world, we propose a KPI anomaly detection framework named iRRCF-Active, which contains an unsupervised and white-box anomaly detector based on Robust Random Cut Forest (RRCF), and an active learning component. Specifically, we novelly propose an improved RRCF (iRRCF) algorithm to overcome the drawbacks of applying original RRCF in KPI anomaly detection. Besides, we also incorporate the idea of active learning to make our model benefit from high-quality labels given by experienced operators. We conduct extensive experiments on a large-scale public dataset and a private dataset collected from a large commercial bank. The experimental resulta demonstrate that iRRCF-Active performs better than existing traditional statistical methods, unsupervised learning methods and supervised learning methods. Besides, each component in iRRCF-Active has also been demonstrated to be effective and indispensable.
Zhaowei Wang 0003, Zejun Xie, Nengwen Zhao, Junjie Chen 0003, Wenchi Zhang, Kaixin Sui, Dan Pei
ICCCN6
2020 Root-Cause Metric Location for Microservice Systems via Log Anomaly Detection
abstract
Microservice systems are typically fragile and failures are inevitable in them due to their complexity and large scale. However, it is challenging to localize the root-cause metric due to its complicated dependencies and the huge number of various metrics. Existing methods are based on either correlation between metrics or correlation between metrics and failures. All of them ignore the key data source in microservice, i.e., logs. In this paper, we propose a novel root-cause metric localization approach by incorporating log anomaly detection. Our approach is based on a key observation, the value of root-cause metric should be changed along with the change of the log anomaly score of the system caused by the failure. Specifically, our approach includes two components, collecting anomaly scores by log anomaly detection algorithm and identifying root-cause metric by robust correlation analysis with data augmentation. Experiments on an open-source benchmark microservice system have demonstrated our approach can identify root-cause metrics more accurately than existing methods and only require a short localization time. Therefore, our approach can assist engineers to save much effort in diagnosing and mitigating failures as soon as possible.
Lingzhi Wang 0002, Nengwen Zhao, Junjie Chen 0003, Pinnong Li, Wenchi Zhang, Kaixin Sui
ICWS5
2020 Automatically and Adaptively Identifying Severe Alerts for Online Service Systems
abstract
In large-scale online service system, to enhance the quality of services, engineers need to collect various monitoring data and write many rules to trigger alerts. However, the number of alerts is way more than what on-call engineers can properly investigate. Thus, in practice, alerts are classified into several priority levels using manual rules, and on-call engineers primarily focus on handling the alerts with the highest priority level (i.e., severe alerts). Unfortunately, due to the complex and dynamic nature of the online services, this rule-based approach results in missed severe alerts or wasted troubleshooting time on non-severe alerts. In this paper, we propose AlertRank, an automatic and adaptive framework for identifying severe alerts. Specifically, AlertRank extracts a set of powerful and interpretable features (textual and temporal alert features, univariate and multivariate anomaly features for monitoring metrics), adopts XGBoost ranking algorithm to identify the severe alerts out of all incoming alerts, and uses novel methods to obtain labels for both training and testing. Experiments on the datasets from a top global commercial bank demonstrate that AlertRank is effective and achieves the F1-score of 0.89 on average, outperforming all baselines. The feedback from practice shows AlertRank can significantly save the manual efforts for on-call engineers.
Nengwen Zhao, Panshi Jin, Xiaoqin Yang, Wenchi Zhang, Kaixin Sui, Dan Pei
INFOCOM6
2020 Real-time incident prediction for online service systems
abstract
Incidents in online service systems could dramatically degrade system availability and destroy user experience. To guarantee service quality and reduce economic loss, it is essential to predict the occurrence of incidents in advance so that engineers can take some proactive actions to prevent them. In this work, we propose an effective and interpretable incident prediction approach, called eWarn, which utilizes historical data to forecast whether an incident will happen in the near future based on alert data in real time. More specifically, eWarn first extracts a set of effective features (including textual features and statistical features) to represent omen alert patterns via careful feature engineering. To reduce the influence of noisy alerts (that are not relevant to the occurrence of incidents), eWarn then incorporates the multi-instance learning formulation. Finally, eWarn builds a classification model via machine learning and generates an interpretable report about the prediction result via a state-of-the-art explanation technique (i.e., LIME). In this way, an early warning signal along with its interpretable report can be sent to engineers to facilitate their understanding and handling for the incoming incident. An extensive study on 11 real-world online service systems from a large commercial bank demonstrates the effectiveness of eWarn, outperforming state-of-the-art alert-based incident prediction approaches and the practice of incident prediction with alerts. In particular, we have applied eWarn to two large commercial banks in practice and shared some success stories and lessons learned from real deployment.
Nengwen Zhao, Junjie Chen 0003, Xiaohui Nie, Wenchi Zhang, Kaixin Sui, Dan Pei
ESEC/SIGSOFT FSE10
2019 Cross-dataset Time Series Anomaly Detection for Cloud Systems
Xu Zhang 0024, Qingwei Lin, Yong Xu 0010, Si Qin, Hongyu Zhang 0002, Bo Qiao 0001, Yingnong Dang, Xinsheng Yang, Murali Chintalapati, Youjiang Wu, Ken Hsieh, Kaixin Sui, Yaohai Xu, Wenchi Zhang, Furao Shen, Dongmei Zhang 0001
USENIX ATC16
2019 Automatic and Generic Periodicity Adaptation for KPI Anomaly Detection
abstract
Key performance indicator (KPI) anomaly detection (AD) is critical to ensure service quality and reliability. Due to the effects of work days, off days, festivals, and business activities on user behavior, KPIs may exhibit different patterns within different days, which we call periodicity profiles of KPIs. However, existing KPI AD approaches have difficulties in adapting to diverse periodicity profiles due to the lack of generality. In this paper, we propose an automatic and generic framework called Period, which can accurately detect the periodicity profiles through daily subsequences clustering, and improve the performance of AD methods by robustly and automatically adapting to different periodicity profiles. In our evaluation using several real-world KPIs with different periodicity profiles from large Internet-based services, the clustering algorithm used to detect periodicity can achieve about 0.95 accuracy on average. More importantly, further evaluation on 56 KPIs shows that Period can significantly improve the best F-score of several widely used AD approaches by up to 0.66.
Nengwen Zhao, Jing Zhu 0007, Minghua Ma, Wenchi Zhang, Dan Pei
IEEE Trans. Netw. Serv. Manag.5
2018 Improving Service Availability of Cloud Systems by Predicting Disk Error
Yong Xu 0010, Kaixin Sui, Randolph Yao, Hongyu Zhang 0002, Qingwei Lin, Yingnong Dang, Peng Li 0062, Keceng Jiang, Wenchi Zhang, Jian-Guang Lou, Murali Chintalapati, Dongmei Zhang 0001
USENIX ATC9