Zhangwei Xu

dblp:214/5675 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
4since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 9 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2023 Incident-aware Duplicate Ticket Aggregation for Cloud Systems
abstract
In cloud systems, incidents are potential threats to customer satisfaction and business revenue. When customers are affected by incidents, they often request customer support service (CSS) from the cloud provider by submitting a support ticket. Many tickets could be duplicate as they are reported in a distributed and uncoordinated manner. Thus, aggregating such duplicate tickets is essential for efficient ticket management. Previous studies mainly rely on tickets' textual similarity to detect duplication; however, duplicate tickets in a cloud system could carry semantically different descriptions due to the complex service dependency of the cloud system. To tackle this problem, we propose iPACK, an incident-aware method for aggregating duplicate tickets by fusing the failure information between the customer side (i.e., tickets) and the cloud side (i.e., incidents). We extensively evaluate iPACK on three datasets collected from the production environment of a large-scale cloud platform, Azure. The experimental results show that iPACK can precisely and comprehensively aggregate duplicate tickets, achieving an F1 score of 0.871~0.935 and outperforming state-of-the-art methods by 12.4%~31.2%.
Jinyang Liu 0002, Shilin He, Zhuangbin Chen, Liqun Li, Yu Kang 0006, Xu Zhang 0024, Pinjia He, Hongyu Zhang 0002, Qingwei Lin, Zhangwei Xu, Saravan Rajmohan, Dongmei Zhang 0001, Michael R. Lyu
ICSE10
2021 Fast Outage Analysis of Large-scale Production Clouds with Service Correlation Mining
abstract
Cloud-based services are surging into popularity in recent years. However, outages, i.e., severe incidents that always impact multiple services, can dramatically affect user experience and incur severe economic losses. Locating the root-cause service, i.e., the service that contains the root cause of the outage, is a crucial step to mitigate the impact of the outage. In current industrial practice, this is generally performed in a bootstrap manner and largely depends on human efforts: the service that directly causes the outage is identified first, and the suspected root cause is traced back manually from service to service during diagnosis until the actual root cause is found. Unfortunately, production cloud systems typically contain a large number of interdependent services. Such a manual root cause analysis is often time-consuming and labor-intensive. In this work, we propose COT, the first outage triage approach that considers the global view of service correlations. COT mines the correlations among services from outage diagnosis data. After learning from historical outages, COT can infer the root cause of emerging ones accurately. We implement COT and evaluate it on a real-world dataset containing one year of data collected from Microsoft Azure, one of the representative cloud computing platforms in the world. Our experimental results show that COT can reach a triage accuracy of 82.1%-83.5%, which outperforms the state-of-the-art triage approach by 28.0%-29.7%.
Yaohui Wang 0003, Guo-Zheng Li 0001, Yu Kang 0006, Yangfan Zhou 0002, Hongyu Zhang 0002, Feng Gao 0022, Jeffrey Sun, Pochian Lee, Zhangwei Xu, Pu Zhao 0004, Bo Qiao 0001, Liqun Li, Xu Zhang 0024, Qingwei Lin
ICSE11
2021 How Long Will it Take to Mitigate this Incident for Online Service Systems?
abstract
Online service systems may encounter a large number of incidents, which should be mitigated as soon as possible to minimize the service disruption time and ensure high service availability. The ability to predict TTM (Time To Mitigation) of incidents can help service teams better organize the mainte-nance efforts. Although there are many traditional bug-fixing time prediction methods, we find that there are not readily available for incident- TTM prediction due to the characteristics of incidents. To better understand how incidents are mitigated, we conduct the first empirical study of incident TTM on 20 large-scale online service systems in Microsoft. We investigate the time distribution in the main stages of the incident life cycle and explore factors affecting TTM. Based on our empirical findings, we propose TTMPred, a deep-learning-based approach for incident- TTM prediction in a continuous triage scenario. Our model designs a two-level attention-based bidirectional GRU model to capture both the semantic information in text data and the temporal information in incremental discussions. And based on a novel continuous loss function, it builds a regression model to achieve accurate TTM prediction as much as possible at each time point of prediction. Our experiments on four large-scale online service systems in Microsoft show that TTMPred is effective and significantly outperforms the compared approaches. For example, TTMPred improves the state-of-the-art regression-based approach by 25.66% on average in terms of MAE (Mean Absolute Error).
Weijing Wang, Junjie Chen 0003, Lin Yang 0030, Hongyu Zhang 0002, Pu Zhao 0004, Bo Qiao 0001, Yu Kang 0006, Qingwei Lin, Saravanakumar Rajmohan, Feng Gao 0022, Zhangwei Xu, Yingnong Dang, Dongmei Zhang 0001
ISSRE11
2021 Fighting the Fog of War: Automated Incident Detection for Cloud Systems
Liqun Li, Xu Zhang 0024, Hongyu Zhang 0002, Yu Kang 0006, Pu Zhao 0004, Bo Qiao 0001, Shilin He, Pochian Lee, Jeffrey Sun, Feng Gao 0022, Qingwei Lin, Saravanakumar Rajmohan, Zhangwei Xu, Dongmei Zhang 0001
USENIX ATC15
2020 How Incidental are the Incidents? Characterizing and Prioritizing Incidents for Large-Scale Online Service Systems
abstract
Although tremendous efforts have been devoted to the quality assurance of online service systems, in reality, these systems still come across many incidents (i.e., unplanned interruptions and outages), which can decrease user satisfaction or cause economic loss. To better understand the characteristics of incidents and improve the incident management process, we perform the first large-scale empirical analysis of incidents collected from 18 real-world online service systems in Microsoft. Surprisingly, we find that although a large number of incidents could occur over a short period of time, many of them actually do not matter, i.e., engineers will not fix them with a high priority after manually identifying their root cause. We call these incidents incidental incidents. Our qualitative and quantitative analyses show that incidental incidents are significant in terms of both number and cost. Therefore, it is important to prioritize incidents by identifying incidental incidents in advance to optimize incident management efforts. In particular, we propose an approach, called DeepIP (Deep learning based Incident Prioritization), to prioritizing incidents based on a large amount of historical incident data. More specifically, we design an attention-based Convolutional Neural Network (CNN) to learn a prediction model to identify incidental incidents. We then prioritize all incidents by ranking the predicted probabilities of incidents being incidental. We evaluate the performance of DeepIP using real-world incident data. The experimental results show that DeepIP effectively prioritizes incidents by identifying incidental incidents and significantly outperforms all the compared approaches. For example, the AUC of DeepIP achieves 0.808, while that of the best compared approach is only 0.624 on average.
Junjie Chen 0003, Xiaoting He 0003, Qingwei Lin, Hongyu Zhang 0002, Dan Hao 0001, Yu Kang 0006, Feng Gao 0022, Zhangwei Xu, Yingnong Dang, Dongmei Zhang 0001
ASE9
2020 Towards intelligent incident management: why we need it and how we make it
abstract
The management of cloud service incidents (unplanned interruptions or outages of a service/product) greatly affects customer satisfaction and business revenue. After years of efforts, cloud enterprises are able to solve most incidents automatically and timely. However, in practice, we still observe critical service incidents that occurred in an unexpected manner and orchestrated diagnosis workflow failed to mitigate them. In order to accelerate the understanding of unprecedented incidents and provide actionable recommendations, modern incident management system employs the strategy of AIOps (Artificial Intelligence for IT Operations). In this paper, to provide a broad view of industrial incident management and understand the modern incident management system, we conduct a comprehensive empirical study spanning over two years of incident management practices at Microsoft. Particularly, we identify two critical challenges (namely, incomplete service/resource dependencies and imprecise resource health assessment) and investigate the underlying reasons from the perspective of cloud system design and operations. We also present IcM BRAIN, our AIOps framework towards intelligent incident management, and show its practical benefits conveyed to the cloud services of Microsoft.
Zhuangbin Chen, Yu Kang 0006, Liqun Li, Xu Zhang 0024, Hongyu Zhang 0002, Hui Xu 0009, Yangfan Zhou 0002, Jeffrey Sun, Zhangwei Xu, Yingnong Dang, Feng Gao 0022, Pu Zhao 0004, Bo Qiao 0001, Qingwei Lin, Dongmei Zhang 0001, Michael R. Lyu
ESEC/SIGSOFT FSE10
2020 Identifying linked incidents in large-scale online service systems
abstract
In large-scale online service systems, incidents occur frequently due to a variety of causes, from updates of software and hardware to changes in operation environment. These incidents could significantly degrade system’s availability and customers’ satisfaction. Some incidents are linked because they are duplicate or inter-related. The linked incidents can greatly help on-call engineers find mitigation solutions and identify the root causes. In this work, we investigate the incidents and their links in a representative real-world incident management (IcM) system. Based on the identified indicators of linked incidents, we further propose LiDAR (Linked Incident identification with DAta-driven Representation), a deep learning based approach to incident linking. More specifically, we incorporate the textual description of incidents and structural information extracted from historical linked incidents to identify possible links among a large number of incidents. To show the effectiveness of our method, we apply our method to a real-world IcM system and find that our method outperforms other state-of-the-art methods.
Yujun Chen, Xian Yang 0001, Hang Dong 0004, Xiaoting He 0003, Hongyu Zhang 0002, Qingwei Lin, Junjie Chen 0003, Pu Zhao 0004, Yu Kang 0006, Feng Gao 0022, Zhangwei Xu, Dongmei Zhang 0001
ESEC/SIGSOFT FSE11
2020 Efficient customer incident triage via linking with system incidents
abstract
In cloud service systems, customers will report the service issues they have encountered to cloud service providers. Despite many issues can be handled by the support team, sometimes the customer issues can not be easily solved, thus raising customer incidents. Quick troubleshooting of a customer incident is critical. To this end, a customer incident should be assigned to its responsible team accurately in a timely manner.
Jiazhen Gu, Jiaqi Wen, Pu Zhao 0004, Chuan Luo 0002, Yu Kang 0006, Yangfan Zhou 0002, Jeffrey Sun, Zhangwei Xu, Bo Qiao 0001, Liqun Li, Qingwei Lin, Dongmei Zhang 0001
ESEC/SIGSOFT FSE10
2020 How to mitigate the incident? an effective troubleshooting guide recommendation technique for online service systems
abstract
In recent years, more and more traditional shrink-wrapped software is provided as 7x24 online services. Incidents (events that lead to service disruptions or outages) could affect service availability and cause great financial loss. Therefore, mitigating the incidents is important and time critical. In practice, a document describing a mitigation process, called a troubleshooting guide (TSG), is usually used to reduce the Time To Mitigate (TTM). To investigate the usage of TSGs in real-world online services, we conduct the first empirical study on 18 real-world, large-scale online service systems in Microsoft. We analyze the distribution and characteristics of TSGs among all incident records in the past two years. According to our study, 27.2% incidents have TSG records and 36.2% of them occurred at least twice. Besides, on average developers spend around 36.3% of the entire mitigation time on locating the desired TSGs.
Jiajun Jiang, Weihai Lu, Junjie Chen 0003, Qingwei Lin, Pu Zhao 0004, Yu Kang 0006, Hongyu Zhang 0002, Yingfei Xiong 0001, Feng Gao 0022, Zhangwei Xu, Yingnong Dang, Dongmei Zhang 0001
ESEC/SIGSOFT FSE10
2019 Continuous Incident Triage for Large-Scale Online Service Systems
abstract
In recent years, online service systems have become increasingly popular. Incidents of these systems could cause significant economic loss and customer dissatisfaction. Incident triage, which is the process of assigning a new incident to the responsible team, is vitally important for quick recovery of the affected service. Our industry experience shows that in practice, incident triage is not conducted only once in the beginning, but is a continuous process, in which engineers from different teams have to discuss intensively among themselves about an incident, and continuously refine the incident-triage result until the correct assignment is reached. In particular, our empirical study on 8 real online service systems shows that the percentage of incidents that were reassigned ranges from 5.43% to 68.26% and the number of discussion items before achieving the correct assignment is up to 11.32 on average. To improve the existing incident triage process, in this paper, we propose DeepCT, a Deep learning based approach to automated Continuous incident Triage. DeepCT incorporates a novel GRU-based (Gated Recurrent Unit) model with an attention-based mask strategy and a revised loss function, which can incrementally learn knowledge from discussions and update incident-triage results. Using DeepCT, the correct incident assignment can be achieved with fewer discussions. We conducted an extensive evaluation of DeepCT on 14 large-scale online service systems in Microsoft. The results show that DeepCT is able to achieve more accurate and efficient incident triage, e.g., the average accuracy identifying the responsible team precisely is 0.641~0.729 with the number of discussion items increasing from 1 to 5. Also, DeepCT statistically significantly outperforms the state-of-the-art bug triage approach.
Junjie Chen 0003, Xiaoting He 0003, Qingwei Lin, Hongyu Zhang 0002, Dan Hao 0001, Feng Gao 0022, Zhangwei Xu, Yingnong Dang, Dongmei Zhang 0001
ASE7
2019 Outage Prediction and Diagnosis for Cloud Service Systems
abstract
With the rapid growth of cloud service systems and their increasing complexity, service failures become unavoidable. Outages, which are critical service failures, could dramatically degrade system availability and impact user experience. To minimize service downtime and ensure high system availability, we develop an intelligent outage management approach, called AirAlert, which can forecast the occurrence of outages before they actually happen and diagnose the root cause after they indeed occur. AirAlert works as a global watcher for the entire cloud system, which collects all alerting signals, detects dependency among signals and proactively predicts outages that may happen anywhere in the whole cloud system. We analyze the relationships between outages and alerting signals by leveraging Bayesian network and predict outages using a robust gradient boosting tree based classification method. The proposed outage management approach is evaluated using the outage dataset collected from a Microsoft cloud system and the results confirm the effectiveness of the proposed approach.
Yujun Chen, Xian Yang 0001, Qingwei Lin, Hongyu Zhang 0002, Feng Gao 0022, Zhangwei Xu, Yingnong Dang, Dongmei Zhang 0001, Hang Dong 0004, Yong Xu 0010, Yu Kang 0006
WWW6