Kanglin Yin

dblp:212/7688 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
6since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 Guardian of the Resiliency: Detecting Erroneous Software Changes Before They Make Your Microservice System Less Fault-Resilient
abstract
The microservice system’s resilience is crucial for ensuring the quality of service. Nowadays, software changes are frequent and error-prone, and erroneous software changes could reduce microservice systems’ resilience to handle faults, leading to service failures and negatively impacting user experience. To better understand erroneous software changes, we conducted an empirical study on 256 real-world incidents from four famous microservice systems. Our quantitative results indicate that 37.87% of erroneous software changes make the microservice systems less fault-resilient; that is, when a fault (e.g., network fluctuation, high CPU usage, etc.) happens in the system after the software change, the services are more likely to experience failures. We refer to these software changes as Erroneous Software Changes that Reduce fault Resilience(ESCR). Traditional methods struggle to detect ESCRs effectively because the occurrence of faults is unpredictable and can hardly be in their post-change monitoring windows. In this paper, we propose a novel framework named ResilienceGuardian, aiming to detect ESCRs before they make microservice systems less fault-resilient. The key idea is utilizing fault injection techniques to evaluate systems’ fault resilience in the staging environment and then training lightweight classifiers of KPI segment pairs to detect ESCRs. The performance of ResilienceGuardian is systematically evaluated on three datasets with various faults and erroneous software changes. The results show that ResilienceGuardian significantly outperforms all the baselines with a 0.9 F1-score in identifying ESCRs and reduces the training time by 56.23% to 97.53%. Besides, ResilienceGuardian can achieve minute-level ESCR detection in large-scale microservice systems.
Guanglei He, Xiaohui Nie, Ruming Tang, Zhaoyang Yu 0002, Xidao Wen, Kanglin Yin, Dan Pei
IWQoS7
2023 LWS: A framework for log-based workload simulation in session-based SUT
Yongqi Han 0001, Qingfeng Du, Jincheng Xu, Shengjie Zhao 0001, Zhekang Chen, Kanglin Yin, Dan Pei
J. Syst. Softw.7
2022 Mining Fluctuation Propagation Graph Among Time Series with Active Learning
Mingjie Li 0005, Minghua Ma, Xiaohui Nie, Kanglin Yin, Xidao Wen, Zhiyun Yuan, Duogang Wu, Guoying Li, Dan Pei
DEXA (1)4
2022 Identifying Erroneous Software Changes through Self-Supervised Contrastive Learning on Time Series Data
abstract
Software changes are frequent and inevitable. How-ever, erroneous software changes may cause failures and incidents, degrading user experience and system stability. Thus, it is critical to distinguish erroneous software changes from normal ones. Our empirical study from a global data center reveals that erroneous software changes have caused nearly one-third of the critical incidents in the last two years. Some quantitative results also imply that the number of software changes and that of the Key Performance Indicator (KPI) time series related to a software change are relatively large. Based on the observations, we propose Kontrast, a self-supervised, generic and adaptive approach using contrastive learning, aiming to identify erroneous software changes on time. Its key idea is to compare pre-change and post-change KPI time series related to the software change, assuring the time series is still in a normal state after the software change. Since contrastive learning approaches need a fully-labeled dataset, we propose a novel data augmentation technique inspired by self-supervised learning to generate data with pseudo labels. Our model significantly outperforms all the compared approaches on two datasets with a millisecond-level speed for each KPI and is proven to obtain cross-dataset adaptability. To better certify our contribution, we also exhibit some success cases of Kontrast from its deployment.
Xuanrun Wang, Kanglin Yin, Qianyu Ouyang, Xidao Wen, Shenglin Zhang, Wenchi Zhang, Jiuxue Han, Dan Pei
ISSRE2
2022 Causal Inference-Based Root Cause Analysis for Online Service Systems with Intervention Recognition
abstract
Fault diagnosis is critical in many domains, as faults may lead to safety threats or economic losses. In the field of online service systems, operators rely on enormous monitoring data to detect and mitigate failures. Quickly recognizing a small set of root cause indicators for the underlying fault can save much time for failure mitigation. In this paper, we formulate the root cause analysis problem as a new causal inference task namedintervention recognition. We proposed a novel unsupervised causal inference-based method namedCausal Inference-based Root Cause Analysis (CIRCA). The core idea is a sufficient condition for a monitoring variable to be a root cause indicator,i.e., the change of probability distribution conditioned on the parents in the Causal Bayesian Network (CBN). Towards the application in online service systems, CIRCA constructs a graph among monitoring metrics based on the knowledge of system architecture and a set of causal assumptions. The simulation study illustrates the theoretical reliability of CIRCA. The performance on a real-world dataset further shows that CIRCA can improve the recall of the top-1 recommendation by 25% over the best baseline method.
Mingjie Li 0005, Zeyan Li 0001, Kanglin Yin, Xiaohui Nie, Wenchi Zhang, Kaixin Sui, Dan Pei
KDD3
2021 On Representing Resilience Requirements of Microservice Architecture Systems
abstract
Together with the spread of DevOps practices and container technologies, Microservice Architecture has become a mainstream architecture style in recent years. Resilience is a key characteristic in Microservice Architecture (MSA) Systems, and it shows the ability to cope with various kinds of system disturbances which cause degradations of services. However, due to lack of consensus definition of resilience in the software field, although a lot of work has been done on resilience for MSA Systems, developers still do not have a clear idea on how resilient an MSA System should be, and what resilience mechanisms are needed. In this paper, by referring to existing systematic studies on resilience in other scientific areas, the definition of microservice resilience is provided and a Microservice Resilience Measurement Model is proposed to measure service resilience. And a requirement model to represent resilience requirements of MSA Systems is given. The requirement model uses elements in KAOS to represent notions in the measurement model, and decompose service resilience goals into system behaviors that can be executed by system components. As a proof of concept, a case study is conducted on an MSA System to illustrate how the proposed models are applied.
Kanglin Yin, Qingfeng Du
Int. J. Softw. Eng. Knowl. Eng.1
2019 Short-Term Performance Metrics Forecasting for Virtual Machine to Support Anomaly Detection Using Hybrid ARIMA-WNN Model
abstract
Anomaly detection is a significant functionality in most cloud monitoring applications. Time-series forecasting model could be easily used for predicting the values of the performance metrics which could be used for representing the performance status of the cloud environment. The proposed hybrid model combines both Autoregressive Integrated Moving Average (ARIMA) and Wavelet Neural Network (WNN) models. Firstly, ARIMA model is employed to firstly predict the linear component and then WNN model is used for the nonlinear residual component prediction. Finally, the results of the two parts are combined into the final prediction value of the performance metric. Finally the experimental results show that the hybrid model could produce more accurate short-term prediction than other models.
Qingfeng Du, Kanglin Yin
COMPSAC (2)4
2018 An Approach of Collecting Performance Anomaly Dataset for NFV Infrastructure
Qingfeng Du, Tiandi Xie, Kanglin Yin
ICA3PP (3)4
2018 Performance Anomaly Detection Models of Virtual Machines for Network Function Virtualization Infrastructure with Machine Learning
Qingfeng Du, YiQun Lin, Jiaye Zhu, Kanglin Yin
ICANN (2)6