Minghua Ma

dblp:164/6802 · DBLP profile ↗
← Back
49ranked-venue papers
5as first author
41since 2021 · last 2026
0000-0002-6303-1731ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 22 · 2 first-author · 21 since 2021Artificial intelligence and machine learning · 11 · 11 since 2021Databases, data management, data science and information retrieval · 11 · 1 first-author · 10 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Computer networks · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 PARIF: Pushing the Pareto Frontier of Instruction Following and Reasoning with Curriculum Reinforcement Learning
abstract
Rongchuan Mu, Zexin Wang, Qianyu Wang, MingHua Ma, Zekun Wang, Ming Liu, Bing Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Rongchuan Mu, Minghua Ma, Zekun Wang 0001, Ming Liu 0004, Bing Qin 0001
ACL (1)4
2026 Learning from Mistakes: Negative Reasoning Samples Enhance Out-of-Domain Generalization
abstract
Tian Xueyun, MingHua Ma, Bingbing Xu, Nuoyan Lyu, Wei Li, Heng Dong, Zheng Chu, Yuanzhuo Wang, Huawei Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tian Xueyun, Minghua Ma, Bingbing Xu 0001, Nuoyan Lyu, Wei Li 0176, Yuanzhuo Wang, Huawei Shen
ACL (1)2
2026 Enhanced Fairness Testing via Generating Effective Initial Individual Discriminatory Instances
abstract
Fairness testing aims at mitigating unintended discrimination in the decision-making process of data-driven AI systems. Individual discrimination may occur when an AI model makes different decisions for two distinct individuals who are distinguishable solely according to protected attributes, such as age and race. Such instances reveal biased AI behavior, and are called Individual Discriminatory Instances (IDIs). In this article, we propose an approach for the selection of the initial seeds to generate IDIs for fairness testing. Previous studies mainly used random initial seeds to this end. However, this phase is crucial, as these seeds are the basis of the follow-up IDI generation. We dubbed our proposed seed selection approach I&D . It generates a large number of initial IDIs exhibiting a great diversity, aiming at improving the overall performance of fairness testing. Our empirical study reveals that I&D is able to produce a larger number of IDIs with respect to four state-of-the-art IDI generation approaches, generating 1.86X more IDIs on average. When using the IDIs generated with I&D for retraining a machine learning model, the percentage of IDIs in the input space \(\mathbb{I}\) is decreased by 24.9% on average, implying that I&D is effective for improving the model’s fairness.
Zhao Tian 0002, Minghua Ma, Max Hort, Federica Sarro, Hongyu Zhang 0002, Junjie Chen 0003
ACM Trans. Softw. Eng. Methodol.2
2026 Failure Diagnosis in Microservice Systems: A Comprehensive Survey and Analysis
abstract
Widely adopted for their scalability and flexibility, modern microservice systems present unique failure diagnosis challenges due to their independent deployment and dynamic interactions. This complexity can lead to cascading failures that negatively impact operational efficiency and user experience. Recognizing the critical role of fault diagnosis in improving the stability and reliability of microservice systems, researchers have conducted extensive studies and achieved a number of significant results. This survey provides an exhaustive review of 98 scientific papers from 2003 to the present, including a thorough examination and elucidation of the fundamental concepts, system architecture, and problem statement. It also includes a qualitative analysis of the dimensions, providing an in-depth discussion of current best practices and future directions, aiming to further its development and application. In addition, this survey compiles publicly available datasets, toolkits, and evaluation metrics to facilitate the selection and validation of techniques for practitioners.
Shenglin Zhang, Sibo Xia, Wenzhao Fan, Binpeng Shi, Xiao Xiong, Zhenyu Zhong, Minghua Ma, Yongqian Sun, Dan Pei
ACM Trans. Softw. Eng. Methodol.7
2026 LLM-Enhanced Failure Localization in Microservices: Integrating Multi-Modal Data and Expert Interpretation
abstract
Failure localization in microservice environments is increasingly challenging. While large language models (LLMs) have shown promise in software engineering tasks, existing approaches struggle to effectively integrate multi-modal telemetry data (e.g., log, metric and trace) and provide interpretable results. This paper presents LocaleXpert, a novel failure localization system that combines specialized LLM-based agents with traditional AIOps methods to diagnose issues in microservice environments. LocaleXpert introduces three key innovations: (1) a modular pipeline that transforms metrics, logs, and traces into natural language descriptions that LLMs can effectively process, (2) specialized expert agents that analyze each data type and collaborate to identify root causes, and (3) an interpretation mechanism that produces clear, actionable explanations of its reasoning process. Evaluation results show that LocaleXpert significantly outperforming baseline approaches in both accuracy and interpretability. The system has been successfully deployed in Microsoft's AIOpsLab benchmark, demonstrating its effectiveness.
Zhenyu Zhong, Ruowei Fu, Minghua Ma, Shenglin Zhang, Yongqian Sun, Chetan Bansal, Dan Pei
IEEE Trans. Serv. Comput.3
2025 EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models
abstract
Zekun Wang, MingHua Ma, Zexin Wang, Rongchuan Mu, Liping Shan, Ming Liu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zekun Wang 0001, Minghua Ma, Rongchuan Mu, Liping Shan, Ming Liu 0004, Bing Qin 0001
ACL (1)2
2025 CFSP: An Efficient Structured Pruning Framework for LLMs with Coarse-to-Fine Activation Information
abstract
The colossal parameters and computational overhead of Large Language Models (LLMs) challenge their real-world applications. Network pruning, which targets unstructured or structured sparsity by removing redundant parameters, has recently been explored for LLM acceleration. Existing LLM pruning works focus on unstructured pruning, which typically requires special hardware support for a practical speed-up. In contrast, structured pruning can reduce latency on general devices. However, it remains a challenge to perform structured pruning efficiently and maintain performance, especially at high sparsity ratios. To this end, we introduce an efficient structured pruning framework named CFSP, which leverages both Coarse (interblock) and Fine-grained (intrablock) activation information as an importance criterion to guide pruning. The pruning is highly efficient, as it only requires one forward pass to compute feature activations. Specifically, we first allocate the sparsity budget across blocks based on their importance and then retain important weights within each block. In addition, we introduce a recovery fine-tuning strategy that adaptively allocates training overhead based on coarse-grained importance to further improve performance. Experimental results demonstrate that CFSP outperforms existing methods on diverse models across various sparsity budgets. Our code will be available at https://github.com/wyxscir/CFSP.
Yuxin Wang 0002, Minghua Ma, Zekun Wang 0001, Jingchang Chen, Liping Shan, Qing Yang 0033, Dongliang Xu, Ming Liu 0004, Bing Qin 0001
COLING2
2025 AllHands :Ask Me Anything on Large-scale Verbatim Feedback via Large Language Models
abstract
Verbatim feedback constitutes a valuable repository of user experiences, opinions, and requirements, crucial for data engineering and software development. Extracting meaningful insights from large-scale feedback data presents a significant challenge. This paper introduces Allhands, an innovative ana-lytic framework that transforms traditional large-scale feedback analysis tasks through a natural language interface, leveraging large language models (LLMs). Allhands performs initial classification and topic modeling on feedback to convert it into a structurally augmented format, enhancing accuracy, robustness and generalization with the aid of LLMs. Subsequently, an LLM-based code-first agent interprets users' diverse natural language questions about the feedback, automatically translates them into executable call of analytic tools or code, and delivers comprehensive multi-modal responses, including text, code, tables, and images. This eliminates the need for developing individual feedback analytic tools for each request, reducing human effort and making the system more accessible and flexible to users. We evaluate Allhands across three diverse feedback datasets, demonstrating its superior efficacy in all stages of analysis, from classification and topic modeling to providing an “ask me anything” experience with comprehensive, accurate, and human-readable responses. To the best of our knowl-edge, Allhands is the first comprehensive feedback analysis framework supporting diverse and customized insight extraction requirements through a natural language interface.
Chaoyun Zhang, Zicheng Ma, Shilin He, Si Qin, Minghua Ma, Xiaoting Qin, Yu Kang 0006, Yuyi Liang, Xiaoyu Gou, Yajie Xue, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001, Qi Zhang 0066
ICDE6
2025 An Empirical Study of Production Incidents in Generative AI Cloud Services
abstract
The ever-increasing demand for generative artificial intelligence (GenAI) has motivated cloud-based GenAI services such as Azure OpenAI Service. Like any large-scale cloud service, failures are inevitable in cloud-based GenAI services, resulting in user dissatisfaction and significant monetary losses. However, GenAI cloud services, featured by their massive parameter scales, hardware demands, and usage patterns, present unique challenges, including generated content quality issues and privacy concerns, compared to traditional cloud services. To understand the production reliability of GenAI cloud services, we analyzed production incidents from Microsoft spanning in the past four years. Our study (1) presents the general characteristics of GenAI cloud service incidents at different stages of the incident life cycle; (2) identifies the symptoms and impacts of these incidents on GenAI cloud service quality and availability; (3) uncovers why these incidents occurred and how they were resolved; (4) discusses open research challenges in terms of incident detection, triage, and mitigation, and sheds light on potential solutions.
Haoran Yan, Yinfang Chen, Minghua Ma, Ming Wen 0001, Shan Lu 0001, Shenglin Zhang, Tianyin Xu, Rujia Wang, Chetan Bansal, Saravan Rajmohan, Qingwei Lin, Chaoyun Zhang, Dongmei Zhang 0001
ISSRE3
2025 Too Many Cooks: Assessing the Need for Multi-Source Data in Microservice Failure Diagnosis
abstract
Microservice systems, characterized by their distributed nature and dynamic environments, pose significant challenges for failure diagnosis. Traditional failure diagnosis methods based on a single source of observability data, such as logs, metrics, or traces, fall short due to their inability to manage the complexity of inter-service communications. Consequently, in recent years, many methods based on multi-source observability data have been proposed, which promise a comprehensive analysis by integrating logs, metrics, traces. However, despite the promising potential of multi-source methods, we find that existing works often overlook a critical question: whether multi-source data is truly necessary for failure diagnosis. To address this, we conduct a systematic evaluation of eleven representative failure diagnosis methods based on multi-source observability data, using three public datasets and our own curated dataset. Our experiments focus on multiple aspects of multi-source datasets and methods. The results indicate that the quality of existing open-source datasets is inconsistent, and not all studied methods consistently perform well. Surprisingly, we find that adding more data sources does not necessarily improve the performance of microservice failure diagnosis in some cases.
Shenglin Zhang, Xiaoyu Feng, Runzhou Wang, Minghua Ma, Wenwei Gu, Yongqian Sun, Zedong Jia, Jinrui Sun, Dan Pei
ISSRE4
2025 Triangle: Empowering Incident Triage with Multi-Agent
abstract
As cloud service systems grow in scale and complexity, incidents that indicate unplanned interruptions and outages become unavoidable. Rapid and accurate triage of these incidents to the appropriate responsible teams is crucial to maintain service reliability and prevent significant financial losses. However, existing incident triage methods relying on manual operations and predefined rules often struggle with efficiency and accuracy due to the heterogeneity of incident data and the dynamic nature of domain knowledge across multiple teams.To solve these issues, we propose Triangle, an end-to-end incident triage system based on a Multi-Agent framework. Triangle leverages a semantic distillation mechanism to tackle the issue of semantic heterogeneity in incident data, enhancing the accuracy of incident triage. Additionally, we introduce multi-role agents and a negotiation mechanism to emulate human engineers’ workflows, effectively handling decentralized and dynamic domain knowledge from multiple teams. Furthermore, our system incorporates an automated troubleshooting information collection and mitigation mechanism, reducing the reliance on human labor and enabling fully automated end-to-end incident triage. Extensive experiments conducted on a real-world cloud production environment demonstrate that Triangle significantly improved incident triage accuracy (up to 97%) and reduced Time to Engage (TTE) by as much as 91%, demonstrating substantial operational impact across diverse cloud services.
Zhaoyang Yu 0002, Aoyang Fang, Minghua Ma, Jaskaran Singh Walia, Chaoyun Zhang, Shu Chi, Ze Li 0005, Murali Chintalapati, Xuchao Zhang, Rujia Wang, Chetan Bansal, Saravan Rajmohan, Qingwei Lin, Shenglin Zhang, Dan Pei, Pinjia He
ASE3
2025 Large Language Models can Deliver Accurate and Interpretable Time Series Anomaly Detection
abstract
Time series anomaly detection (TSAD) plays a crucial role in various industrial applications. Traditional deep learning TSAD models require extensive training data and operate as black boxes, lacking interpretability for detected anomalies. To address these challenges, we propose LLMAD, a novel TSAD method that employs Large Language Models (LLMs) to deliver accurate and interpretable TSAD results. LLMAD applies in-context anomaly detection by retrieving both positive and negative similar time series segments, significantly enhancing LLMs' effectiveness. Furthermore, LLMAD employs the Anomaly Detection Chain-of-Thought approach to mimic expert logic for its decision-making process. This further enhances its performance and enables LLMAD to provide explanations for their detections through versatile perspectives.
Chaoyun Zhang, Jiaxu Qian, Minghua Ma, Si Qin, Chetan Bansal, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001
KDD (2)4
2025 UFO: A UI-Focused Agent for Windows OS Interaction
abstract
Chaoyun Zhang, Liqun Li, Shilin He, Xu Zhang, Bo Qiao, Si Qin, Minghua Ma, Yu Kang, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Qi Zhang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Chaoyun Zhang, Liqun Li, Shilin He, Xu Zhang 0024, Bo Qiao 0001, Si Qin, Minghua Ma, Yu Kang 0006, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001, Qi Zhang 0066
NAACL (Long Papers)7
2025 AIOpsArena: Scenario-Oriented Evaluation and Leaderboard for AIOps Algorithms in Microservices
abstract
AIOps algorithms playa crucial role in the mainte-nance of microservice systems. Many previous benchmarks' per-formance leaderboard provides valuable guidance for selecting appropriate algorithms. However, existing AIOps benchmarks mainly utilize offline static datasets to evaluate algorithms. They cannot consistently evaluate the performance of algorithms using real-time datasets, and the operation scenarios for evaluation are static, which is insufficient for effective algorithm selection. To address these issues, we propose an evaluation-consistent and scenario-oriented evaluation framework named AIOpsArena. The core idea is to build a live microservice benchmark to generate real-time datasets and consistently simulate the specific operation scenarios on it. AIOpsArena supports different leaderboards by selecting specific algorithms and datasets according to the operation scenarios. It also supports the deployment of various types of algorithms, enabling algorithms hot-plugging. At last, we test AIOpsArena with typical microservice operation scenarios to demonstrate its efficiency and usability. Platform and a video demonstrating the functioning of AIOpsArena is available from https://github.com/AIOpsArena/aiopsarena.
Yongqian Sun, Jiaju Wang, Zhengdan Li, Xiaohui Nie, Minghua Ma, Shenglin Zhang, Yuhe Ji, Wen Long, Hengmao Chen, Yongnan Luo, Dan Pei
SANER5
2025 Bridging Edge and Cloud: A Knowledge-Enhanced Framework for Efficient Time Series Anomaly Detection
Shenglin Zhang, Minghua Ma, Yongqian Sun, Dan Pei
IEEE Trans. Serv. Comput.6
2024 Building AI Agents for Autonomous Clouds: Challenges and Design Principles
abstract
The rapid growth in the use of Large Language Models (LLMs) and AI Agents as part of software development and deployment is revolutionizing the information technology landscape. While code generation receives significant attention, a higher-impact application lies in using agents for the operational resilience of cloud services, which currently require significant human effort and domain knowledge. There is a growing interest in AI for IT Operations (AIOps), which aims to automate complex operational tasks, like fault localization and root cause analysis, reducing human intervention and customer impact. However, achieving the vision of autonomous and self-healing clouds through AIOps is hampered by the lack of standardized frameworks for building, evaluating, and improving AIOps agents. This vision paper lays the groundwork for such a framework by framing the requirements and then discussing design decisions that satisfy them. We also propose AIOpsLab, a prototype implementation leveraging agent-cloud-interface that orchestrates an application, injects real-time faults using chaos engineering, and interfaces with an agent to localize and resolve the faults. We report promising results and lay the groundwork to build a modular and robust framework for building, evaluating, and improving agents for autonomous clouds.
Manish Shetty, Yinfang Chen, Gagan Somashekar, Minghua Ma, Yogesh L. Simmhan, Xuchao Zhang, Jonathan Mace, Dax Vandevoorde, Pedro Henrique B. Las-Casas, Shachee Mishra Gupta, Suman Nath, Chetan Bansal, Saravan Rajmohan
SoCC4
2024 Deoxys: A Causal Inference Engine for Unhealthy Node Mitigation in Large-scale Cloud Infrastructure
abstract
The presence of unhealthy nodes in cloud infrastructure signals the potential failure of machines, which can significantly impact the availability and reliability of cloud services, resulting in negative customer experiences. Effectively addressing unhealthy node mitigation is therefore vital for sustaining cloud system performance. This paper introduces Deoxys, a causal inference engine tailored to recommending mitigation actions for unhealthy node in cloud systems to minimize virtual machine downtime and interruptions during unhealthy events. It employs double machine learning combined with causal forest to produce precise and reliable mitigation recommendations based solely on limited observational data collected from the historical unhealthy events. To enhance the causal inference model, Deoxys further incorporates a policy fallback mechanism based on model uncertainty and action overriding mechanisms to (i) improve the reliability of the system, and (ii) strike a good tradeoff between downtime reduction and resource utilization, thereby enhancing the overall system performance.
Chaoyun Zhang, Randolph Yao, Si Qin, Ze Li 0005, Shekhar Agrawal, Binit R. Mishra, Minghua Ma, Qingwei Lin, Murali Chintalapati, Dongmei Zhang 0001
SoCC8
2024 Automatic Root Cause Analysis via Large Language Models for Cloud Incidents
abstract
Ensuring the reliability and availability of cloud services necessitates efficient root cause analysis (RCA) for cloud incidents. Traditional RCA methods, which rely on manual investigations of data sources such as logs and traces, are often laborious, error-prone, and challenging for on-call engineers. In this paper, we introduce RCACopilot, an innovative on-call system empowered by the large language model for automating RCA of cloud incidents. RCACopilot matches incoming incidents to corresponding incident handlers based on their alert types, aggregates the critical runtime diagnostic information, predicts the incident's root cause category, and provides an explanatory narrative. We evaluate RCACopilot using a real-world dataset consisting of a year's worth of incidents from Microsoft. Our evaluation demonstrates that RCACopilot achieves RCA accuracy up to 0.766. Furthermore, the diagnostic information collection component of RCACopilot has been successfully in use at Microsoft for over four years.
Yinfang Chen, Huaibing Xie, Minghua Ma, Yu Kang 0006, Xin Gao 0017, Liu Shi, Yunjie Cao, Xuedong Gao, Ming Wen 0001, Jun Zeng 0006, Supriyo Ghosh, Xuchao Zhang, Chaoyun Zhang, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001, Tianyin Xu
EuroSys3
2024 Xpert: Empowering Incident Management with Query Recommendations via Large Language Models
abstract
Large-scale cloud systems play a pivotal role in modern IT infrastructure. However, incidents occurring within these systems can lead to service disruptions and adversely affect user experience. To swiftly resolve such incidents, on-call engineers depend on crafting domain-specific language (DSL) queries to analyze telemetry data. However, writing these queries can be challenging and time-consuming. This paper presents a thorough empirical study on the utilization of queries of KQL, a DSL employed for incident management in a large-scale cloud management system at Microsoft. The findings obtained underscore the importance and viability of KQL queries recommendation to enhance incident management.
Yuxuan Jiang 0016, Chaoyun Zhang, Shilin He, Minghua Ma, Si Qin, Yu Kang 0006, Yingnong Dang, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001
ICSE5
2024 Can We Trust Auto-Mitigation? Improving Cloud Failure Prediction with Uncertain Positive Learning
abstract
In the rapidly expanding domain of cloud computing, a variety of software services have been deployed in the cloud. To ensure the reliability of cloud services, prior studies focus on the prediction of failure instances, such as disks, nodes, switches, etc. The mitigation actions are initiated to resolve the underlying issue once the prediction output is positive. However, our real-world practice in Microsoft Azure revealed a decline in prediction accuracy, approximate 9%, after model retraining. The decrease is attributed to the mitigation actions, which can result in uncertain positive instances. Since these instances cannot be verified after mitigation, they may introduce additional noise into the model updating process. To the best of our knowledge, we are the first to identify this Uncertain Positive Learning (UPLearning) issue in the real-world cloud failure prediction scenario, and we design an Uncertain Positive Learning Risk Estimator (Uptake) approach to address this problem. By utilizing two real-world datasets for disk failure prediction and conducting node prediction experiments in Azure, which is a top-tier cloud provider serving millions of users. We demonstrate that our Uptake method can significantly enhance failure prediction accuracy by an average of 5%.
Minghua Ma, Pu Zhao 0004, Shuo Li 0013, Ze Li 0005, Murali Chintalapati, Yingnong Dang, Chetan Bansal, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001
ISSRE2
2024 Large Language Models Can Provide Accurate and Interpretable Incident Triage
abstract
Large-scale cloud services frequently experience incidents that can have a significant impact on their stability. Incident triage is a critical process that assigns incidents to dedicated teams for resolution. However, traditional rule-based methods, commonly employed in various systems, have limitations due to a finite set of rules that necessitate continuous updates, leading to suboptimal performance. Current state-of-the-art approaches primarily rely on textual information, utilizing classifiers or unsupervised clustering. Unfortunately, the abundance of textual information, combined with considerable noise, presents a significant challenge to the accuracy of these methods. To tackle these challenges, we introduce COMET, an innovative system that utilizes an AutoExtractor to filter out non-critical logs and employs a Large Language Model (LLM) for keyword extraction. This approach effectively mitigates the complexity arising from disordered textual information. Additionally, COMET incorporates significant domain knowledge during keyword extraction, enhancing the LLM’s comprehension of the text. We deployed COMET on multiple cloud services within Microsoft, where it has operated continuously for over six months. Offline and online evaluations have shown that COMET achieves enhanced accuracy and reduced Time to Mitigation (TTM).
Minghua Ma, Ze Li 0005, Yu Kang 0006, Chaoyun Zhang, Chetan Bansal, Murali Chintalapati, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001, Changhua Pei, Gaogang Xie
ISSRE3
2024 ART: A Unified Unsupervised Framework for Incident Management in Microservice Systems
abstract
Automated incident management is critical for large-scale microservice systems, including tasks such as anomaly detection (AD), failure triage (FT), and root cause localization (RCL). Currently, most techniques focus only on a single task, overlooking shared knowledge across closely related tasks. However, employing isolated models for managing multiple tasks may result in inefficiencies, delayed responses, a lack of systemic perspective, and complexity in updates and operations. Therefore we propose ART, an unsupervised framework that integrates a full-process solution covering Anomaly detection, failure Triage, and Root cause localization. It reaches the unification of multiple tasks by extracting the shared knowledge. Specifically, we first conduct an empirical study to analyze how the shared knowledge embedded in anomalous deviations manifests in AD, FT, and RCL. To better calculate deviations and extract shared knowledge, we sequentially model channel, temporal, and call dependencies using Transformer Encoder, GRU, and GraphSAGE, respectively. Then unified failure representations enhance the interpretability of abstract features with explicit semantic information, serving as the basis for unsupervised multitask solutions. Our evaluations on the datasets generated from two benchmark microservice systems demonstrate that ART outperforms existing methods in terms of AD (improving by 5.65% to 60.8%), FT (improving by 13.2% to 95.7%), and RCL (improving by 13.3% to 205%).
Yongqian Sun, Binpeng Shi, Mingyu Mao, Minghua Ma, Sibo Xia, Shenglin Zhang, Dan Pei
ASE4
2024 Giving Every Modality a Voice in Microservice Failure Diagnosis via Multimodal Adaptive Optimization
abstract
Microservice systems are inherently complex and prone to failures, which can significantly impact user experience. Existing diagnostic approaches based on single-modal data such as logs, metrics, or traces cannot comprehensively capture failure patterns. For those multimodal data-based failure diagnosis methods, the dominant modality can overshadow others, hindering low-yield modalities from fully leveraging their characteristics. This paper proposes Medicine, a modal-independent microservice failure diagnosis framework based on multimodal adaptive optimization. It encodes different modalities separately to retain their unique features and employs adaptive optimization to adjust the learning pace between modalities, thereby enhancing overall diagnostic performance. Experimental results demonstrate that Medicine outperforms existing single-modal and multimodal diagnostic approaches on three public datasets, with F1-score improving by 15.72% to 70.84%. Even in cases where individual modal data is missing or of lower quality, Medicine maintains high diagnostic accuracy.
Shenglin Zhang, Zedong Jia, Jinrui Sun, Minghua Ma, Zhengdan Li, Yongqian Sun, Canqun Yang, Dan Pei
ASE5
2024 End-to-End AutoML for Unsupervised Log Anomaly Detection
abstract
As modern software systems evolve towards greater complexity, ensuring their reliable operation has become a critical challenge. Log data analysis is vital in maintaining system stability, with anomaly detection being a key aspect. However, existing log anomaly detection methods heavily rely on manual effort from experts, lacking transferability across systems. This has led to the situation where to perform anomaly detection on a new dataset, the operators must have a high level of understanding of the dataset, make multiple attempts, and spend a lot of time to deploy an algorithm that performs well successfully. This paper proposes LogCraft, an end-to-end unsupervised log anomaly detection framework based on automated machine learning (AutoML). LogCraft automates feature engineering, model selection, and anomaly detection, reducing the need for specialized knowledge and lowering the threshold for algorithm deployment. Extensive evaluations on five public datasets demonstrate LogCraft's effectiveness, achieving an average F1 score of 0.899, which outperforms the second-best average F1 score of 0.847 obtained by existing unsupervised algorithms. According to our knowledge, LogCraft is the first attempt to extract fixed-dimensional vectors as latent representations from a complete log dataset. The proposed meta-feature extractor also exhibits promising potential for measuring log dataset similarity and guiding future log analytics research.
Shenglin Zhang, Yuhe Ji, Jiaqi Luan, Xiaohui Nie, Minghua Ma, Yongqian Sun, Dan Pei
ASE6
2024 Microservice Root Cause Analysis With Limited Observability Through Intervention Recognition in the Latent Space
abstract
Many failure root cause analysis (RCA) algorithms for microservices have been proposed with the widespread adoption of microservices systems. Existing algorithms generally focus on RCA with ranking single-level (e.g. metric-level or service-level) root cause candidates (RCCs) with comprehensive monitoring metrics. However, many heterogeneous RCCs exist with limited observability in real-world microservices systems. Further, we find that the limited observability may result in inaccurate RCA through real-world failures in eBay. In this paper, for the first time, we propose to "model RCCs as latent variables". The core idea is to infer the status of RCCs as latent variables with related monitoring metrics instead of directly extracting features from only the observable metrics. Based on this, we propose LatentScope, an unsupervised RCA framework with heterogeneous RCCs under limited observability. A dual-space graph is proposed to model both observable and unobservable variables, with many-to-many relationships between spaces. To achieve fast inference of latent variables and RCA, we propose the LatentRegressor algorithm, which includes Regression-based Latent-space Intervention Recognition (RLIR) to achieve intervention recognition-based RCA in latent space. LatentScope has been deployed in eBay's production environment and evaluated on both eBay's real-world failures and a testbed dataset. The evaluation results show that, compared with baseline algorithms, our model significantly improves the Top-1 recall by 9.7%-57.9%. The source code of LatentScope and the dataset are available at https://github.com/NetManAIOps/LatentScope.
Zhe Xie, Shenglin Zhang, Yitong Geng, Yao Zhang 0009, Minghua Ma, Xiaohui Nie, Zhenhe Yao, Longlong Xu, Yongqian Sun, Dan Pei
KDD5
2024 Pre-trained KPI Anomaly Detection Model Through Disentangled Transformer
abstract
In large-scale online service systems, numerous Key Performance Indicators (KPIs), such as service response time and error rate, are gathered in a time-series format. KPI Anomaly Detection (KAD) is a critical data mining problem due to its widespread applications in real-world scenarios. However, KAD faces the challenges of dealing with KPI heterogeneity and noisy data. We propose KAD-Disformer, a KPI Anomaly Detection approach through Disentangled Transformer. KAD-Disformer pre-trains a model on existing accessible KPIs, and the pre-trained model can be effectively "fine-tuned" to unseen KPI using only a handful of samples from the unseen KPI. We propose a series of innovative designs, including disentangled projection for transformer, unsupervised few-shot fine-tuning (uTune), and denoising modules, each of which significantly contributes to the overall performance. Our extensive experiments demonstrate that KAD-Disformer surpasses the state-of-the-art universal anomaly detection model by 13% in F1-score and achieves comparable performance using only 1/8 of the finetuning samples saving about 25 hours. KAD-Disformer has been successfully deployed in the real-world cloud system serving millions of users, attesting to its feasibility and robustness. Our code is available at https://github.com/NetManAIOps/KAD-Disformer.
Zhaoyang Yu 0002, Changhua Pei, Xin Wang 0001, Minghua Ma, Chetan Bansal, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001, Xidao Wen, Gaogang Xie, Dan Pei
KDD4
2024 Revisiting VAE for Unsupervised Time Series Anomaly Detection: A Frequency Perspective
abstract
Time series Anomaly Detection (AD) plays a crucial role for web systems. Various web systems rely on time series data to monitor and identify anomalies in real time, as well as to initiate diagnosis and remediation procedures. Variational Autoencoders (VAEs) have gained popularity in recent decades due to their superior de-noising capabilities, which are useful for anomaly detection. However, our study reveals that VAE-based methods face challenges in capturing long-periodic heterogeneous patterns and detailed short-periodic trends simultaneously. To address these challenges, we propose Frequency-enhanced Conditional Variational Autoencoder (FCVAE), a novel unsupervised AD method for univariate time series. To ensure an accurate AD, FCVAE exploits an innovative approach to concurrently integrate both the global and local frequency features into the condition of Conditional Variational Autoencoder (CVAE) to significantly increase the accuracy of reconstructing the normal data. Together with a carefully designed "target attention" mechanism, our approach allows the model to pick the most useful information from the frequency domain for better short-periodic trend construction. Our FCVAE has been evaluated on public datasets and a large-scale cloud system, and the results demonstrate that it outperforms state-of-the-art methods. This confirms the practical applicability of our approach in addressing the limitations of current VAE-based anomaly detection models.
Changhua Pei, Minghua Ma, Xin Wang 0001, Zhihan Li 0002, Dan Pei, Saravan Rajmohan, Dongmei Zhang 0001, Qingwei Lin, Haiming Zhang 0002, Gaogang Xie
WWW3
2023 How Different are the Cloud Workloads? Characterizing Large-Scale Private and Public Cloud Workloads
abstract
With the rapid development of cloud systems, an increasing number of service workloads are deployed in the private cloud and/or public cloud. Although large cloud providers such as Azure and Google have published workload traces in the past, prior work has not focused on analyzing and characterizing the differences between private and public cloud workloads in detail. Based on our experience working with Azure, one of the most widely used cloud platforms in the world, we find that the workload characteristics are different between the private and public cloud workloads. Specifically, compared with the public cloud workloads, the private cloud workloads tend to be more homogeneous in both deployment sizes and utilization patterns, more static with occasional bursts in deployment characteristics, and more region-agnostic regarding the sensitivity to deployed regions. Our findings gain several insights and implications on cloud management and motivate us to build a centralized workload knowledge base.
Xiaoting Qin, Minghua Ma, Yuheng Zhao, Anjaly Parayil, Chetan Bansal, Saravan Rajmohan, Íñigo Goiri, Eli Cortez, Si Qin, Qingwei Lin, Dongmei Zhang 0001
DSN2
2023 CODEC: Cost-Effective Duration Prediction System for Deadline Scheduling in the Cloud
abstract
Modern cloud platforms allow customers to flexibly allocate or release computing resources. One crucial scenario is how to drive existing VMs to a specific state by a given deadline in a reliable and cost-effective manner. These state transition requests could involve starting or reallocating VMs. Performing millions of these requests per day can be challenging because the throughput of cloud platforms is not consistently deterministic. To meet customer service level agreements, cloud providers often trade-off cost for reliability and use oversized estimates to ensure that they initiate state transitions against VMs well ahead of time. In this paper, we propose a COst-effective Duration prEdiCtion system (CODEC) that solves the deadline scheduling problem for cloud providers by building intelligent automation to discover and execute optimal strategies when performing concurrent requests. In the CODEC, the core is to categorize durations of requests into buckets in real-time and the buffer time of each bucket is then predicted by extreme value theory. Extensive experiments show the proposed approach can guarantee the success rate of requests, while significantly saving cost with 38.71% and 86.97% compared to the baseline approach and the static buffer time, respectively. The results obtained show that CODEC can effectively and efficiently solve the deadline scheduling problem, and lead a worldwide cloud provider to deploy its prototype.
Minghua Ma, Si Qin, Bo Qiao 0001, Randolph Yao, Harshwardhan Chaturvedi, Murali Chintalapati, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001
ISSRE2
2023 Robust Multimodal Failure Detection for Microservice Systems
abstract
Proactive failure detection of instances is vitally essential to microservice systems because an instance failure can propagate to the whole system and degrade the system's performance. Over the years, many single-modal (i.e., metrics, logs, or traces) databased anomaly detection methods have been proposed. However, they tend to miss a large number of failures and generate numerous false alarms because they ignore the correlation of multimodal data. In this work, we propose AnoFusion, an unsupervised failure detection approach, to proactively detect instance failures through multimodal data for microservice systems. It applies a Graph Transformer Network (GTN) to learn the correlation of the heterogeneous multimodal data and integrates a Graph Attention Network (GAT) with Gated Recurrent Unit (GRU) to address the challenges introduced by dynamically changing multimodal data. We evaluate the performance of AnoFusion through two datasets, demonstrating that it achieves the F1-score of 0.857 and 0.922, respectively, outperforming the state-of-the-art failure detection approaches.
Minghua Ma, Zhenyu Zhong, Shenglin Zhang, Zhiyuan Tan 0005, Xiao Xiong, LuLu Yu, Yongqian Sun, Dan Pei, Qingwei Lin, Dongmei Zhang 0001
KDD2
2023 TraceDiag: Adaptive, Interpretable, and Efficient Root Cause Analysis on Large-Scale Microservice Systems
abstract
Root Cause Analysis (RCA) is becoming increasingly crucial for ensuring the reliability of microservice systems. However, performing RCA on modern microservice systems can be challenging due to their large scale, as they usually comprise hundreds of components, leading significant human effort. This paper proposes TraceDiag, an end-to-end RCA framework that addresses the challenges for large-scale microservice systems. It leverages reinforcement learning to learn a pruning policy for the service dependency graph to automatically eliminates redundant components, thereby significantly improving the RCA efficiency. The learned pruning policy is interpretable and fully adaptive to new RCA instances. With the pruned graph, a causal-based method can be executed with high accuracy and efficiency. The proposed TraceDiag framework is evaluated on real data traces collected from the Microsoft Exchange system, and demonstrates superior performance compared to state-of-the-art RCA approaches. Notably, TraceDiag has been integrated as a critical component in the Microsoft M365 Exchange, resulting in a significant improvement in the system's reliability and a considerable reduction in the human effort required for RCA.
Ruomeng Ding, Chaoyun Zhang, Lu Wang 0029, Yong Xu 0010, Minghua Ma, Meng Zhang 0025, Qingjun Chen, Xin Gao 0017, Xuedong Gao, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001
ESEC/SIGSOFT FSE5
2023 Detection Is Better Than Cure: A Cloud Incidents Perspective
abstract
Cloud providers use automated watchdogs or monitors to continuously observe service availability and to proactively report incidents when system performance degrades. Improper monitoring can lead to delays in the detection and mitigation of production incidents, which can be extremely expensive in terms of customer impacts and manual toil from engineering resources. Therefore, a systematic understanding of the pitfalls in current monitoring practices and how they can lead to production incidents is crucial for ensuring continuous reliability of cloud services.
Vaibhav Ganatra, Anjaly Parayil, Supriyo Ghosh, Yu Kang 0006, Minghua Ma, Chetan Bansal, Suman Nath, Jonathan Mace
ESEC/SIGSOFT FSE5
2023 Assess and Summarize: Improve Outage Understanding with Large Language Models
abstract
Cloud systems have become increasingly popular in recent years due to their flexibility and scalability. Each time cloud computing applications and services hosted on the cloud are affected by a cloud outage, users can experience slow response times, connection issues or total service disruption, resulting in a significant negative business impact. Outages are usually comprised of several concurring events/source causes, and therefore understanding the context of outages is a very challenging yet crucial first step toward mitigating and resolving outages. In current practice, on-call engineers with in-depth domain knowledge, have to manually assess and summarize outages when they happen, which is time-consuming and labor-intensive. In this paper, we first present a large-scale empirical study investigating the way on-call engineers currently deal with cloud outages at Microsoft, and then present and empirically validate a novel approach (dubbed Oasis) to help the engineers in this task. Oasis is able to automatically assess the impact scope of outages as well as to produce human-readable summarization. Specifically, Oasis first assesses the impact scope of an outage by aggregating relevant incidents via multiple techniques. Then, it generates a human-readable summary by leveraging fine-tuned large language models like GPT-3.x. The impact assessment component of Oasis was introduced in Microsoft over three years ago, and it is now widely adopted, while the outage summarization component has been recently introduced, and in this article we present the results of an empirical evaluation we carried out on 18 real-world cloud systems as well as a human-based evaluation with outage owners. The results obtained show that Oasis can effectively and efficiently summarize outages, and lead Microsoft to deploy its first prototype which is currently under experimental adoption by some of the incident teams.
Pengxiang Jin, Shenglin Zhang, Minghua Ma, Yu Kang 0006, Liqun Li, Bo Qiao 0001, Chaoyun Zhang, Pu Zhao 0004, Shilin He, Federica Sarro, Yingnong Dang, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001
ESEC/SIGSOFT FSE3
2023 ImDiffusion: Imputed Diffusion Models for Multivariate Time Series Anomaly Detection
abstract
Anomaly detection in multivariate time series data is of paramount importance for large-scale systems. However, accurately detecting anomalies in such data poses significant challenges due to the need for precise data modeling capability. Existing forecasting and reconstruction-based methods struggle to address these challenges effectively. To overcome these limitations, we propose a novel anomaly detection framework named ImDiffusion, which combines time series imputation and diffusion models to achieve accurate and robust anomaly detection. The imputation-based approach employed by ImDiffusion leverages the information from neighboring values in the time series, enabling precise modeling of temporal and inter-correlated dependencies, reducing uncertainty in the data, thereby enhancing the robustness of the anomaly detection process. ImDiffusion further leverages diffusion models as time series imputers to accurately capture complex dependencies. We leverage the step-by-step denoised outputs generated during the inference process to serve as valuable signals for anomaly prediction, resulting in improved accuracy and robustness of the detection process. We evaluate the performance of ImDiffusion via extensive experiments on benchmark datasets. The results demonstrate that our proposed framework significantly outperforms state-of-the-art approaches in terms of detection accuracy and timeliness. ImDiffusion is further integrated into the real production system in Microsoft and observes a remarkable 11.4% increase in detection F1 score compared to the legacy approach. To the best of our knowledge, ImDiffusion represents a pioneering approach that combines imputation-based techniques with time series anomaly detection, while introducing the novel use of diffusion models to the field.
Chaoyun Zhang, Minghua Ma, Ruomeng Ding, Bowen Li 0002, Shilin He, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001
Proc. VLDB Endow.3
2023 Robust Failure Diagnosis of Microservice System Through Multimodal Data
abstract
Automatic failure diagnosis is crucial for large microservice systems. Currently, most failure diagnosis methods rely solely on single-modal data (i.e., using either metrics, logs, or traces). In this study, we conduct an empirical study using real-world failure cases to show that combining these sources of data (multimodal data) leads to a more accurate diagnosis. However, effectively representing these data and addressing imbalanced failures remain challenging. To tackle these issues, we proposeDiagFusion, a robust failure diagnosis approach that uses multimodal data. It leverages embedding techniques and data augmentation to represent the multimodal data of service instances, combines deployment data and traces to build a dependency graph, and uses a graph neural network to localize the root cause instance and determine the failure type. Our evaluations using real-world datasets show thatDiagFusionoutperforms existing methods in terms of root cause instance localization (improving by 20.9% to 368%) and failure type determination (improving by 11.0% to 169%).
Shenglin Zhang, Pengxiang Jin, Yongqian Sun, Bicheng Zhang, Sibo Xia, Zhengdan Li, Zhenyu Zhong, Minghua Ma, Wa Jin, Dan Pei
IEEE Trans. Serv. Comput.9
2022 Mining Fluctuation Propagation Graph Among Time Series with Active Learning
Mingjie Li 0005, Minghua Ma, Xiaohui Nie, Kanglin Yin, Xidao Wen, Zhiyun Yuan, Duogang Wu, Guoying Li, Dan Pei
DEXA (1)2
2022 Multi-task Hierarchical Classification for Disk Failure Prediction in Online Service Systems
abstract
One of the most common threats to online service system's reliability is disk failure. Many disk failure prediction techniques have been developed to predict failures before they actually occur, allowing proactive steps to be taken to minimize service disruption and increase service reliability. Existing approaches for disk failure prediction do not differentiate among various types of disk failure. In industrial practice, however, different product teams treat distinct types of disk failures as different prediction tasks in large-scale online service systems like Microsoft 365. For example, hardware operation team is concerned with physical disk errors, while database service team focuses on I/O delay. In this paper, we propose MTHC (Multi-Task Hierarchical Classification) to enhance the performance of disk failure prediction for each task via multi-task learning. In addition, MTHC introduces a novel hierarchy-aware mechanism to deal with the data imbalance problem, which is a severe issue in the area of disk failure prediction. We show that MTHC can be easily utilized to enhance most state-of-the-art disk failure prediction models. Our experiments on both industrial and public datasets demonstrate that such disk failure prediction models enhanced by MTHC performs much better than those models working without MTHC. Furthermore, our experiments also present that the hierarchical-aware mechanism underlying MTHC can alleviate the data imbalance problem and thus improve the practical performance of various disk failure prediction models. More encouragingly, the proposed MTHC has been successfully applied to Microsoft 365 online service systems, and averagely reduces the number of virtual machine interruptions by 10% per month.
Hailan Yang, Pu Zhao 0004, Minghua Ma, Chengwu Wen, Hongyu Zhang 0002, Chuan Luo 0002, Qingwei Lin, Chang Yi, Jiaojian Wang, Chenjian Zhang, Yingnong Dang, Saravan Rajmohan, Dongmei Zhang 0001
KDD4
2022 An empirical study of log analysis at Microsoft
abstract
Logs are crucial to the management and maintenance of software systems. In recent years, log analysis research has achieved notable progress on various topics such as log parsing and log-based anomaly detection. However, the real voices from front-line practitioners are seldom heard. For example, what are the pain points of log analysis in practice? In this work, we conduct a comprehensive survey study on log analysis at Microsoft. We collected feedback from 105 employees through a questionnaire of 13 questions and individual interviews with 12 employees. We summarize the format, scenario, method, tool, and pain points of log analysis. Additionally, by comparing the industrial practices with academic research, we discuss the gaps between academia and industry, and future opportunities on log analysis with four inspiring findings. Particularly, we observe a huge gap exists between log anomaly detection research and failure alerting practices regarding the goal, technique, efficiency, etc. Moreover, data-driven log parsing, which has been widely studied in recent research, can be alternatively achieved by simply logging template IDs during software development. We hope this paper could uncover the real needs of industrial practitioners and the unnoticed yet significant gap between industry and academia, and inspire interesting future directions that converge efforts from both sides.
Shilin He, Xu Zhang 0024, Pinjia He, Yong Xu 0010, Liqun Li, Yu Kang 0006, Minghua Ma, Yining Wei, Yingnong Dang, Saravanakumar Rajmohan, Qingwei Lin
ESEC/SIGSOFT FSE7
2022 An empirical investigation of missing data handling in cloud node failure prediction
abstract
Cloud computing systems have become increasingly popular in recent years. A typical cloud system utilizes millions of computing nodes as the basic infrastructure. Node failure has been identified as one of the most prevalent causes of cloud system downtime. To improve the reliability of cloud systems, many previous studies collected monitoring metrics from nodes and built models to predict node failures before the failures happen. However, based on our experience with large-scale real-world cloud systems in Microsoft, we find that the task of predicting node failure is severely hampered by missing data. There is a large amount of missing data, and the online latest data utilized for prediction is even worse. As a result, the real-time performance of the node prediction model is limited. In this paper, we first characterize the missing data problem for node failure prediction. Then, we evaluate several existing data interpolation approaches, and find that node dimension interpolation approaches outperform time dimension ones and deep learning based interpolation is the best for early prediction. Our findings can help academics and engineers address the missing data problem in cloud node failure prediction and other data-driven software engineering scenarios.
Minghua Ma, Yuang Tong, Pu Zhao 0004, Yong Xu 0010, Hongyu Zhang 0002, Shilin He, Lu Wang 0029, Yingnong Dang, Saravanakumar Rajmohan, Qingwei Lin
ESEC/SIGSOFT FSE1
2022 UniParser: A Unified Log Parser for Heterogeneous Log Data
abstract
Logs provide first-hand information for engineers to diagnose failures in large-scale online service systems. Log parsing, which transforms semi-structured raw log messages into structured data, is a prerequisite of automated log analysis such as log-based anomaly detection and diagnosis. Almost all existing log parsers follow the general idea of extracting the common part as templates and the dynamic part as parameters. However, these log parsing methods, often neglect the semantic meaning of log messages. Furthermore, high diversity among various log sources also poses an obstacle in the generalization of log parsing across different systems. In this paper, we propose UniParser to capture the common logging behaviours from heterogeneous log data. UniParser utilizes a Token Encoder module and a Context Encoder module to learn the patterns from the log token and its neighbouring context. A Context Similarity module is specially designed to model the commonalities of learned patterns. We have performed extensive experiments on 16 public log datasets and our results show that UniParser outperforms state-of-the-art log parsers by a large margin. 1
Xu Zhang 0024, Shilin He, Hongyu Zhang 0002, Liqun Li, Yu Kang 0006, Yong Xu 0010, Minghua Ma, Qingwei Lin, Yingnong Dang, Saravan Rajmohan, Dongmei Zhang 0001
WWW8
2021 Jump-Starting Multivariate Time Series Anomaly Detection for Online Service Systems
Minghua Ma, Shenglin Zhang, Junjie Chen 0003, Jim Xu, Yongliang Lin, Xiaohui Nie, Dan Pei
USENIX ATC1
2020 Diagnosing Root Causes of Intermittent Slow Queries in Large-Scale Cloud Databases
abstract
With the growing market of cloud databases, careful detection and elimination of slow queries are of great importance to service stability. Previous studies focus on optimizing the slow queries that result from internal reasons (e.g., poorly-written SQLs). In this work, we discover a different set of slow queries which might be more hazardous to database users than other slow queries. We name such queries Intermittent Slow Queries (iSQs), because they usually result from intermittent performance issues that are external (e.g., at database or machine levels). Diagnosing root causes of iSQs is a tough but very valuable task. This paper presents iSQUAD, Intermittent Slow QUery Anomaly Diagnoser, a framework that can diagnose the root causes of iSQs with a loose requirement for human intervention. Due to the complexity of this issue, a machine learning approach comes to light naturally to draw the interconnection between iSQs and root causes, but it faces challenges in terms of versatility, labeling overhead and interpretability. To tackle these challenges, we design four components, i.e., Anomaly Extraction, Dependency Cleansing, Type-Oriented Pattern Integration Clustering (TOPIC) and Bayesian Case Model. iSQUAD consists of an offline clustering & explanation stage and an online root cause diagnosis & update stage. DBAs need to label each iSQ cluster only once at the offline stage unless a new type of iSQs emerges at the online stage. Our evaluations on real-world datasets from Alibaba OLTP Database show that iSQUAD achieves an iSQ root cause diagnosis average F1-score of 80.4%, and outperforms existing diagnostic tools in terms of accuracy and efficiency.
Minghua Ma, Zheng Yin, Shenglin Zhang, Sheng Wang 0011, Christopher Zheng, Xinhao Jiang, Hanwen Hu, Nengjun Qiu, Feifei Li 0001, Changcheng Chen, Dan Pei
Proc. VLDB Endow.1
2019 Automatic and Generic Periodicity Adaptation for KPI Anomaly Detection
abstract
Key performance indicator (KPI) anomaly detection (AD) is critical to ensure service quality and reliability. Due to the effects of work days, off days, festivals, and business activities on user behavior, KPIs may exhibit different patterns within different days, which we call periodicity profiles of KPIs. However, existing KPI AD approaches have difficulties in adapting to diverse periodicity profiles due to the lack of generality. In this paper, we propose an automatic and generic framework called Period, which can accurately detect the periodicity profiles through daily subsequences clustering, and improve the performance of AD methods by robustly and automatically adapting to different periodicity profiles. In our evaluation using several real-world KPIs with different periodicity profiles from large Internet-based services, the clustering algorithm used to detect periodicity can achieve about 0.95 accuracy on average. More importantly, further evaluation on 56 KPIs shows that Period can significantly improve the best F-score of several widely used AD approaches by up to 0.66.
Nengwen Zhao, Jing Zhu 0007, Minghua Ma, Wenchi Zhang, Dan Pei
IEEE Trans. Netw. Serv. Manag.4
2018 Robust and Rapid Adaption for Concept Drift in Software System Anomaly Detection
abstract
Anomaly detection is critical for web-based software systems. Anecdotal evidence suggests that in these systems, the accuracy of a static anomaly detection method that was previously ensured is bound to degrade over time. It is due to the significant change of data distribution, namely concept drift, which is caused by software change or personal preferences evolving. Even though dozens of anomaly detectors have been proposed over the years in the context of software system, they have not tackled the problem of concept drift. In this paper, we present a framework, StepWise, which can detect concept drift without tuning detection threshold or per-KPI (Key Performance Indicator) model parameters in a large scale KPI streams, take external factors into account to distinguish the concept drift which under operators' expectations, and help any kind of anomaly detection algorithm to handle it rapidly. For the prototype deployed in Sogou, our empirical evaluation shows StepWise improve the average F-score by 206% for many widely-used anomaly detectors over a baseline without any concept drift detection.
Minghua Ma, Shenglin Zhang, Dan Pei, Hongwei Dai
ISSRE1
2017 You can hide, but your periodic schedule can't
abstract
The enterprise Wi-Fi networks enable the collection of large-scale users' trajectory datasets, which are highly desired for both research and commercial purposes. Meanwhile, releasing these mobility data also raises serious privacy concerns. A large body of work tries to achieve k-anonymity as the first step to solve the privacy problem and it has been qualitatively recognized that k-anonymity is still risky when the diversity of sensitive information in the k-anonymity set is low. However, there lacks a study that provides a quantitative understanding for trajectory data. In this work, we investigate the schedule-leakage risk for the first time, by presenting a large-scale measurement based analysis of the high schedule-leakage risk over sixteen weeks of trajectory data collected from Tsinghua University, a campus with 2,670 access points deployed in 111 buildings. Using this dataset, we recognize the high risk of the schedule-leakage, i.e., even when 4-anonymity is satisfied, 28% of individuals' schedules are totally disclosed, and 56% are partly disclosed.
Minghua Ma, Kaixin Sui, Yong Li 0008, Dan Pei
IWQoS1
2016 EDUM: classroom education measurements via large-scale WiFi networks
abstract
Behavior in classroom-based courses is hard to measure at large-scale. In this paper, we propose the EDUM (EDUcation Measurement) system to help characterize educational behavior through data collected from WLANs (WiFi networks) on campuses. EDUM characterizes students' punctuality (attendances, late arrivals, and early departures) for lectures using longitudinal WLAN data, and further characterizes the attractiveness of lectures using mobile phone's interactive states at minute-scale granularity. EDUM is easy to deploy and extensible for new types of data. We deploy EDUM at Tsinghua University where ~700 volunteer students' data are measured during a 9-week period by ~2,800 APs and two popular mobile apps. Our results show that EDUM makes it possible to obtain large-scale observations on punctuality, distraction and study performance, and quantitatively confirm or disprove numerous assumptions about educational behavior.
Mengyu Zhou, Minghua Ma, Yangkun Zhang, Kaixin Sui, Dan Pei, Thomas Moscibroda
UbiComp2
2016 WiFi can be the weakest link of round trip network latency in the wild
abstract
As mobile Internet is now indispensable in our daily lives, WiFi's latency performance has become critical to mobile applications' quality of experience. Unfortunately, WiFi hop latency in the wild remains largely unknown. In this paper, we first propose an effective approach to break down the round trip network latency. Then we provide the first systematic study on WiFi hop latency in the wild based on the latency and WiFi factors collected from 47 APs on T university campus for two months. We observe that WiFi hop can be the weakest link in the round trip network latency: more than 50% (10%) of TCP packets suffer from WiFi hop latency larger than 20ms (100ms), and WiFi hop latency occupies more than 60% in more than half of the round trip network latency. To help understand, troubleshoot, and optimize WiFi hop latency for WiFi APs in general, we train a decision tree model. Based on the model's output, we are able to reduce the median latency by 80% from 50ms to 10ms in one real case, and reduce the maximum latency from 250ms to 50ms in another real case.
Changhua Pei, Youjian Zhao, Guo Chen 0001, Ruming Tang, Yuan Meng 0002, Minghua Ma, Ken Ling, Dan Pei
INFOCOM6
2016 Your trajectory privacy can be breached even if you walk in groups
abstract
The enterprise Wi-Fi networks enable the collection of large-scale users' mobility information at an indoor level. The collected trajectory data is very valuable for both research and commercial purposes, but the use of the trajectory data also raises serious privacy concerns. A large body of work tries to achieve k-anonymity (hiding each user in an anonymity set no smaller than k) as the first step to solve the privacy problem. Yet it has been qualitatively recognized that k-anonymity is still risky when the diversity of the sensitive information in the k-anonymity set is low. There, however, still lacks a study that provides a quantitative understanding of that risk in the trajectory dataset. In this work, we present a large-scale measurement based analysis of the low-diversity risk over four weeks of trajectory data collected from Tsinghua, a campus that covers an area of 4 km2, on which 2,670 access points are deployed in 111 buildings. Using this dataset, we highlight the high risk of the low diversity. For example, we find that even when 5-anonymity is satisfied, the sensitive attributes of 25% of individuals can be easily guessed. We also find that although a larger k increases the size of anonymity sets, the corresponding improvement on the diversity of anonymity sets is very limited (decayed exponentially). These results suggest that diversity-oriented solutions are necessary.
Kaixin Sui, Youjian Zhao, Minghua Ma, Zimu Li, Dan Pei
IWQoS4
2016 Characterizing and Improving WiFi Latency in Large-Scale Operational Networks
abstract
WiFi latency is a key factor impacting the user experience of modern mobile applications, but it has not been well studied at large scale. In this paper, we design and deploy WiFiSeer, a framework to measure and characterize WiFi latency at large scale. WiFiSeer comprises a systematic methodology for modeling the complex relationships between WiFi latency and a diverse set of WiFi performance metrics, device characteristics, and environmental factors. WiFiSeer was deployed on Tsinghua campus to conduct a WiFi latency measurement study of unprecedented scale with more than 47,000 unique user devices. We observe that WiFi latency follows a long tail distribution and the 90th (99th) percentile is around 20 ms (250 ms). Furthermore, our measurement results quantitatively confirm some anecdotal perceptions about impacting factors and disapprove others. We deploy three practical solutions for improving WiFi latency in Tsinghua, and the results show significantly improved WiFi latencies. In particular, over 1,000 devices use our AP selection service based on a predictive WiFi latency model for 2.5 months, and 72% of their latencies are reduced by over half after they re-associate to the suggested APs.
Kaixin Sui, Mengyu Zhou, Minghua Ma, Dan Pei, Youjian Zhao, Zimu Li, Thomas Moscibroda
MobiSys4