EDBT 2026 Demo / reviewers in the wild / expert
Guangba Yu
dblp:248/2943
· DBLP profile ↗
40ranked-venue papers
9as first author
37since 2021 · last 2026
0000-0001-6195-9088ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 25 · 7 first-author · 23 since 2021Systems, architecture and hardware · 8 · 1 first-author · 7 since 2021Computer networks · 4 · 4 since 2021Security and privacy · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLMGuard: Multi-Agent Fault Diagnosis for Reliable Language-Model-as-a-Service
Yuedong Zhong, Guangba Yu, QunChao Fu, Yongqiang Yang, Michael R. Lyu |
DSN | 2 |
| 2026 | KPIRoot+: An efficient integrated framework for anomaly detection and root cause analysis in large-scale cloud systems
Wenwei Gu, Renyi Zhong, Guangba Yu, Xinying Sun, Jinyang Liu 0002, Yintong Huo, Zhuangbin Chen, Jianping Zhang 0002, Jiazhen Gu, Yongqiang Yang, Michael R. Lyu |
Empir. Softw. Eng. | 3 |
| 2026 | Logfun: An efficient function-Level log management framework for systems implemented with python
Min Li 0065, Gou Tan, Mingdong He, Guangba Yu, Pengfei Chen 0002, Chuanfu Zhang |
J. Syst. Softw. | 4 |
| 2026 | Fine-grained Tracing for Performance Anomaly Diagnosis of Serverless FunctionsabstractServerless function compositions subject to unpredictable faults are challenging to evaluate for root cause analysis. Even though distributed tracing provides observations at multiple levels of granularity for troubleshooting, excessive code instrumentation increases the tracing overheads in terms of both computation and storage. Therefore, developers face the challenge of where and how to instrument serverless functions to maximize the likelihood of locating faults based on tracing data while minimizing tracing overhead and costs. In this article, we propose a methodology to instrument an application with code-level tracing to infer the location of faults, taking into account constraints in terms of the maximum cost of the instrumentation and testing. We encode the tracing probe placement based on the control flow graph of the application and devise heuristics-based tracing data collection strategies to relate possible probe placements with their ability to locate a fault. Then we train novelty detection models to identify the internal anomalies and present an enhanced global search algorithm that automatically computes a probe placement with optimal fault localization ability versus cost. Experimental results show high performance in locating single and multiple faults with over 90% recall score for up to 15% latency anomalies, with minimal instrumentation overhead. Runan Wang, Guangba Yu, Giuliano Casale, Pengfei Chen 0002, Antonio Filieri |
ACM Trans. Auton. Adapt. Syst. | 2 |
| 2026 | A Survey on Failure Analysis and Fault Injection in AI SystemsabstractThe rapid advancement of AI has led to its integration into various areas, especially with Large Language Models (LLMs) significantly enhancing capabilities in Artificial Intelligence Generated Content (AIGC). However, the complexity of AI systems has also exposed their vulnerabilities, necessitating robust methods for Failure Analysis (FA) and Fault Injection (FI) to ensure resilience and reliability. Despite the importance of these techniques, there lacks a comprehensive review of FA and FI methodologies in AI systems. This study fills this gap by presenting a detailed survey of existing FA and FI approaches across six layers of AI systems. We systematically analyze 142 studies to answer three research questions including (1) what are the prevalent failures in AI systems, (2) what types of faults can current FI tools simulate, (3) what gaps exist between the simulated faults and real-world failures. Our findings reveal a taxonomy of AI system failures, assess the capabilities of existing FI tools, and highlight discrepancies between real-world and simulated failures. Moreover, this survey contributes to the field by providing a framework for fault diagnosis, evaluating the state-of-the-art in FI, and identifying areas for improvement in FI techniques to enhance the resilience of AI systems. Guangba Yu, Gou Tan, Haojia Huang, Pengfei Chen 0002, Roberto Natella, Zibin Zheng, Michael R. Lyu |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2025 | Mint: Cost-Efficient Tracing with All Requests Collection via Commonality and Variability AnalysisabstractDistributed traces contain valuable information but are often massive in volume, posing a core challenge in tracing framework design: balancing the tradeoff between preserving essential trace information and reducing trace volume. To address this tradeoff, previous approaches typically used a '1 or 0' sampling strategy: retaining sampled traces while completely discarding unsampled ones. However, based on an empirical study on real-world production traces, we discover that the '1 or 0' strategy actually fails to effectively balance this tradeoff. Haiyu Huang 0002, Cheng Chen 0056, Kunyi Chen, Pengfei Chen 0002, Guangba Yu, Yilun Wang 0001, Huxing Zhang, Qi Zhou 0001 |
ASPLOS (1) | 5 |
| 2025 | COCA: Generative Root Cause Analysis for Distributed Systems with Code KnowledgeabstractRuntime failures are commonplace in modern distributed systems. When such issues arise, users often turn to platforms such as Github or JIRA to report them and request assistance. Automatically identifying the root cause of these failures is critical for ensuring high reliability and availability. However, prevailing automatic root cause analysis (RCA) approaches rely significantly on comprehensive runtime monitoring data, which is often not fully available in issue platforms. Recent methods leverage large language models (LLMs) to analyze issue reports, but their effectiveness is limited by incomplete or ambiguous user-provided information. To obtain more accurate and comprehensive RCA results, the core idea of this work is to extract additional diagnostic clues from code to supplement data-limited issue reports. Specifically, we propose COCA, a code knowledge enhanced root cause analysis approach for issue reports. Based on the data within issue reports, COCA intelligently extracts relevant code snippets and reconstructs execution paths, providing a comprehensive execution context for further RCA. Subsequently, COCA constructs a prompt combining historical issue reports along with profiled code knowledge, enabling the LLMs to generate detailed root cause summaries and localize responsible components. Our evaluation on datasets from five real-world distributed systems demonstrates that COCA significantly outperforms existing methods, achieving a$\mathbf{2 8. 3 \%}$improvement in root cause localization and a 22.0 % improvement in root cause summarization. Furthermore, COCA's performance consistency across various LLMs underscores its robust generalizability. Yichen Li 0003, Jinyang Liu 0002, Zhuangbin Chen, Guangba Yu, Michael R. Lyu |
ICSE | 6 |
| 2025 | Take Kernel Stack Overhead Out: eBPF-Enhanced Network Acceleration for Distributed Training within EthernetabstractAs deep neural networks (DNN) continue to scale up in size to achieve greater capabilities, distributed training (DT) has become the prevailing approach to accelerate the training process.Through measurements and analysis of network communication overheads in DT scenarios within traditional Ethernet-based data centers, we observe that the Linux kernel network stack accounts for 30% to 40% of the total communication time, posing a significant bottleneck to training efficiency.As deep neural networks (DNN) continue to scale up in size to achieve greater capabilities, distributed training (DT) has become the prevailing approach to accelerate the training process.However, according to our observation on the network communication overheads in DT within Ethernet, the Linux kernel network stack accounts for 30% to 40% of the total communication time, posing a significant bottleneck to training efficiency.To mitigate the overhead introduced by the kernel network stack, we propose eRAR, an eBPF-based gradient aggregation over Ring-AR for DT tasks in traditional Ethernet-based data centers.eRAR offloads gradient aggregation to kernel using eBPF and avoids the overhead of network stack.eRAR has the advantages of hardwareagnostic, network-topology-independent, and resource-efficient.Our experimental results on four popular DNN models demonstrate that, compared to traditional TCP-based aggregation, eRAR improves the gradient aggregation throughput by 77.2%.Furthermore, eRAR reduces the communication time by up to 37.4% compared to existing systems.To mitigate the overhead introduced by the kernel network stack, we propose eRAR, an eBPF-based gradient aggregation over Ring-AR for DT tasks in commodity data centers.eRAR exploits Pengfei Chen 0002, Guangba Yu |
Internetware | 3 |
| 2025 | Causelens: Causality-Based Interpretable Root Cause Analysis for Microservice SystemsabstractMicroservice applications consist of complex API invocation relationships, where a single fault can propagate through multiple paths, leading to widespread failures. The diverse propagation patterns of different faults make efficient and interpretable root cause analysis (RCA) crucial. We propose CauseLens, a causality-based unsupervised RCA framework that improves both accuracy and interpretability. The key insight is that fine-grained causal modeling enhances root cause localization. CauseLens constructs a heterogeneous causal diagram at the operation and entity levels using normal monitoring data (i.e., metrics and traces) and trains a structural causal model. It then integrates reconstruction error and counterfactual analysis to identify root causes while revealing fault propagation paths. Experiments on two microservice datasets demonstrate that CauseLens outperforms state-of-the-art methods in RCA accuracy. Further ablation studies and parameter experiments validate its design, while overhead analysis confirms its feasibility for real-time RCA in production environments. Qihan Liu, Pengfei Chen 0002, Guangba Yu, Yuanhao Lai |
IWQoS | 3 |
| 2025 | iKnow: an Intent-Guided Chatbot for Cloud Operations with Retrieval-Augmented GenerationabstractManaging complex cloud services requires standard operational documentation, but its sheer volume often hinders cloud engineers from efficient knowledge acquisition. Retrieval-Augmented Generation (RAG) can streamline this process by retrieving relevant knowledge and generating concise, referenced answers. However, deploying a reliable RAG-based chatbot for cloud operation remains a challenge. In this experience paper, we analyze the development and deployment of RAG-based chatbots for operational question answering (OpsQA) at a large-scale cloud vendor. Through an empirical study of 2,000 real-world queries across three operational teams, we identify five unique OpsQA intent types (e.g., symptom analysis and terminology explanation) and their corresponding requirements for a satisfactory answer, which differ from general software engineering queries. Our analysis further uncovers six root causes leading to chatbot failures—over half stem from query issues (i.e., incompleteness, out-of-scope, or invalid queries), while others are from retrieval or generation issues. To address these issues, we propose iKnow, an intent-guided RAG-based chatbot that integrates intent detection, query rewriting tailored to each intent, and missing knowledge detection to enhance answer quality. In internal evaluations, iKnow improves average answer accuracy from 65.8% to 81.3% with only a modest increase in latency. iKnow has been deployed for six months at CloudA, supporting thousands of cloud engineers in daily operations. We discuss lessons learned from real-world deployment, providing valuable insights for future research and practical implementations in similar domains. Junjie Huang 0008, Yuedong Zhong, Guangba Yu, Minzhi Yan, Wenfei Luan, Michael R. Lyu |
ASE | 3 |
| 2025 | AlertGuardian: Intelligent Alert Life-Cycle Management for Large-scale Cloud SystemsabstractAlerts are critical for detecting anomalies in large-scale cloud systems, ensuring reliability and user experience. However, current systems generate overwhelming volumes of alerts, degrading operational efficiency due to ineffective alert life-cycle management. This paper details the efforts of Company-X to optimize alert life-cycle management, addressing alert fatigue in cloud systems. We propose AlertGuardian, a framework collaborating large language models (LLMs) and lightweight graph models to optimize the alert life-cycle through three phases: Alert Denoise uses graph learning model with virtual noise to filter noise, Alert Summary employs Retrieval Augmented Generation (RAG) with LLMs to create actionable summary, and Alert Rule Refinement leverages multi-agent iterative feedbacks to improve alert rule quality. Evaluated on four real-world datasets from Company-X’s services, AlertGuardian significantly mitigates alert fatigue (94.8% alert reduction ratios) and accelerates fault diagnosis (90.5% diagnosis accuracy). Moreover, AlertGuardian improves 1,174 alert rules, with 375 accepted by SREs (32% acceptance rate). Finally, we share success stories and lessons learned about alert life-cycle management after the deployment of AlertGuardian in Company-X. Guangba Yu, Genting Mai, Pengfei Chen 0002, Long Pan |
ASE | 1 |
| 2025 | Conan: Uncover Consensus Issues in Distributed Databases Using Fuzzing-Driven Fault InjectionabstractConsensus is critical for distributed databases as it ensures the consistency of states across nodes, reinforcing the robustness of the overall system. However, faults related to the consensus protocols such as Paxos can lead to serious issues in distributed databases. Such consensus issues impact the correctness and availability of these databases. Therefore, to automatically uncover consensus issues in distributed databases, we propose Conan, a framework designed with fuzzing-driven fault injection. Conan applies a state-guided fuzzing algorithm to effectively explore the fault search space. Moreover, Conan employs hybrid fault sequences that combines fine-grained message-level faults and coarse-grained system-level faults to enhance fault injection. We implement and evaluate Conan on 3 widely-used distributed databases, including etcd, rqlite and openGauss. Finally, Conan has successfully uncovered previously unknown consensus issues, some of which are not detected by existing approaches. Haojia Huang, Pengfei Chen 0002, Guangba Yu, Haiyu Huang 0002, Jia Chang |
SANER | 3 |
| 2025 | NetScope: Fault Localization in Programmable Networking Systems With Low-Cost In-Band Network Telemetry and In-Network DetectionabstractRecently, Software Defined Networking (SDN) has gained widespread adoption as a network infrastructure. Although the openness and programmability of SDN facilitate large complex network construction, diagnosing faults in datacenter-scale network remains challenging. Previous network diagnosis tools pose significant overhead in fine-grained telemetry and typically lack automated fine-grained fault diagnosis capabilities. Although on-demand monitoring methods have been proposed to reduce telemetry overhead, they struggle with effectively setting fixed thresholds, which requires expert experience. This paper presents NetScope, a lightweight system for real-time anomaly detection with self-adaptive thresholds and automatic root cause localization in programmable networking systems. NetScope estimates latency medians for each Flow (i.e., a pair of source and sink switches) within the switch using the proposed per-Flow quantile sketch and calculates the threshold accordingly for anomaly detection. Upon detecting anomalies, NetScope collects aggregated packet-level telemetry on demand and generates a ranked list of fine-grained fault culprits at multiple levels, including port-level, Flow-level, and switch-level. Extensive experiments demonstrate the effectiveness and efficiency of NetScope in anomaly detection and fault localization. Specifically, NetScope achieves a 32%~116% relative improvement in anomaly detection and 6%~197% improvement in root cause analysis compared with other baselines without causing any network bandwidth in anomaly detection while consuming 64.2% less telemetry bandwidth for localization. Hongyang Chen 0002, Benran Wang, Guangba Yu, Pengfei Chen 0002, Chen Sun 0005, Zibin Zheng |
IEEE Trans. Netw. | 3 |
| 2025 | Subgraphs as First-Class Citizens in Incident Management for Large-Scale Online Systems: An Evolution-Aware Framework
Pengfei Chen 0002, Yu Luo 0019, Qiuyu Yan, Hongyang Chen 0002, Guangba Yu, Zibin Zheng |
IEEE Trans. Software Eng. | 6 |
| 2024 | CTuner: Automatic NoSQL Database Tuning with Causal Reinforcement LearningabstractThe rapid development of information technology has necessitated the management of large volumes of data in modern society, leading to the emergence of NoSQL databases (e.g., MongoDB). To meet the huge demand for efficient data management and query, optimizing the performance of these databases has become crucial. Currently, some reinforcement learning-based methods have been used to improve the efficiency of databases by tuning customizable database configurations. However, these methods have limitations: they ignore operating system configurations, incur high training costs with more knobs, and adapt poorly to new environments with varying workloads and hardware. To address these issues, we propose a novel and effective approach named CTuner for the online performance tuning of NoSQL databases. CTuner skips cold start by Bayesian optimization-based learning, and improves the exploitation strategy of the Twin Delayed Deep Deterministic Policy Gradient (TD3) model with causal inference. Practical implementation and experimental evaluations on three prominent NoSQL databases show that CTuner can find a better configuration at the same time cost than state-of-the-art approaches, with up to a 27.4% improvement in throughput and up to 13.2 % reduction in 95 %-tail latency. Moreover, we introduce meta-learning to enhance the adaptability of CTuner and confirm that it is able to reliably improve performance under new environments and workloads. Genting Mai, Guangba Yu, Pengfei Chen 0002 |
Internetware | 3 |
| 2024 | FaaSRCA: Full Lifecycle Root Cause Analysis for Serverless ApplicationsabstractServerless becomes popular as a novel computing paradigms for cloud native services. However, the complexity and dynamic nature of serverless applications present significant challenges to ensure system availability and performance. There are many root cause analysis (RCA) methods for microservice systems, but they are not suitable for precise modeling serverless applications. This is because: (1) Compared to microservice, serverless applications exhibit a highly dynamic nature. They have short lifecycle and only generate instantaneous pulse-like data, lacking long-term continuous information. (2) Existing methods solely focus on analyzing the running stage and overlook other stages, failing to encompass the entire lifecycle of serverless applications. To address these limitations, we propose FaaSRCA, a full lifecycle root cause analysis method for serverless applications. It integrates multi-modal observability data generated from platform and application side by using Global Call Graph. We train a Graph Attention Network (GAT) based graph autoencoder to compute reconstruction scores for the nodes in global call graph. Based on the scores, we determine the root cause at the granularity of the lifecycle stage of serverless functions. We conduct experimental evaluations on two serverless benchmarks, the results show that FaaSRCA outperforms other baseline methods with a top-k precision improvement ranging from 21.25% to 81.63%. Pengfei Chen 0002, Guangba Yu, Yilun Wang 0001, Haiyu Huang 0002 |
ISSRE | 3 |
| 2024 | FaaSConf: QoS-aware Hybrid Resources Configuration for Serverless WorkflowsabstractServerless computing, also known as Function-as-a-Service (FaaS), is a significant development trend in modern software system architecture. The workflow composition of multiple short-lived functions has emerged as a prominent pattern in FaaS, exposing a considerable resources configuration challenge compared to individual independent serverless functions. This challenge unfolds in two ways. Firstly, workflows frequently encounter dynamic and concurrent user workloads, increasing the risk of QoS violations. Secondly, the performance of a function can be affected by the resource reprovision of other functions within the workflow. Yilun Wang 0001, Pengfei Chen 0002, Yiwen Zhang 0001, Guangba Yu, Haiyu Huang 0002 |
ASE | 5 |
| 2024 | Network shortcut in data plane of service mesh with eBPF
Wanqi Yang, Pengfei Chen 0002, Guangba Yu, Huxing Zhang |
J. Netw. Comput. Appl. | 3 |
| 2024 | MicroFI: Non-Intrusive and Prioritized Request-Level Fault Injection for Microservice ApplicationsabstractMicroservice is a widely-adopted architecture for constructing cloud-native applications. To test application resiliency, chaos engineering is widely used to inject faults proactively in applications. However, the searching space formed by possible injection locations is huge due to the scale and complexity of the application. Although some methods are proposed to effectively explore injection space, they cannot prioritize high-impact injection solutions. Additionally, the blast radius of faults injected by existing methods is typically full of uncertainty, causing faults of multiple application functions. Although some tools are designed to conduct request-level injection, they require instrumentation on application code. To tackle these problems, this paper presents MicroFI, a non-intrusive fault injection framework, aiming to efficiently test different application functions with request-level injection. Request-level injection limits the blast radius to specified requests without any source code modification. Additionally, MicroFI leverages historical injection results and parallel technique to accelerate the searching. Moreover, An enhanced PageRank is used to measure the impact of faults and prioritize high-impact faults that fail more functions. Evaluations on three microservice applications show that MicroFI precisely injects faults and reduces up to 91% redundant faults on average. Additionally, by employing prioritization, MicroFI reduces an average of 47.3% injection budgets to cover all high-impact faults. Hongyang Chen 0002, Pengfei Chen 0002, Guangba Yu |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2023 | MARS: Fault Localization in Programmable Networking Systems with Low-cost In-Band Network TelemetryabstractRecently, the adoption of Software Defined Networking (SDN) as a network infrastructure has gained significant popularity. Although the openness and programmability of SDN ease the construction of large complex networks, it is still challenging to diagnose faults in a complex datacenter-scale network, which is crucial to guarantee rigorous service level agreement (SLA) of upper-layer applications. Previous network diagnosis tools incur significant overhead in fine-grained telemetry, and usually lack the ability to automatically diagnose fine-grained faults. Although on-demand monitoring methods is proposed to reduce telemetry overhead, they struggle to effectively set static thresholds, which requires expert experience. In this paper, we present MARS, a lightweight system for anomaly detection with dynamic threshold and automatic root cause localization in programmable networking systems. MARS collects aggregated packet-level telemetry on demand and generates a ranked list of fine-grained fault culprits at multiple levels, including port-level, switch-level, and flow-level. Experimental evaluations show the cost-effectiveness of MARS, both in terms of network bandwidth and switch memory usage. Moreover, MARS achieves a 0.97 F1 score in anomaly detection, and 0.95 Recall at Top-2 and an overall 0.3 Exam Score in root cause localization. Benran Wang, Hongyang Chen 0002, Pengfei Chen 0002, Guangba Yu |
ICPP | 5 |
| 2023 | DeepPower: Deep Reinforcement Learning based Power Management for Latency Critical Applications in Multi-core SystemsabstractLatency-critical (LC) applications are widely deployed in modern datacenters. Effective power management for LC applications can yield significant cost savings. However, it poses a significant challenge in maintaining the desired Service Level Aggrement (SLA) levels. Prior researches have mainly emphasized predicting the service time of request and utilize heuristic algorithms for CPU frequency adjustment. Unfortunately, the control granularity is limited to the request level and manual feature selection is needed. Jingrun Zhang, Guangba Yu, Liang Ai, Pengfei Chen 0002 |
ICPP | 2 |
| 2023 | LogReducer: Identify and Reduce Log Hotspots in Kernel on the FlyabstractModern systems generate a massive amount of logs to detect and diagnose system faults, which incurs expensive storage costs and runtime overhead. After investigating real-world production logs, we observe that most of the logging overhead is due to a small number of log templates, referred to as log hotspots. Therefore, we conduct a systematical study about log hotspots in an industrial system WeChat, which motivates us to identify log hotspots and reduce them on the fly. In this paper, we propose LogReducer, a non-intrusive and language-independent log reduction framework based on eBPF (Extended Berkeley Packet Filter), consisting of both online and offline processes. After two months of serving the offline process of LogReducer in WeChat, the log storage overhead has dropped from 19.7 PB per day to 12.0 PB (i.e., about a 39.08% decrease). Practical implementation and experimental evaluations in the test environment demonstrate that the online process of LogReducer can control the logging overhead of hotspots while preserving logging effectiveness. Moreover, the log hotspot handling time can be reduced from an average of 9 days in production to 10 minutes in the test with the help of LogReducer, Guangba Yu, Pengfei Chen 0002, Pairui Li, Tianjun Weng, Haibing Zheng, Yuetang Deng, Zibin Zheng |
ICSE | 1 |
| 2023 | MARS: Fault Localization in Programmable Networking Systems with Low-cost In-Band Network TelemetryabstractThis paper presents MARS, a lightweight system for anomaly detection with dynamic threshold and automatic root cause localization in programmable networking systems. MARS collects aggregated packet-level telemetry on demand and generates a ranked list of fine-grained fault culprits at port-level, switch-level, and flow-level. Benran Wang, Hongyang Chen 0002, Pengfei Chen 0002, Guangba Yu |
IWQoS | 5 |
| 2023 | DiagConfig: Configuration Diagnosis of Performance Violations in Configurable Software SystemsabstractPerformance degradation due to misconfiguration in software systems that violates SLOs (service-level objectives) is commonplace. Diagnosing and explaining the root causes of such performance violations in configurable software systems is often challenging due to their increasing complexity. Although there are many tools and techniques for diagnosing performance violations, they provide limited evidence to attribute causes of observed performance violations to specific configurations. This is because the configuration is not originally considered in those tools. This paper proposes DiagConfig, specifically designed to conduct configuration diagnosis of performance violations. It leverages static code analysis to track configuration option propagation, identifies performance-sensitive options, detects performance violations, and constructs cause-effect chains that help stakeholders better understand the relationship between configuration and performance violations. Experimental evaluations with eight real-world software demonstrate that DiagConfig produces fewer false positives than a state-of-the-art documentation analysis-based tool (i.e., 5 vs 41) in the identification of performance-sensitive options, and outperforms a statistics-based debugging tool in the diagnosis of performance violations caused by configuration changes, offering more comprehensive results (recall: 0.892 vs 0.289). Moreover, we also show that DiagConfig can accelerate auto-tuning by compressing configuration space. Pengfei Chen 0002, Guangba Yu, Genting Mai |
ESEC/SIGSOFT FSE | 4 |
| 2023 | Nezha: Interpretable Fine-Grained Root Causes Analysis for Microservices on Multi-modal Observability DataabstractRoot cause analysis (RCA) in large-scale microservice systems is a critical and challenging task. To understand and localize root causes of unexpected faults, modern observability tools collect and preserve multi-modal observability data, including metrics, traces, and logs. Since system faults may manifest as anomalies in different data sources, existing RCA approaches that rely on single-modal data are constrained in the granularity and interpretability of root causes. In this study, we present Nezha, an interpretable and fine-grained RCA approach that pinpoints root causes at the code region and resource type level by incorporative analysis of multi-modal data. Nezha transforms heterogeneous multi-modal data into a homogeneous event representation and extracts event patterns by constructing and mining event graphs. The core idea of Nezha is to compare event patterns in the fault-free phase with those in the fault-suffering phase to localize root causes in an interpretable way. Practical implementation and experimental evaluations on two microservice applications show that Nezha achieves a high top1 accuracy (89.77%) on average at the code region and resource type level and outperforms state-of-the-art approaches by a large margin. Two ablation studies further confirm the contributions of incorporating multi-modal data. Guangba Yu, Pengfei Chen 0002, Hongyang Chen 0002, Zibin Zheng |
ESEC/SIGSOFT FSE | 1 |
| 2023 | TraceRank: Abnormal service localization with dis-aggregated end-to-end tracing data in cloud native systemsabstractAbstract Modern cloud native applications are generally built with a microservice architecture. To tackle various performance problems among a large number of services and machines, an end‐to‐end tracing tool is always equipped in these systems to track the execution path of every single request. However, it is nontrivial to conduct root cause analysis of anomalies with such a large volume of tracing data. This paper proposes a novel system named TraceRank to identify and locate abnormal services causing performance problems with dis‐aggregated end‐to‐end traces. TraceRank mainly includes an anomaly detection module and a root cause analysis module. The root cause analysis procedure is triggered when an anomaly is detected. To fully leverage the information provided by the tracing data, both the spectrum analysis and the PageRank‐based random walk methods are introduced to pinpoint abnormal services. The experiments in TrainTicket and Bookinfo microservice benchmarks and a real‐world system show that TraceRank can locate root causes with 90% in Precision and 86% in Recall. TraceRank has up to 10% improvement compared with several state‐of‐the‐art approaches in both Precision and Recall. Finally, TraceRank has good scalability and a low overhead to adapt to large‐scale microservice systems. Guangba Yu, Pengfei Chen 0002 |
J. Softw. Evol. Process. | 1 |
| 2023 | SwissLog: Robust Anomaly Detection and Localization for Interleaved Unstructured LogsabstractModern distributed systems generate interleaved logs when running in parallel. Identifiers (ID) are always attached to them to trace running instances or entities in logs. Therefore, log messages can be grouped by the same IDs to help anomaly detection and localization. The existing approaches to achieve this still fall short meeting these challenges: 1) Log is solely processed in single components without mining log dependencies. 2) Log formats are continually changing in modern software systems. 3) It is challenging to detect latent performance issues non-intrusively by trivial monitoring tools. To remedy the above shortcomings, we propose SwissLog, a robust anomaly detection and localization tool for interleaved unstructured logs. SwissLog focuses on log sequential anomalies and tries to dig out possible performance issues. SwissLog constructs ID relation graphs across distributed components and groups log messages by IDs. Moreover, we propose an online data-driven log parser without parameter tuning. The grouped log messages are parsed via the novel log parser and transformed with semantic and temporal embedding. Finally, SwissLog utilizes an attention-based Bi-LSTM model and a heuristic searching algorithm to detect and localize anomalies in instance-granularity, respectively. The experiments on real-world and synthetic datasets confirm the effectiveness, efficiency, and robustness of SwissLog. Pengfei Chen 0002, Linxiao Jing, Guangba Yu |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2023 | A Spatiotemporal Deep Learning Approach for Unsupervised Anomaly Detection in Cloud SystemsabstractAnomaly detection is a critical task for maintaining the performance of a cloud system. Using data-driven methods to address this issue is the mainstream in recent years. However, due to the lack of labeled data for training in practice, it is necessary to enable an anomaly detection model trained on contaminated data in an unsupervised way. Besides, with the increasing complexity of cloud systems, effectively organizing data collected from a wide range of components of a system and modeling spatiotemporal dependence among them become a challenge. In this article, we propose TopoMAD, a stochastic seq2seq model which can robustly model spatial and temporal dependence among contaminated data. We include system topological information to organize metrics from different components and apply sliding windows over metrics collected continuously to capture the temporal dependence. We extract spatial features with the help of graph neural networks and temporal features with long short-term memory networks. Moreover, we develop our model based on variational auto-encoder, enabling it to work well robustly even when trained on contaminated data. Our approach is validated on the run-time performance data collected from two representative cloud systems, namely, a big data batch processing system and a microservice-based transaction processing system. The experimental results show that TopoMAD outperforms some state-of-the-art methods on these two data sets. Pengfei Chen 0002, Yongfeng Wang, Guangba Yu, Cailin Chen, Zibin Zheng |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | FaaSDeliver: Cost-Efficient and QoS-Aware Function Delivery in Computing ContinuumabstractServerless Function-as-a-Service (FaaS) is a rapidly growing computing paradigm in the cloud era. To provide rapid service response and save network bandwidth, traditional cloud-based FaaS platforms have been extended to the edge. However, launching functions in a heterogeneous computing continuum (HCC) that includes the cloud, fog, and the edge brings new challenges: determiningwhere functions should be delivered and how many resources should be allocated.To optimize the cost of running functions in the HCC, we propose an adaptive and efficient function delivery engine, namedFaaSDeliver, which automatically unearths a cost-efficient function delivery policy (FDP) for each function, including the FaaS platform selection and resource allocation. Real system implementation and evaluations in a practical HCC demonstrate thatFaaSDelivercan unearth the most cost-efficient FDPs from among 180,200 FDPs after a few trials.FaaSDeliverreduces the average cost of function execution from 38% to 78% compared to some state-of-the-art approaches. Guangba Yu, Pengfei Chen 0002, Zibin Zheng, Jingrun Zhang |
IEEE Trans. Serv. Comput. | 1 |
| 2022 | MicroSketch: Lightweight and Adaptive Sketch Based Performance Issue Detection and Localization in Microservice Systems
Guangba Yu, Pengfei Chen 0002, Chuanfu Zhang, Zibin Zheng |
ICSOC | 2 |
| 2022 | TS-InvarNet: Anomaly Detection and Localization based on Tempo-spatial KPI Invariants in Distributed ServicesabstractModern industrial systems are often large-scale distributed systems composed of dozens to thousands of services, leading to difficulty in anomaly detection and localization. KPIs (Key Performance Indicators) record the states of different services and are presented as time series, which reflect the status of the system. However, due to the dynamic and complex periodic patterns embedded in KPIs, pinpointing anomalous behavior of these multivariate time series data quickly and accurately is a challenging problem. The current state-of-the-art deep-learning-based anomaly detection methods model global inter-KPI dependency, causing the limited ability to detect local subtle anomalies and poor interpretability. In practice, interpreting anomalies can accelerate problem localization and further troubleshooting. In this study, we propose TS-InvarNet, an interpretable end-to-end anomaly detection and diagnosis framework based on tempo-spatial KPI invariants. Extensive empirical studies on three real-world industrial datasets and a widely-used open-source system demonstrate that TS-InvarNet can outperform state-of-the-art baseline methods in detection and diagnosis performance. Specifically, TS-InvarNet increases best F1-scores by up to 27% compared to the baselines. Zijun Hu, Pengfei Chen 0002, Guangba Yu |
ICWS | 3 |
| 2022 | Going through the Life Cycle of Faults in Clouds: Guidelines on Fault HandlingabstractFaults are the primary culprits of breaking the high availability of cloud systems, even leading to costly outages. As the scale and complexity of clouds increase, it becomes extraordinarily difficult to understand, detect and diagnose faults. During outages, engineers record the detailed information of the whole life cycle of faults (i.e., fault occurrence, fault detection, fault identification, and fault mitigation) in the form of postmortems. In this paper, we conduct a quantitative and qualitative study on 354 public post-mortems collected in three popular large-scale clouds, 97.7% of which spans from 2015 to 2021. By reviewing and analyzing post-mortems, we go through the life cycle of faults in clouds and obtain 10 major findings. Based on these findings, we further reach a series of actionable guidelines for better fault handling. Guangba Yu, Pengfei Chen 0002, Hongyang Chen 0002, Zhekang Chen |
ISSRE | 2 |
| 2022 | Graph based Incident Extraction and Diagnosis in Large-Scale Online SystemsabstractWith the ever increasing scale and complexity of online systems, incidents are gradually becoming commonplace. Without appropriate handling, they can seriously harm the system availability. However, in large-scale online systems, these incidents are usually drowning in a slew of issues (i.e., something abnormal, while not necessarily an incident), rendering them difficult to handle. Typically, these issues will result in a cascading effect across the system, and a proper management of the incidents depends heavily on a thorough analysis of this effect. Therefore, in this paper, we propose a method to automatically analyze the cascading effect of availability issues in online systems and extract the corresponding graph based issue representations incorporating both of the issue symptoms and affected service attributes. With the extracted representations, we train and utilize a graph neural networks based model to perform incident detection. Then, for the detected incident, we leverage the PageRank algorithm with a flexible transition matrix design to locate its root cause. We evaluate our approach using real-world data collected from the WeChat online service system, the largest instant message system in China. The results confirm the effectiveness of our approach. Moreover, our approach is successfully deployed in the company and eases the burden of operators in the face of a flood of issues and related alert signals. Pengfei Chen 0002, Yu Luo 0019, Qiuyu Yan, Hongyang Chen 0002, Guangba Yu |
ASE | 6 |
| 2022 | Microscaler: Cost-Effective Scaling for Microservice Applications in the Cloud With an Online Learning ApproachabstractRecently, the microservice becomes a popular architecture to construct cloud native systems due to its agility. In cloud native systems, autoscaling is a key enabling technique to adapt to workload changes by acquiring or releasing the right amount of computing resources. However, it becomes a challenging problem in microservice applications, since such an application usually comprises a large number of different microservices with complex interactions. When the performance decreases due to an unpredictable workload peak, it is difficult to pinpoint the scaling-needed services which need to scale out and evaluate how many resources they need. In this article, we present a novel system namedMicroscalerto automatically identify the scaling-needed services and scale them to meet the Service Level Agreement (SLA) with an optimal cost for microservice applications.Microscalerfirst collects the quality of service (QoS) metrics in the service mesh enabled microservice infrastructure. Then, it determines under-provisioning or over-provisioning service instances along the service dependency graph with a novel scaling-needed service criterion namedservice power. The service dependency graph could be obtained by correlating each request flow in the service mesh. By combining an online learning approach and a step-by-step heuristic approach,Microscalercan precisely reach the optimal service scale meeting the SLA requirements. The experimental evaluations in a microservice benchmark show thatMicroscalerachieves an average 93 percent precision in scaling-needed service determination and converges to the optimal service scale faster than several state-of-the-art methods. Moreover,Microscaleris lightweight and flexible enough to work in a large-scale microservice system. Guangba Yu, Pengfei Chen 0002, Zibin Zheng |
IEEE Trans. Cloud Comput. | 1 |
| 2021 | T-Rank: A Lightweight Spectrum based Fault Localization Approach for Microservice SystemsabstractThe cloud-native system is shifting from traditional monolithic architecture to microservice architecture because of loosely coupling, better maintainability and availability, faster deployment, and richer ecology brought by it. Except for these advantages, it still has an inevitable weakness-the communication over RPC (Remote Procedure Call) between services makes the system performance more unpredictable. Moreover, the complex interactions amongst services make it hard to reveal the root cause of performance issues. To address this challenge, we propose a lightweight spectrum-based performance diagnosis tool, named T-Rank. T-Rank provides the ranked suspicious score in a list of microservices to localize root causes with very few resources. We demonstrate the high accuracy and the low cost of T-Rank by conducting experiments with the data collected from a real-world production microservice system. Moreover, comparison results show that T-Rank outperforms other state-of-the-art approaches. Pengfei Chen 0002, Guangba Yu |
CCGRID | 3 |
| 2021 | Sieve: Attention-based Sampling of End-to-End Trace Data in Distributed Microservice SystemsabstractEnd-to-end tracing plays an important role in understanding and monitoring distributed microservice systems. The trace data are valuable to help find out the anomalous or erroneous behavior of the system. However, the volume of trace data is huge leading to a heavy burden on analyzing and storing them. To reduce the volume of trace data, the sampling technique is widely adopted. However, existing uniform sampling approaches are unable to capture uncommon traces that are more interesting and informative. To tackle this problem, we design and implement Sieve, an online sampler that aims to bias sampling towards uncommon traces by taking advantage of the attention mechanism. The evaluation results on the trace datasets collected from real-world and experimental microservice systems show that Sieve is effective to increase sampling probabilities of the structurally and temporally uncommon traces and reduce the storage space to a large extent by taking a low sampling rate. Pengfei Chen 0002, Guangba Yu, Hongyang Chen 0002, Zibin Zheng |
ICWS | 3 |
| 2021 | MicroRank: End-to-End Latency Issue Localization with Extended Spectrum Analysis in Microservice EnvironmentsabstractWith the advantages of flexible scalability and fast delivery, microservice has become a popular software architecture in the modern IT industry. However, the explosion in the number of service instances and complex dependencies make the troubleshooting extremely challenging in microservice environments. To help understand and troubleshoot a microservice system, the end-to-end tracing technology has been widely applied to capture the execution path of each request. Nevertheless, the tracing data are not fully leveraged by cloud and application providers when conducting latency issue localization in the microservice environment. Guangba Yu, Pengfei Chen 0002, Hongyang Chen 0002, Zijie Guan, Linxiao Jing, Tianjun Weng, Xinmeng Sun |
WWW | 1 |
| 2020 | A Learning-based Dynamic Load Balancing Approach for Microservice Systems in Multi-cloud EnvironmentabstractMulti-cloud environment has become common since companies manage to prevent cloud vendor lock-in for security and cost concerns. Meanwhile, the microservice architecture is often considered for its flexibility. Combining multi-cloud with microservice, the problem of routing requests among all possible microservice instances in multi-cloud environment arises. This paper presents a learning-based approach to route requests in order to balance the load. In our approach, the performance of microservice is modeled explicitly through machine learning models. The model can derive the response time from request volume, route decision, and other cloud metrics. Then the balanced route decision is obtained from optimizing the model with Bayesian Optimization. With this approach, the request route decision can adjust to dynamic runtime metrics instead of remaining static for all different circumstances. Explicit performance modeling avoids searching on an actual microservice system which is time-consuming. Experiments show that our approach reduces average response time by 10% at least. Jieqi Cui, Pengfei Chen 0002, Guangba Yu |
ICPADS | 3 |
| 2020 | SwissLog: Robust and Unified Deep Learning Based Log Anomaly Detection for Diverse FaultsabstractLog-based anomaly detection has been widely studied and achieves a satisfying performance on stable log data. But, the existing approaches still fall short meeting these challenges: 1) Log formats are changing continually in practice in those software systems under active development and maintenance. 2) Performance issues are latent causes that may not be detected by trivial monitoring tools. We thus propose SwissLog, namely a robust and unified deep learning based anomaly detection model for detecting diverse faults. SwissLog targets at those faults resulting in log sequence order changes and log time interval changes. To achieve that, an advanced log parser is introduced. Moreover, the semantic embedding and the time embedding approaches are combined to train a unified attention based BiLSTM model to detect anomalies. The experiments on real-world datasets and synthetic datasets show that SwissLog is robust to the changing log data and effective for diverse faults. Pengfei Chen 0002, Linxiao Jing, Guangba Yu |
ISSRE | 5 |
| 2019 | Microscaler: Automatic Scaling for Microservices with an Online Learning ApproachabstractRecently, the microservice becomes a popular architecture to construct cloud native systems due to its agility. In cloud native systems, autoscaling is a core enabling technique to adapt to workload changes by scaling out/in. However, it becomes a challenging problem in a microservice system, since such a system usually comprises a large number of different micro services with complex interactions. When bursty and unpredictable workloads arrive, it is difficult to pinpoint the scaling-needed services which need to scale and evaluate how much resource they need. In this paper, we present a novel system named Microscaler to automatically identify the scaling-needed services and scale them to meet the service level agreement (SLA) with an optimal cost for micro-service systems. Microscaler collects the quality of service metrics (QoS) with the help of the service mesh enabled infrastructure. Then, it determines the under-provisioning or over-provisioning services with a novel criterion named service power. By combining an online learning approach and a step-by-step heuristic approach, Microscaler could achieve the optimal service scale satisfying the SLA requirements. The experimental evaluations in a micro-service benchmark show that Microscaler converges to the optimal service scale faster than several state-of-the-art methods. Guangba Yu, Pengfei Chen 0002, Zibin Zheng |
ICWS | 1 |