EDBT 2026 Demo / reviewers in the wild / expert
Yifan Wu 0002
dblp:25/7019-2
· DBLP profile ↗
23ranked-venue papers
3as first author
22since 2021 · last 2026
0000-0001-5847-3132ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 19 · 3 first-author · 18 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UDA-RCL: Unsupervised Domain Adaptation for Microservice Root Cause Localization Utilizing Multimodal DataabstractRoot cause Localization methods play a crucial role in ensuring the stability of large-scale microservice systems. However, existing methods either rely on unsupervised approaches with limited localization accuracy or on supervised learning that requires large volumes of historical anomaly data. Such data is often unavailable in newly deployed systems. To address this limitation, we attempt to use the labeled data from mature systems to help new systems build root cause localization models. Specifically, we propose UDA-RCL, an unsupervised domain adaptation root cause localization method using multimodal data (log, metric and trace data). UDA-RCL first incorporates an aggregation based event extraction module to standardize the format of multimodal data from different systems. Then, it utilizes a multimodal event encoder and multimodal domain adversarial adaptation module to narrow the feature distribution gap between different systems. Furthermore, taking into account the situation of sparse anomaly samples, existing methods' classifiers are either hard to transfer to other systems or struggle to capture the process of anomaly propagation, we propose a PageRank classifier module. This module employs a neural network embedded with anomaly propagation rules to output the final root cause ranking results, alleviating the issue of sparse anomaly samples. Extensive experiments have proven that our method achieves the best results in both supervised and transfer learning scenarios. Xiaosong Huang, Yifan Wu 0002, Lingzhe Zhang, Ying Li 0012, Zhonghai Wu |
IEEE Trans. Serv. Comput. | 3 |
| 2025 | ScalaLog: Scalable Log-Based Failure Diagnosis Using LLMabstractAs Industrial Internet of Things (IIoT) software systems become increasingly complex, precise failure diagnosis has become both essential and challenging. Current log-based failure diagnosis methods lack scalability for different failure types. In IIoT software systems, the number of failure types is constantly growing, and retraining the model each time a new failure type is introduced is highly resource-intensive. Additionally, traditional log-based failure diagnosis models often require log parsing as a preliminary step, which can also be resource-consuming. To address these challenges, we propose a scalable log-based failure diagnosis method named ScalaLog. ScalaLog builds on RAG by utilizing LLM-based summarization to extract key log information, applying sample augmentation to increase the number of samples, and using CoT prompts to guide the LLM in failure diagnosis. Experiments on various public and real-world datasets demonstrate that ScalaLog significantly enhances failure diagnosis accuracy without the need for training or log parsing. Lingzhe Zhang, Mengxi Jia, Yifan Wu 0002, Ying Li 0012 |
ICASSP | 4 |
| 2025 | An Empirical Study on Commit Message Generation Using LLMs via In-Context LearningabstractCommit messages concisely describe code changes in natural language and are important for software maintenance. Several approaches have been proposed to automatically generate commit messages, but they still suffer from critical limitations, such as time-consuming training and poor generalization ability. To tackle these limitations, we propose to borrow the weapon of large language models (LLMs) and in-context learning (ICL). Our intuition is based on the fact that the training corpora of LLMs contain extensive code changes and their pairwise commit messages, which makes LLMs capture the knowledge about commits, while ICL can exploit the knowledge hidden in the LLMs and enable them to perform downstream tasks without model tuning. However, it remains unclear how well LLMs perform on commit message generation via ICL. In this paper, we conduct an empirical study to investigate the capability of LLMs to generate commit messages via ICL. Specifically, we first explore the impact of different settings on the performance of ICL-based commit message generation. We then compare ICL-based commit message generation with state-of-the-art approaches on a popular multilingual dataset and a new dataset we created to mitigate potential data leakage. The results show that ICL-based commit message generation significantly outperforms state-of-the-art approaches on subjective evaluation and achieves better generalization ability. We further analyze the root causes for LLM's underperformance and propose several implications, which shed light on future research directions for using LLMs to generate commit messages. Yifan Wu 0002, Ying Li 0012, Siyu Yu, Wei Jiang 0041 |
ICSE | 1 |
| 2025 | CSLParser: A Collaborative Framework Using Small and Large Language Models for Log ParsingabstractLog parsing is a prerequisite for log analysis. Recently, large language models (LLMs) have demonstrated high accuracy in log parsing. However, their frequent invocations incur substantial costs. To address this issue, some methods have turned to small language models (SLMs), which offer improved efficiency but suffer from reduced accuracy due to limited model capacity. To achieve both high accuracy and efficiency, we propose CSLParser, a collaborative log parsing framework using SLMs and LLMs. CSLParser delegates most log parsing tasks to SLMs and selectively invokes LLMs to correct parsing results generated by SLMs, thereby effectively reducing the invocation cost of LLMs while maintaining high accuracy. Specifically, to enhance the accuracy of SLMs, we propose a diversified sampling strategy to select diverse samples for training, enabling SLMs to effectively handle diverse log patterns. To efficiently invoke LLMs, we design a rule-based selection strategy to identify hard cases that are challenging for SLMs to correctly parse, which are subsequently corrected by LLMs. Additionally, we propose a dynamic template updating mechanism that merges similar templates based on structural and semantic information to further enhance parsing accuracy. Extensive experiments on public large-scale log datasets show that CSLParser outperforms state-of-the-art baselines in both accuracy and efficiency. Weijie Hong, Yifan Wu 0002, Lingzhe Zhang, Chiming Duan, Pei Xiao 0005, Minghua He, Xixuan Yang, Ying Li 0012 |
ISSRE | 2 |
| 2025 | Log Parsing Using LLMs with Self-Generated In-Context Learning and Self-CorrectionabstractLog parsing transforms log messages into structured formats, serving as a crucial step for log analysis. Despite a variety of log parsers that have been proposed, their performance on evolving log data remains unsatisfactory due to reliance on human-crafted rules or learning-based models with limited training data. The recent emergence of large language models (LLMs) has demonstrated strong abilities in understanding natural language and code, making it promising to apply LLMs for log parsing. Consequently, several studies have proposed LLM-based log parsers. However, LLMs may produce inaccurate templates, and existing LLM-based log parsers directly use the template generated by the LLM as the parsing result, hindering the accuracy of log parsing. Furthermore, these log parsers depend heavily on historical log data as demonstrations, which poses challenges in maintaining accuracy when dealing with scarce historical log data or evolving log data. To address these challenges, we propose AdaParser, an effective and adaptive log parsing framework using LLMs with self-generated in-context learning (SG-ICL) and self-correction. To facilitate accurate log parsing, AdaParser incorporates a novel component, a template corrector, which utilizes the LLM to correct potential parsing errors in the templates it generates. In addition, AdaParser maintains a dynamic candidate set composed of previously generated templates as demonstrations to adapt evolving log data. Extensive experiments on public large-scale datasets indicate that AdaParser outperforms state-of-the-art methods across all metrics, even in zero-shot scenarios. Moreover, when integrated with different LLMs, AdaParser consistently enhances the performance of the utilized LLMs by a large margin. Yifan Wu 0002, Siyu Yu, Ying Li 0012 |
ICPC | 1 |
| 2025 | LogAction: Consistent Cross-system Anomaly Detection through Logs via Active Domain AdaptationabstractLog-based anomaly detection is a essential task for ensuring the reliability and performance of software systems. However, the performance of existing anomaly detection methods heavily relies on labeling, while labeling a large volume of logs is highly challenging. To address this issue, many approaches based on transfer learning and active learning have been proposed. Nevertheless, their effectiveness is hindered by issues such as the gap between source and target system data distributions and cold-start problems. In this paper, we propose LogAction, a novel log-based anomaly detection model based on active domain adaptation. LogAction integrates transfer learning and active learning techniques. On one hand, it uses labeled data from a mature system to train a base model, mitigating the cold-start issue in active learning. On the other hand, LogAction utilize free energy-based sampling and uncertainty-based sampling to select logs located at the distribution boundaries for manual labeling, thus addresses the data distribution gap in transfer learning with minimal human labeling efforts. Experimental results on six different combinations of datasets demonstrate that LogAction achieves an average 93.01% F1 score with only 2% of manual labels, outperforming some state-of-the-art methods by 26.28%. Website: https://logaction.github.io Chiming Duan, Minghua He, Pei Xiao 0005, Zhewei Zhong, Yan Niu, Lingzhe Zhang, Siyu Yu, Yifan Wu 0002, Weijie Hong, Ying Li 0012, Gang Huang 0001 |
ASE | 11 |
| 2025 | United We Stand: Towards End-to-End Log-based Fault Diagnosis via Interactive Multi-Task LearningabstractLog-based fault diagnosis is essential for maintaining software system availability. However, existing fault diagnosis methods are built using a task-independent manner, which fails to bridge the gap between anomaly detection and root cause localization in terms of data form and diagnostic objectives, resulting in three major issues: 1) Diagnostic bias accumulates in the system; 2) System deployment relies on expensive monitoring data; 3) The collaborative relationship between diagnostic tasks is overlooked. Facing this problems, we propose a novel end-to-end log-based fault diagnosis method, Chimera, whose key idea is to achieve end-to-end fault diagnosis through bidirectional interaction and knowledge transfer between anomaly detection and root cause localization. Chimera is based on interactive multitask learning, carefully designing interaction strategies between anomaly detection and root cause localization at the data, feature, and diagnostic result levels, thereby achieving both sub-tasks interactively within a unified end-to-end framework. Evaluation on two public datasets and one industrial dataset shows that Chimera outperforms existing methods in both anomaly detection and root cause localization, achieving improvements of over 2.92%~5.00% and 19.01% ~ 37.09%, respectively. It has been successfully deployed in production, serving an industrial cloud platform. Minghua He, Chiming Duan, Pei Xiao 0005, Siyu Yu, Lingzhe Zhang, Weijie Hong, Yifan Wu 0002, Ying Li 0012, Gang Huang 0001 |
ASE | 9 |
| 2025 | Walk the Talk: Is Your Log-based Software Reliability Maintenance System Really Reliable?abstractLog-based software reliability maintenance systems are crucial for sustaining stable customer experience. However, existing deep learning-based methods represent a black box for service providers, making it impossible for providers to understand how these methods detect anomalies, thereby hindering trust and deployment in real production environments. To address this issue, this paper defines a trustworthiness metric—diagnostic faithfulness—for models to gain service providers’ trust, based on surveys of SREs at a major cloud provider. We design two evaluation tasks: attention-based root cause localization and event perturbation. Empirical studies demonstrate that existing methods perform poorly in diagnostic faithfulness. Consequently, we propose FaithLog, a faithful log-based anomaly detection system, which achieves faithfulness through a carefully designed causality-guided attention mechanism and adversarial consistency learning. Evaluation results on two public datasets and one industrial dataset demonstrate that the proposed method achieves state-of-the-art performance in diagnostic faithfulness. Minghua He, Chiming Duan, Pei Xiao 0005, Lingzhe Zhang, Kangjin Wang, Yifan Wu 0002, Ying Li 0012, Gang Huang 0001 |
ASE | 7 |
| 2025 | CoorLog: Efficient-Generalizable Log Anomaly Detection via Adaptive Coordinator in Software EvolutionabstractFrequent software updates lead to log evolution, posing generalization challenges for current log anomaly detection. Traditional log anomaly detection research focuses on using small deep learning models (SMs), but these models inherently lack generalization due to their closed-world assumption. Large language models (LLMs) exhibit strong semantic understanding and generalization capabilities, making them promising for log anomaly detection. However, they suffer from computational inefficiencies. To balance efficiency and generalization, we propose a collaborative log anomaly detection scheme (CoorLog) that uses an adaptive coordinator to integrate SM and LLM. The coordinator determines if incoming logs have evolved. Non-evolved logs are routed to the SM, while evolved logs are directed to the LLM for detailed inference using the constructed Evol-CoT. To gradually adapt to evolution, we introduce the adaptive evolution mechanism (AEM), which updates the coordinator to redirect evolved logs identified by the LLM to the SM. Simultaneously, the SM is fine-tuned to inherit the LLM’s judgment on these logs. Extensive experiments on real-world datasets demonstrate that CoorLog achieves superior F1-scores in both intra-version and inter-version anomaly detection. Additionally, CoorLog reduces processing time by 91.63% and token consumption by 85.59% compared to using an LLM alone. Pei Xiao 0005, Chiming Duan, Minghua He, Yifan Wu 0002, Gege Gao, Lingzhe Zhang, Weijie Hong, Ying Li 0012, Gang Huang 0001 |
ASE | 5 |
| 2024 | CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window ExtendingabstractSelf-attention and position embedding are two crucial modules in transformer-based Large Language Models (LLMs).However, the potential relationship between them is far from well studied, especially for long context window extending.In fact, anomalous behaviors that hinder long context extrapolation exist between Rotary Position Embedding (RoPE) and vanilla self-attention.Incorrect initial angles between Q and K can cause misestimation in modeling rotary position embedding of the closest tokens.To address this issue, we propose Collinear Constrained Attention mechanism, namely CoCA.Specifically, we enforce a collinear constraint between Q and K to seamlessly integrate RoPE and self-attention.While only adding minimal computational and spatial complexity, this integration significantly enhances long context window extrapolation ability.We provide an optimized implementation, making it a drop-in replacement for any existing transformer-based models.Extensive experiments demonstrate that CoCA excels in extending context windows.A CoCAbased GPT model, trained with a context length of 512, can extend the context window up to 32K (60×) without any fine-tuning.Additionally, incorporating CoCA into LLaMA-7B achieves extrapolation up to 32K within a training length of only 2K.Our code is publicly available at: https://github.com/codefuse- ai/Collinear-Constrained-Attention Shiyi Zhu, Wei Jiang 0041, Siqiao Xue, Yifan Wu 0002 |
ACL (1) | 6 |
| 2024 | Unlocking the Power of Numbers: Log Compression via Numeric Token ParsingabstractParser-based log compressors have been widely explored in recent years because the explosive growth of log volumes makes the compression performance of general-purpose compressors unsatisfactory. These parser-based compressors preprocess logs by grouping the logs based on the parsing result and then feed the preprocessed files into a general-purpose compressor. However, parser-based compressors have their limitations. First, the goals of parsing and compression are misaligned, so the inherent characteristics of logs were not fully utilized. In addition, the performance of parser-based compressors depends on the sample logs and thus it is very unstable. Moreover, parser-based compressors often incur a long processing time. To address these limitations, we propose Denum, a simple, general log compressor with high compression ratio and speed. The core insight is that a majority of the tokens in logs are numeric tokens (i.e. pure numbers, tokens with only numbers and special characters, and numeric variables) and effective compression of them is critical for log compression. Specifically, Denum contains a Numeric Token Parsing module, which extracts all numeric tokens and applies tailored processing methods (e.g. store the differences of incremental numbers like timestamps), and a String Processing module, which processes the remaining log content without numbers. The processed files of the two modules are then fed as input to a general-purpose compressor and it outputs the final compression results. Denum has been evaluated on 16 log datasets and it achieves an 8.7% -- 434.7% higher average compression ratio and 2.6× -- 37.7× faster average compression speed (i.e. 26.2 MB/S) compared to the baselines. Moreover, integrating Denum's Numeric Token Parsing module into existing log compressors can provide a 11.8% improvement in their average compression ratio and achieve 37% faster average compression speed. Siyu Yu, Yifan Wu 0002, Ying Li 0012, Pinjia He |
ASE | 2 |
| 2024 | MMDL-Based Data Augmentation with Domain Knowledge for Time Series Classification
Xiaosheng Li, Yifan Wu 0002, Wei Jiang 0041, Ying Li 0012 |
ECML/PKDD (3) | 2 |
| 2024 | OCRCL: Online Contrastive Learning for Root Cause Localization of Business IncidentsabstractMicroservices architecture has garnered extensive attention for its stability and scalability. However, in the complex and dynamic landscape of microservices systems, a incident in one service can propagate to others, resulting in significant economic losses and degraded user experiences. Therefore, the effective and precise localization of incidents in microservices systems becomes a critical concern. Previous research has leveraged runtime data (logs, metrics, call traces) and historical incident data to assist in root cause localization. However, due to the scarcity of business incidents (those causing severe impacts on business operations) and the fact that many incidents are reported by users, relevant run-time data and sufficient historical data are often unavailable, rendering previous methods impractical. In response to this challenge, we propose an online contrastive learning-based method for root cause localization of business incidents(OCRCL). We fully exploit incident tickets and the static dependency graph of services, integrating both textual semantic information and structural information from the dependency graph to discover root causes. Furthermore, we suggest that online contrastive learning can exhibit excellent performance with limited data and enable real-time model updates, making it better suited for industrial scenarios. Our approach demonstrates significant improvements over baseline methods across three real-world industrial datasets, highlighting its effectiveness in root cause localization. Xiaosong Huang, Yifan Wu 0002, Yujin Zhao, Changlong Wu, Songlin Zhang, Ying Li 0012, Zhonghai Wu |
SANER | 3 |
| 2024 | Understanding and Improving Change Risk Detection in PracticeabstractChanges are inevitable and frequent in large-scale online service systems, which has been one of the leading causes that induce incidents. Change risk detection (CRD) aims to help engineers detect high-risk changes so that proactive actions can be taken to avoid incidents, which is vital for the availability and reliability of online service systems. Though some efforts have been dedicated to CRD, their performances are still far from satisfactory in practice. To better understand the practical challenges of CRD, we conducted the first empirical study on a large-scale online service system in Ant Group. Through this study, we identified four critical challenges, including poor interpretability, adaptation to diverse change types, indirect anomaly factors, and expected but false alarm anomalies. To address these challenges, we propose an effective and eXplainable Change Risk Detection framework named XCRD. XCRD can detect change-induced unexpected anomalies using multi-source data and provide explainable alerts for engineers to facilitate anomaly diagnosis and mitigation. We have successfully deployed XCRD in Ant Group for the past 14 months, demonstrating a significant performance improvement in CRD. We also discuss some successful cases and lessons learned during our study. To our knowledge, we are the first to deeply investigate CRD in industrial scenarios. We believe that our work can provide valuable insights for engineers and researchers to understand and improve CRD in practice. Yifan Wu 0002, Ying Li 0012, Bingxu Chai, Wei Jiang 0041 |
SANER | 1 |
| 2024 | XDrain: Effective log parsing in log streams using fixed-depth forest
Yang Tian 0008, Siyu Yu, Donghui Gao, Yifan Wu 0002, Suqun Huang, Xiaochun Hu, Ningjiang Chen |
Inf. Softw. Technol. | 5 |
| 2023 | How to Manage Change-Induced Incidents? Lessons from the Study of Incident Life CycleabstractIn online service systems, software changes cause a majority of incidents (i.e., unplanned interruptions and outages). Managing change-induced incidents efficiently is crucial for ensuring the reliability and availability of online service systems. Understanding the incidents can help improve change-induced incident management. The task is challenging because the life cycle of change-induced incidents is complicated due to diverse change deployment and incident resolution procedures. Detailed records of the incidents and changes, together with a comprehensive analysis, are needed to gain an in-depth understanding. In this paper, we conduct a qualitative and quantitative study on 231 change-induced incidents in a real-world, large-scale online service system. Detailed change tickets and incident timeline in the post-mortems provides extensive information about the incident life cycle, enabling us to understand each incident in depth. Based on the data, we give a generic model of the complicated life cycle of change-induced incidents. Following the model, we systematically study the whole life cycle of the incident, including the introduction and resolution stages, and answer what affects the efficiency of resolution. We obtain 9 major findings from our study. Based on the findings, we discuss existing techniques and promising future directions for improving change-induced incident management. Yujin Zhao, Ye Tao 0011, Songlin Zhang, Changlong Wu, Yifan Wu 0002, Ying Li 0012, Zhonghai Wu |
ISSRE | 6 |
| 2023 | Log Parsing with Generalization Ability under New Log TypesabstractLog parsing, which converts semi-structured logs into structured logs, is the first step for automated log analysis. Existing parsers are still unsatisfactory in real-world systems due to new log types in new-coming logs. In practice, available logs collected during system runtime often do not contain all the possible log types of a system because log types related to infrequently activated system states are unlikely to be recorded and new log types are frequently introduced with system updates. Meanwhile, most existing parsers require preprocessing to extract variables in advance, but preprocessing is based on the operator’s prior knowledge of available logs and therefore may not work well on new log types. In addition, parser parameters set based on available logs are difficult to generalize to new log types. To support new log types, we propose a variable generation imitation strategy to craft a novel log parsing approach with generalization ability, called Log3T. Log3T employs a pre-trained transformer encoder-based model to extract log templates and can update parameters at parsing time to adapt to new log types by a modified test-time training. Experimental results on 16 benchmark datasets show that Log3T outperforms the state-of-the-art parsers in terms of parsing accuracy. In addition, Log3T can automatically adapt to new log types in new-coming logs. Siyu Yu, Yifan Wu 0002, Zhijing Li 0007, Pinjia He, Ningjiang Chen |
ESEC/SIGSOFT FSE | 2 |
| 2023 | UDA-DP: Unsupervised Domain Adaptation for Software Defect PredictionabstractSoftware defect prediction can automatically locate defective code modules to focus testing resources better. Traditional defect prediction methods mainly focus on manually designing features, which are input into machine learning classifiers to identify defective code. However, there are mainly two problems in prior works. First manually designing features is time consuming and unable to capture the semantic information of programs, which is an important capability for accurate defect prediction. Second the labeled data is limited along with severe class imbalance, affecting the performance of defect prediction.In response to the above problems, we first propose a new unsupervised domain adaptation method using pseudo labels for defect prediction(UDA-DP). Compared to manually designed features, it can automatically extract defective features from source programs to save time and contain more semantic information of programs. Moreover, unsupervised domain adaptation using pseudo labels is a kind of transfer learning, which is effective in leveraging rich information of limited data, alleviating the problem of insufficient data.Experiments with 10 open source projects from the PROMISE data set show that our proposed UDA-DP method outperforms the state-of-the-art methods for both within-project and cross-project defect predictions. Our code and data are available at https://github.com/xsarvin/UDA-DP. Xiaosong Huang, Yifan Wu 0002, Ying Li 0012, Hao Yu 0016, Dadi Guo, Zhonghai Wu |
SANER | 2 |
| 2023 | Self-supervised log parsing using semantic contribution difference
Siyu Yu, Ningjiang Chen, Yifan Wu 0002, Wensheng Dou |
J. Syst. Softw. | 3 |
| 2023 | Brain: Log Parsing With Bidirectional Parallel TreeabstractAutomated log analysis can facilitate failure diagnosis for developers and operators using a large volume of logs. Log parsing is a prerequisite step for automated log analysis, which parses semi-structured logs into structured logs. However, existing parsers are difficult to apply to software-intensive systems, due to their unstable parsing accuracy on various software. Although neural network-based approaches are stable, their inefficiency makes it challenging to keep up with the speed of log production. In this work, we found that the longest common pattern among logs is likely to be part of the log template. Inspired by this key insight, we propose a new stable log parsing approach, called Brain, which creates initial groups according to the longest common pattern. Then a bidirectional tree is used to hierarchically complement the constant words to the longest common pattern to form the complete log template efficiently. Experimental results on 16 benchmark datasets show that our approach outperforms the state-of-the-art parsers on two widely-used parsing accuracy metrics, and it only takes around 46 seconds to process one million lines of logs. Siyu Yu, Pinjia He, Ningjiang Chen, Yifan Wu 0002 |
IEEE Trans. Serv. Comput. | 4 |
| 2022 | ByteGNN: Efficient Graph Neural Network Training at Large ScaleabstractGraph neural networks (GNNs) have shown excellent performance in a wide range of applications such as recommendation, risk control, and drug discovery. With the increase in the volume of graph data, distributed GNN systems become essential to support efficient GNN training. However, existing distributed GNN training systems suffer from various performance issues including high network communication cost, low CPU utilization, and poor end-to-end performance. In this paper, we propose ByteGNN, which addresses the limitations in existing distributed GNN systems with three key designs: (1) an abstraction of mini-batch graph sampling to support high parallelism, (2) a two-level scheduling strategy to improve resource utilization and to reduce the end-to-end GNN training time, and (3) a graph partitioning algorithm tailored for GNN workloads. Our experiments show that ByteGNN outperforms the state-of-the-art distributed GNN systems with up to 3.5--23.8 times faster end-to-end execution, 2--6 times higher CPU utilization, and around half of the network communication cost. Chenguang Zheng, Yuxuan Cheng, Zhezheng Song, Yifan Wu 0002, Changji Li, James Cheng |
Proc. VLDB Endow. | 5 |
| 2021 | LogFlash: Real-time Streaming Anomaly Detection and Diagnosis from System Logs for Large-scale Software SystemsabstractToday, software systems are getting increasingly large and complex and a short failure time may cause huge loss. Therefore, it is important to detect and diagnose anomalies accurately and timely. System logs are a straightforward and important source of information for anomaly detection and diagnosis. However, existing log-based approaches have three key limitations. First, they are not designed for processing real-time log streams. Second, they require restrictions on training log data. Third, they lack the adaptiveness to system update. To break through these limitations, we propose LogFlash, a real-time streaming anomaly detection and diagnosis approach that enables both training and detection in a real-time streaming processing manner. By assigning a dynamic pairwise transition rate to each template pair and model the transition possibility as typical power-law distribution, our approach achieves real-time model construction and updates. Experiment results show that it reduces over 5 times of training and detection time compared with the state-of-art works while maintaining the capability of accurate anomaly diagnosis. Yifan Wu 0002, Chuanjia Hou, Ying Li 0012 |
ISSRE | 2 |
| 2020 | How Far Have We Come in Detecting Anomalies in Distributed Systems? An Empirical Study with a Statement-level Fault Injection MethodabstractAnomaly detection in distributed systems has been a fertile research area, and a range of anomaly detectors have been proposed for distributed systems. Unfortunately, there is no systematic quantitative study of the efficacy of different anomaly detectors, which is of great importance to reveal the deficiencies of existing anomaly detectors and shed light on future research directions. In this paper, we investigate how various anomaly detectors behave on anomalies of different types and the reasons for the same, by extensively injecting software faults into three widely-used distributed systems. We use a statementlevel fault injection method to observe the anomalies, characterize these anomalies, and analyze the detection results from anomaly detectors of three categories. We find that: (1) the distributed systems' own error reporting mechanisms are able to report most of the anomalies (from 82.1% to 92.8%) but they incur a high false alarm rate of 26.6%. (2) State-of-the-art anomaly detectors are able to detect the existence of anomalies with 99.08% precision and 90.60% recall, but there is still a long way to go to pinpoint the accurate location of the detected anomalies, and (3) Log-based anomaly detection techniques outperform other anomaly detection techniques, but not for all anomaly types. Yong Yang 0011, Yifan Wu 0002, Karthik Pattabiraman, Long Wang 0003, Ying Li 0012 |
ISSRE | 2 |