EDBT 2026 Demo / reviewers in the wild / expert
Yong Yang 0011
dblp:11/357-11
· DBLP profile ↗
14ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0001-9667-2423ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EagerLog: Active Learning Enhanced Retrieval Augmented Generation for Log-based Anomaly DetectionabstractLogs record essential information about system operations and serve as a critical source for anomaly detection, which has generated growing research interest. Utilizing large language models (LLMs) within a retrieval-augmented generation (RAG) framework for log-based anomaly detection is an effective approach due to its strong generalization capabilities and efficient few-shot performance. However, the effectiveness of this method hinges on the quality of the knowledge source, which can be impacted by noise and changes within the software systems. Facing these problems, in this paper, we propose a novel log-based anomaly detection method named EagerLog, employing active learning to choose the logs for humans to label, thereby adding them to the knowledge source, thus enhancing the knowledge source and maintaining its quality. Our experiments on three open datasets (BGL, Thunderbird, Zookeeper) and one industrial dataset demonstrate that EagerLog can achieve 93.65% F1 score with approximately 10 labeled log sequences, surpassing existing methods by 15.32%. Chiming Duan, Yong Yang 0011, Guiyang Liu, Jinbu Liu, Huxing Zhang, Qi Zhou 0001, Ying Li 0012, Gang Huang 0001 |
ICASSP | 3 |
| 2025 | Famos: Fault Diagnosis for Microservice Systems Through Effective Multi-Modal Data FusionabstractAccurately diagnosing the fault that causes the failure is crucial for maintaining the reliability of a microservice system after a failure occurs. Mainstream fault diagnosis approaches are data-driven and mainly rely on three modalities of runtime data: traces, logs, and metrics. Diagnosing faults with multiple modalities of data in microservice systems has been a clear trend in recent years because different types of faults and corresponding failures tend to manifest in data of various modalities. Accurately diagnosing faults by fully leveraging multiple modalities of data is confronted with two challenges: 1) how to minimize information loss when extracting features for data of each modality; 2) how to correctly capture and utilize the relationships among data of different modalities. To address these challenges, we propose FAMOS, a Fault diagnosis Approach for MicrOservice Systems through effective multi-modal data fusion. On the one hand, FAMOS employs independent feature extractors to preserve the intrinsic features for each modality. On the other hand, FAMOS introduces a new Gaussian-attention mechanism to accurately correlate data of different modalities and then captures the inter-modality relationship with a crossattention mechanism. We evaluated FAMOS on two datasets constructed by injecting comprehensive and abundant faults into an open-source microservice system and a real-world industrial microservice system. Experimental results demonstrate the FAMOS's effectiveness in fault diagnosis, achieving significant improvements in F1 scores compared to state-of-the-art (SOTA) methods, with an increase of 20.33 %. Chiming Duan, Yong Yang 0011, Guiyang Liu, Jinbu Liu, Huxing Zhang, Qi Zhou 0001, Ying Li 0012, Gang Huang 0001 |
ICSE | 2 |
| 2025 | Tracing Service Request Processing in CloudabstractABSTRACT Nowadays, more and more IT services are being hosted on cloud systems, which render cloud systems to grow into a huge complex with millions of physical servers, multi‐layer software stacks and the processing of cloud service requests across many servers and software layers. It is highly demanded for cloud service providers to have the capability of getting the knowledge on cloud service behaviour directly from the service execution instead of from people's expertise. This paper studies the problem of tracing cloud service's processing of requests across components in cloud environments and proposes cloud tracing mechanisms for this purpose. We also developed model‐based studies of our proposed mechanisms for analysing certain designs of the mechanisms. The implementation of the proposed cloud tracing is deployed onto the environments of OpenStack, Kubernetes and Hadoop, and the experiments on these environments demonstrate that our mechanisms effectively trace cloud service behaviour and generate a single complete request execution path, while without our mechanisms the cloud tracing either fails to work or results in thousands of path segments. Our mechanisms have a low performance overhead (2.3%) in the experiments. Yinqin Zhao, Long Wang 0003, Xuanqing Shi, Yong Yang 0011, Ying Li 0012, Zhengang Wang, Dongdong Shangguan |
Softw. Test. Verification Reliab. | 6 |
| 2025 | Towards Close-to-Zero Runtime Collection Overhead: Raft-Based Anomaly Diagnosis on System Faults for Distributed Storage SystemabstractDistributed storage systems are fundamental infrastructures of today’s large-scale software systems such as cloud systems. Diagnosing anomalies in distributed storage systems is essential for maintaining software availability. Existing anomaly diagnosis approaches mainly rely on the run-time data including monitoring data and application logs. However, collecting and analyzing the run-time data requires huge computing, storage, and management costs. Typically, more fine-grained run-time data can reveal more symptoms of anomalies, but on the contrary, requires more computing, storage, and management costs. As a result, solving the anomaly diagnosis problem is a balancing between the quality of run-time data and system overhead or cost. In this paper, we take into account both data quality and system overhead or cost by introducing a new type of run-time data-Raft logs. Raft logs are naturally produced by distributed storage systems and collecting raft logs will not bring any extra system overhead. To verify the ability of Raft logs in reflecting anomalies, we conduct a comprehensive study on the interconnection between the anomalies and Raft logs. Based on the study, we propose an effectiveRaft-BasedAnomalyDiagnosis approach namedRBAD. For evaluation, we expose the first open-sourced comprehensive dataset with multiple runtime data containing both Raft logs, application logs and monitoring data. Experiments based on this dataset demonstrate RBAD’s superiority, outperforming monitoring-based methods by 15.38% and log-based methods by 53.10%. Lingzhe Zhang, Mengxi Jia, Yong Yang 0011, Zhonghai Wu, Ying Li 0012 |
IEEE Trans. Serv. Comput. | 5 |
| 2024 | Reducing Events to Augment Log-based Anomaly Detection Models: An Empirical StudyabstractAs software systems grow increasingly intricate, the precise detection of anomalies have become both essential and challenging. Current log-based anomaly detection methods depend heavily on vast amounts of log data leading to inefficient inference and potential misguidance by noise logs. However, the quantitative effects of log reduction on the effectiveness of anomaly detection remain unexplored. Therefore, we first conduct a comprehensive study on six distinct models spanning three datasets. Through the study, the impact of log quantity and their effectiveness in representing anomalies is qualifies, uncovering three distinctive log event types that differently influence model performance. Drawing from these insights, we propose LogCleaner: an efficient methodology for the automatic reduction of log events in the context of anomaly detection. Serving as middleware between software systems and models, LogCleaner continuously updates and filters anti-events and duplicative-events in the raw generated logs. Experimental outcomes highlight LogCleaner’s capability to reduce over 70% of log events in anomaly detection, accelerating the model’s inference speed by approximately 300%, and universally improving the performance of models for anomaly detection. Lingzhe Zhang, Kangjin Wang, Mengxi Jia, Yong Yang 0011, Ying Li 0012 |
ESEM | 5 |
| 2024 | Multivariate Log-based Anomaly Detection for Distributed DatabaseabstractDistributed databases are fundamental infrastructures of today's large-scale software systems such as cloud systems. Detecting anomalies in distributed databases is essential for maintaining software availability. Existing approaches, predominantly developed using Loghub-a comprehensive collection of log datasets from various systems-lack datasets specifically tailored to distributed databases, which exhibit unique anomalies. Additionally, there's a notable absence of datasets encompassing multi-anomaly, multi-node logs. Consequently, models built upon these datasets, primarily designed for standalone systems, are inadequate for distributed databases, and the prevalent method of deeming an entire cluster anomalous based on irregularities in a single node leads to a high false-positive rate. This paper addresses the unique anomalies and multivariate nature of logs in distributed databases. We expose the first open-sourced, comprehensive dataset with multivariate logs from distributed databases. Utilizing this dataset, we conduct an extensive study to identify multiple database anomalies and to assess the effectiveness of state-of-the-art anomaly detection using multivariate log data. Our findings reveal that relying solely on logs from a single node is insufficient for accurate anomaly detection on distributed database. Leveraging these insights, we propose MultiLog, an innovative multivariate log-based anomaly detection approach tailored for distributed databases. Our experiments, based on this novel dataset, demonstrate MultiLog's superiority, outperforming existing state-of-the-art methods by approximately 12%. Lingzhe Zhang, Mengxi Jia, Ying Li 0012, Yong Yang 0011, Zhonghai Wu |
KDD | 5 |
| 2024 | Hilogx: noise-aware log-based anomaly detection with human feedback
Ying Li 0012, Yong Yang 0011, Gang Huang 0001 |
VLDB J. | 3 |
| 2023 | Capturing Request Execution Path for Understanding Service Behavior and Detecting Anomalies Without Code InstrumentationabstractWith the increasing scale and complexity of cloud platforms and big-data analytics platforms, it is becoming more and more challenging to understand and diagnose the processing of a service request across multi-layer software stacks of such platforms. One way that helps to deal with this problem is to accurately capture the complete end-to-end execution path of service requests among all involved components. This paper presents REPTrace, a generic methodology for capturing such execution paths in a transparent fashion. Moreover, this paper demonstrates the effectiveness of REPTrace by presenting how REPTrace can be leveraged for knowledge extraction and anomaly detection on the platforms’ request processing. Our experimental results show that, REPTrace enables capturing a holistic view of the request processing across multiple layers of the platforms (which is missing in official documentation) and discovering important undocumented features of the platforms. Fault injection experiments show execution anomalies are detected with 93% precision and 96% recall with aid of REPTrace. Yong Yang 0011, Long Wang 0003, Ying Li 0012 |
IEEE Trans. Serv. Comput. | 1 |
| 2022 | Augmenting Log-based Anomaly Detection Models to Reduce False Anomalies with Human FeedbackabstractWith the increasing complexity of modern software systems, it is essential yet hard to detect anomalies and diagnose problems precisely. Existing log-based anomaly detection approaches rely on a few key assumptions on system logs and perform well in some experimental systems. However, real-world industrial systems are often with poor logging quality, in which system logs are noisy and often violate the assumptions of existing approaches. This makes these approaches inefficient. This paper first conducts a comprehensive study on the system logs of three large-scale industrial software systems. Through the study, we identify four typical anti-patterns that affect the detection results the most. Based on these patterns, we propose HiLog, an effective human-in-the-loop log-based anomaly detection approach that integrates human knowledge to augment anomaly detection models. With little human labeling effort, our approach can significantly improve the effectiveness of existing models. Experiment results on three large-scale industrial software systems show that our method improves over 50% precision rate on average. Ying Li 0012, Yong Yang 0011, Gang Huang 0001, Zhonghai Wu |
KDD | 3 |
| 2022 | Tracing Processing of Service Requests in Cloud EnvironmentsabstractCloud computing is growingly popular for hosting IT services, and is also growing into a huge complex with millions of physical servers, multi-layer software stacks and the processing of cloud service requests across many servers and software layers. It is highly demanded for cloud service providers to have the capability of getting the knowledge on cloud service behavior directly from the service execution instead of from people's expertise. This paper studies the problem of tracing cloud services' processing of requests across components in cloud environments, and proposes cloud tracing mechanisms for this purpose. The implementation of the proposed cloud tracing is deployed onto an OpenStack cloud environment, and the experiments performed on the cloud environment shows that our mechanisms effectively trace cloud service behavior and generate a single complete request execution path, while without our mechanisms the cloud tracing either could not work or results in thousands of path segments. Our mechanisms' performance overhead is low (2.3%). Yinqin Zhao, Long Wang 0003, Xuanqing Shi, Yong Yang 0011, Ying Li 0012, Zhengang Wang, Dongdong Shangguan |
PRDC | 6 |
| 2020 | How Far Have We Come in Detecting Anomalies in Distributed Systems? An Empirical Study with a Statement-level Fault Injection MethodabstractAnomaly detection in distributed systems has been a fertile research area, and a range of anomaly detectors have been proposed for distributed systems. Unfortunately, there is no systematic quantitative study of the efficacy of different anomaly detectors, which is of great importance to reveal the deficiencies of existing anomaly detectors and shed light on future research directions. In this paper, we investigate how various anomaly detectors behave on anomalies of different types and the reasons for the same, by extensively injecting software faults into three widely-used distributed systems. We use a statementlevel fault injection method to observe the anomalies, characterize these anomalies, and analyze the detection results from anomaly detectors of three categories. We find that: (1) the distributed systems' own error reporting mechanisms are able to report most of the anomalies (from 82.1% to 92.8%) but they incur a high false alarm rate of 26.6%. (2) State-of-the-art anomaly detectors are able to detect the existence of anomalies with 99.08% precision and 90.60% recall, but there is still a long way to go to pinpoint the accurate location of the detected anomalies, and (3) Log-based anomaly detection techniques outperform other anomaly detection techniques, but not for all anomaly types. Yong Yang 0011, Yifan Wu 0002, Karthik Pattabiraman, Long Wang 0003, Ying Li 0012 |
ISSRE | 1 |
| 2018 | Transparently Capturing Execution Path of Service/Job Request Processing
Yong Yang 0011, Long Wang 0003, Ying Li 0012 |
ICSOC | 1 |
| 2016 | An Approach for Cross-Community Content Recommendation: A Case Study on Docker
Yong Yang 0011, Ying Li 0012, Hongyan Tang, Wenlong Shao |
APWeb (2) | 1 |
| 2016 | CUT: A Combined Approach for Tag Recommendation in Software Information Sites
Yong Yang 0011, Ying Li 0012, Zhonghai Wu, Wenlong Shao |
KSEM | 1 |