EDBT 2026 Demo / reviewers in the wild / expert
Chenxi Zhang 0003
dblp:37/3579-3
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0007-1432-9957ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TraceLLM: Evaluating and Exploring Large Language Models on Trace Analysis in Microservice-based Web ApplicationsabstractTrace analysis is essential for understanding system behaviors, detecting anomalies, and diagnosing faults in complex microservice-based web applications. Existing trace analysis approaches face several challenges in industrial microservice-based systems, including high manual overhead, limited functionality, unfriendly interaction mechanisms, and difficulties in deployment and integration. The strong capabilities of large language models (LLMs) in natural language understanding, reasoning, and multi-task generalization provide new opportunities for a more intelligent and flexible trace analysis approach. However, the trace analysis capabilities of LLMs remain underexplored and underdeveloped. To bridge this gap, we conduct the first comprehensive evaluation on the trace analysis capabilities of LLMs. In particular, we construct the first instruction&response benchmark dataset for trace analysis, named TraceBench. It involves a wide range of trace analysis tasks, allowing us to systematically evaluate the capabilities of LLMs in this area. Experimental results show that LLMs have potential in handling trace analysis tasks, but there leaves room for improvement. To this end, we propose TraceLLM, an approach that significantly enhances the capabilities of LLMs via fine-tuning, outperforming the open-source LLMs by 34.77% on average in terms of accuracy, and outperforming the closed-source model by 21.66% in the best case. The generalization and robustness of TraceLLM are also confirmed in our experiments. To the best of our knowledge, TraceLLM is the first LLM which is specialized for handling various types of trace analysis tasks. This work provides a foundation for future research to further explore the trace analysis capabilities of LLMs. Xin Peng 0001, Chaofeng Sha, Chenxi Zhang 0003, Zicheng Yuan, Senyu Xie |
WWW | 5 |
| 2025 | DistriAD: Distributed Anomaly Detection for Large-Scale Microservice SystemsabstractMicroservice architecture is used by leading companies to develop their large-scale software systems. These systems comprise numerous nodes, diverse service types and instances, and substantial volumes of data. Current research usually requires a central node to collect massive data from the system to build an anomaly detection model, encountering two significant limitations: 1) Most research trains a model for the entire system, ignoring the unique characteristics of individual nodes. Additionally, processing vast system-wide data in a single node imposes significant resource demands. 2) Microservice systems change frequently, and the historical data distribution differs significantly from the real data distribution, resulting in concept drift. Thus, we proposes DistriAD, a distributed anomaly detection method specifically designed for large-scale microservice systems. DistriAD involves a lightweight anomaly detection model deployed on each distributed node for precise anomaly detection, thus enhancing its accuracy. Furthermore, DistriAD utilizes a federated learning framework and a continuous updating method incorporating human feedback to update model parameters and address concept drift. Experimental validation on public datasets, e.g., TrainTicket-based and GAIA, and a proprietary test system dataset demonstrate that DistriAD outperforms baseline methods, improving F1-score up to 39.2 %. We believe that this work can provide insights into distributed anomaly detection in large-scale microservice systems, thereby improving their performance. Yaxiao Li, Qingshan Li, Chenxi Zhang 0003, Lu Wang 0014, Chenyi Wang 0001, Zhongliang Bai, Haixing Luo, Tianyuan Gao, Lingfeng Pan |
ICWS | 3 |
| 2025 | Hypergraph Neural Network-based Multi-Granular Root Cause Localization for Microservice SystemsabstractModern enterprises are increasingly adopting microservice architectures to enhance system flexibility and scalability. However, in the face of ever-changing business requirements, the relationships between system components have become increasingly complex, resulting in significant challenges in maintaining system robustness. In recent years, multimodal data-driven approaches based on graph neural networks have emerged as a predominant solution for root cause localization in microservice systems. Our detailed analysis of architectural characteristics and existing research reveals two critical limitations. First, simple graph is insufficient to represent the one-to-many relationships inherent in microservice component interactions, such as deployment, subordinate, and dependency. Second, the current multimodal data-based method has difficulty in performing localization on faults occurring on hosts, services, and instances at the same time.To address these challenges, we propose HyperRCA, a novel multi-granular root cause analysis approach based on hypergraph neural networks. Our approach models system states during faults via a hypergraph with instances as graph nodes, explicitly capturing heterogeneous relationships through three innovative hyperedge designs: deployment hyperedges for infrastructure relationships, subordinate hyperedges for service hierarchies, and dependency hyperedges for inter-component interactions. We used hypergraph neural networks and multi-layer perceptrons to train a root cause localization model based on hyperedge features to achieve multi-granularity root cause localization. Experimental evaluations demonstrate significant performance improvements over state-of-the-art approaches. HyperRCA achieves a maximum HR@5 improvement of 112.62% on single-granularity datasets and 466.43% in multi-granularity scenarios. Yaxiao Li, Lu Wang 0014, Chenxi Zhang 0003, Qingshan Li, Siming Rong, Baiyang Wen, Quanwei Du, KeYang Li, Lingfeng Pan, Mingxuan Hui |
ASE | 3 |
| 2024 | Trace-based Multi-Dimensional Root Cause Localization of Performance Issues in Microservice SystemsabstractModern microservice systems have become increasingly complicated due to the dynamic and complex interactions and runtime environment. It leads to the system vulnerable to performance issues caused by a variety of reasons, such as the runtime environments, communications, coordinations, or implementations of services. Traces record the detailed execution process of a request through the system and have been widely used in performance issues diagnosis in microservice systems. By identifying the execution processes and attribute value combinations that are common in anomalous traces but rare in normal traces, engineers may localize the root cause of a performance issue into a smaller scope. However, due to the complex structure of traces and the large number of attribute combinations, it is challenging to find the root cause from the huge search space. In this paper, we propose TraceContrast, a trace-based multi-dimensional root cause localization approach. TraceContrast uses a sequence representation to describe the complex structure of a trace with attributes of each span. Based on the representation, it combines contrast sequential pattern mining and spectrum analysis to localize multi-dimensional root causes efficiently. Experimental studies on a widely used microservice benchmark show that TraceContrast outperforms existing approaches in both multi-dimensional and instance-dimensional root cause localization with significant accuracy advantages. Moreover, Trace-Contrast is efficient and its efficiency can be further improved by parallel execution. Chenxi Zhang 0003, Xin Peng 0001, Bicheng Zhang |
ICSE | 1 |
| 2023 | TraceStream: Anomalous Service Localization based on Trace Stream Clustering with Online FeedbackabstractModern large-scale service-based systems such as microservice systems have become increasingly complex, making it hard to localize anomalous services when various issues emerge. Traces record the workflows of requests through service instances and have been widely used in anomaly detection and root cause analysis. Existing trace-based approaches widely use statistical methods or learning-based techniques to detect trace anomalies and localize anomalous services. However, these approaches often suffer from the concept drift problem, i.e., the statistical properties of traces change over time in unforeseen ways. In this paper, we propose TraceStream, an anomalous service localization approach based on trace data stream clustering. TraceStream uses data stream clustering to discover potential anomalous trace clusters in evolving trace data and uses spectrum analysis to localize anomalous services based on the clusters. Moreover, TraceStream can effectively incorporate the online feedback of operation engineers based on the trace clusters to improve the accuracy for localizing anomalous services. Our evaluation confirms that TraceStream can effectively detect anomalies and localize anomalous services in an evolving microservice system. It can effectively incorporate human feedback to further improve the performance of anomalous service localization. Moreover, TraceStream is efficient and its efficiency can be further improved by sampling a small portion of traces by cluster. Chenxi Zhang 0003, Xin Peng 0001, Zhenghui Yan, Pairui Li, Jianming Liang, Haibing Zheng, Wujie Zheng, Yuetang Deng |
ISSRE | 2 |
| 2023 | Dynamic Graph Neural Networks-Based Alert Link Prediction for Online Service SystemsabstractA fault in large online service systems often triggers numerous alerts due to the complex business and component dependencies among services, which is known as “alert storm”. In a short time, an online service system may generate a huge amount of alert data. This poses a challenge for on-call engineers to identify alerts that are associated with a system failure for root cause analysis. In this paper, we propose DyAlert, a dynamic graph neural networks-based approach for linking alerts that might be triggered by a same fault to reduce the burden of on-call engineers in the fault analysis. Our insight is that alerts are often triggered by alert propagation when a system failure occurs, e.g., alert$a$would lead to the occurrence of alert$b$. Whether two alerts should be linked depends on if one alert is triggered by the propagation of the other. Leveraging this insight, we design a dynamic graph (namely Alert-Metric Dynamic Graph) that describes the propagation process of alerts. Based on the dynamic graph, we train a neural networks-based model to predict alert links. We evaluate DyAlert with real-world data collected from an online service system running 85 business units and about 30,000 different services in a large enterprise. The results show that DyAlert is effective in predicting alert links and it outperforms the state-of-the-art approaches with an average increase of 0.259 in F1-score. Chenxi Zhang 0003, Dingyu Yang, Xin Peng 0001, Jiayu Ou, Zheshun Wu, Xiaojun Qu, Wei Li 0075 |
ASE | 2 |
| 2022 | DeepTraLog: Trace-Log Combined Microservice Anomaly Detection through Graph-based Deep LearningabstractA microservice system in industry is usually a large-scale distributed system consisting of dozens to thousands of services running in different machines. An anomaly of the system often can be reflected in traces and logs, which record inter-service interactions and intra-service behaviors respectively. Existing trace anomaly detection approaches treat a trace as a sequence of service invocations. They ignore the complex structure of a trace brought by its invocation hierarchy and parallel/asynchronous invocations. On the other hand, existing log anomaly detection approaches treat a log as a sequence of events and cannot handle microservice logs that are distributed in a large number of services with complex interactions. In this paper, we propose DeepTraLog, a deep learning based microservice anomaly detection approach. DeepTraLog uses a unified graph representation to describe the complex structure of a trace together with log events embedded in the structure. Based on the graph representation, DeepTraLog trains a GGNNs based deep SVDD model by combing traces and logs and detects anomalies in new traces and the corresponding logs. Evaluation on a microservice benchmark shows that DeepTraLog achieves a high precision (0.93) and recall (0.97), outperforming state-of-the-art trace/log anomaly detection approaches with an average increase of 0.37 in F1-score. It also validates the efficiency of DeepTraLog, the contribution of the unified graph representation, and the impact of the configurations of some key parameters. Chenxi Zhang 0003, Xin Peng 0001, Chaofeng Sha, Zhenqing Fu, Xiya Wu, Qingwei Lin, Dongmei Zhang 0001 |
ICSE | 1 |
| 2022 | PUTraceAD: Trace Anomaly Detection with Partial Labels based on GNN and PU LearningabstractDistributed tracing has been an important part of microservice infrastructure and learning-based trace analysis has been used to detect anomalies in microservice systems. Existing learning-based trace anomaly detection approaches ei-ther assume that trace patterns can be learned from normal execution or rely on fault injection to produce labeled traces (i.e., normal/anomalous ones). However, in practice it is often difficult to ensure that the normal execution does not involve anomalous traces or obtain a large variety of normal and anomalous traces through fault injection. In this paper, we propose PUTraceAD, a trace anomaly detection approach that can alleviate the above problems. PUTraceAD represents a trace as a span causal graph with node features such as operation name, response code, duration time. Based on the graph representation, PUTraceAD trains a GNN- and PU learning-based trace anomaly detection model. During the process, PU (Positive and Unlabeled) learning optimizes model parameters through estimating the data distribution. Therefore, PUTraceAD can train the model based on a small set of labeled anomalous traces and a large set of unlabeled traces. Our evaluation shows that PUTraceAD outperforms existing unsupervised trace anomaly detection approaches and only slightly underperforms a supervised learning-based approach that takes full advantage of labeled traces. Chenxi Zhang 0003, Xin Peng 0001, Chaofeng Sha |
ISSRE | 2 |
| 2022 | Trace analysis based microservice architecture measurementabstractMicroservice architecture design highly relies on expert experience and may often result in improper service decomposition. Moreover, a microservice architecture is likely to degrade with the continuous evolution of services. Architecture measurement is thus important for the long-term evolution of microservice architectures. Due to the independent and dynamic nature of services, source code analysis based approaches cannot well capture the interactions between services. In this paper, we propose a trace analysis based microservice architecture measurement approach. We define a trace data model for microservice architecture measurement, which enables fine-grained analysis of the execution processes of requests and the interactions between interfaces and services. Based on the data model, we define 14 architectural metrics to measure the service independence and invocation chain complexity of a microservice system. We implement the approach and conduct three case studies with a student course project, an open-source microservice benchmark system, and three industrial microservice systems. The results show that our approach can well characterize the independence and invocation chain complexity of microservice architectures and help developers to identify microservice architecture issues caused by improper service decomposition and architecture degradation. Xin Peng 0001, Chenxi Zhang 0003, Akasaka Isami, Yunna Cui |
ESEC/SIGSOFT FSE | 2 |
| 2022 | TraceCRL: contrastive representation learning for microservice trace analysisabstractDue to the large amount and high complexity of trace data, microservice trace analysis tasks such as anomaly detection, fault diagnosis, and tail-based sampling widely adopt machine learning technology. These trace analysis approaches usually use a preprocessing step to map structured features of traces to vector representations in an ad-hoc way. Therefore, they may lose important information such as topological dependencies between service operations. In this paper, we propose TraceCRL, a trace representation learning approach based on contrastive learning and graph neural network, which can incorporate graph structured information in the downstream trace analysis tasks. Given a trace, TraceCRL constructs an operation invocation graph where nodes represent service operations and edges represent operation invocations together with predefined features for invocation status and related metrics. Based on the operation invocation graphs of traces TraceCRL uses a contrastive learning method to train a graph neural network-based model for trace representation. In particular, TraceCRL employs six trace data augmentation strategies to alleviate the problems of class collision and uniformity of representation in contrastive learning. Our experimental studies show that TraceCRL can significantly improve the performance of trace anomaly detection and offline trace sampling. It also confirms the effectiveness of the trace augmentation strategies and the efficiency of TraceCRL. Chenxi Zhang 0003, Xin Peng 0001, Chaofeng Sha, Zhenghui Yan |
ESEC/SIGSOFT FSE | 1 |