Lingzhe Zhang

dblp:325/9388 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models
abstract
Leyi Pan, Shuchang Tao, Yunpeng Zhai, Zheyu Fu, Liancheng Fang, Minghua He, Lingzhe Zhang, Zhaoyang Liu, Bolin Ding, Aiwei Liu, Lijie Wen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Leyi Pan, Shuchang Tao, Yunpeng Zhai, Zheyu Fu, Liancheng Fang, Minghua He, Lingzhe Zhang, Zhaoyang Liu 0003, Bolin Ding, Aiwei Liu, Lijie Wen 0001
ACL (1)7
2026 UDA-RCL: Unsupervised Domain Adaptation for Microservice Root Cause Localization Utilizing Multimodal Data
abstract
Root cause Localization methods play a crucial role in ensuring the stability of large-scale microservice systems. However, existing methods either rely on unsupervised approaches with limited localization accuracy or on supervised learning that requires large volumes of historical anomaly data. Such data is often unavailable in newly deployed systems. To address this limitation, we attempt to use the labeled data from mature systems to help new systems build root cause localization models. Specifically, we propose UDA-RCL, an unsupervised domain adaptation root cause localization method using multimodal data (log, metric and trace data). UDA-RCL first incorporates an aggregation based event extraction module to standardize the format of multimodal data from different systems. Then, it utilizes a multimodal event encoder and multimodal domain adversarial adaptation module to narrow the feature distribution gap between different systems. Furthermore, taking into account the situation of sparse anomaly samples, existing methods' classifiers are either hard to transfer to other systems or struggle to capture the process of anomaly propagation, we propose a PageRank classifier module. This module employs a neural network embedded with anomaly propagation rules to output the final root cause ranking results, alleviating the issue of sparse anomaly samples. Extensive experiments have proven that our method achieves the best results in both supervised and transfer learning scenarios.
Xiaosong Huang, Yifan Wu 0002, Lingzhe Zhang, Ying Li 0012, Zhonghai Wu
IEEE Trans. Serv. Comput.4
2025 ScalaLog: Scalable Log-Based Failure Diagnosis Using LLM
abstract
As Industrial Internet of Things (IIoT) software systems become increasingly complex, precise failure diagnosis has become both essential and challenging. Current log-based failure diagnosis methods lack scalability for different failure types. In IIoT software systems, the number of failure types is constantly growing, and retraining the model each time a new failure type is introduced is highly resource-intensive. Additionally, traditional log-based failure diagnosis models often require log parsing as a preliminary step, which can also be resource-consuming. To address these challenges, we propose a scalable log-based failure diagnosis method named ScalaLog. ScalaLog builds on RAG by utilizing LLM-based summarization to extract key log information, applying sample augmentation to increase the number of samples, and using CoT prompts to guide the LLM in failure diagnosis. Experiments on various public and real-world datasets demonstrate that ScalaLog significantly enhances failure diagnosis accuracy without the need for training or log parsing.
Lingzhe Zhang, Mengxi Jia, Yifan Wu 0002, Ying Li 0012
ICASSP1
2025 AAAD: Asynchronous Inter-Variable Relationship-Aware Anomaly Detection for Multivariate Time Series
abstract
Anomaly detection in multivariate time series (MTS) plays a crucial role in various domains, particularly in multimedia. While significant progress has been made in modeling normal data patterns and detecting anomalies based on deviations, existing methods face challenges in capturing complex inter-variable relationships, particularly in the presence of asynchronous dependencies. To address the challenges of detecting anomalies in MTS with complex inter-variable relationships, we propose an asynchronous inter-variable relationship-aware anomaly detection method. This approach simultaneously extracts both synchronous and asynchronous feature pairs between variables and leverages attention mechanisms to automatically learn unified inter-variable relationships across these dependencies. Additionally, we utilize a memory network to store and dynamically update the normal patterns of inter-variable relationships, computing anomaly scores based on deviations in these relationships. Extensive experiments demonstrate that our method outperforms existing baseline approaches, achieving an average F1 score of 97.05% across five benchmark datasets. Ablation studies further validate the effectiveness of each component of our method.
Xiaosong Huang, Mengxi Jia, Lingzhe Zhang, Zhonghai Wu, Ying Li 0012
ICME4
2025 ORA: Job Runtime Prediction for High-Performance Computing Platforms Using the Online Retrieval-Augmented Language Model
abstract
Accurate job runtime prediction is critical for efficient scheduling in high-performance computing (HPC) platforms.For instance, precise predictions enable techniques such as backfilling, where small jobs are executed ahead of schedule to maximize resource utilization and enhance computational efficiency.However, existing runtime prediction methods primarily rely on job metadata (e.g., submission time, requested runtime, and required memory) while ignoring the content of job scripts, which limits their accuracy.To address this issue, we propose an Online Retrieval-Augmented Language Model (ORA) for job runtime prediction.ORA encodes both metadata and script information from historical jobs into feature vectors to form a database, enabling similarity-based retrieval to assist in predicting the runtime of new jobs.To address distribution shifts, ORA incrementally updates the database without requiring model retraining.Additionally, personalized retrieval mechanisms are employed to mitigate the * Co-corresponding author.
Yinping Ma, Xiaosong Huang, Lingzhe Zhang, Ying Li 0012
ICS4
2025 CSLParser: A Collaborative Framework Using Small and Large Language Models for Log Parsing
abstract
Log parsing is a prerequisite for log analysis. Recently, large language models (LLMs) have demonstrated high accuracy in log parsing. However, their frequent invocations incur substantial costs. To address this issue, some methods have turned to small language models (SLMs), which offer improved efficiency but suffer from reduced accuracy due to limited model capacity. To achieve both high accuracy and efficiency, we propose CSLParser, a collaborative log parsing framework using SLMs and LLMs. CSLParser delegates most log parsing tasks to SLMs and selectively invokes LLMs to correct parsing results generated by SLMs, thereby effectively reducing the invocation cost of LLMs while maintaining high accuracy. Specifically, to enhance the accuracy of SLMs, we propose a diversified sampling strategy to select diverse samples for training, enabling SLMs to effectively handle diverse log patterns. To efficiently invoke LLMs, we design a rule-based selection strategy to identify hard cases that are challenging for SLMs to correctly parse, which are subsequently corrected by LLMs. Additionally, we propose a dynamic template updating mechanism that merges similar templates based on structural and semantic information to further enhance parsing accuracy. Extensive experiments on public large-scale log datasets show that CSLParser outperforms state-of-the-art baselines in both accuracy and efficiency.
Weijie Hong, Yifan Wu 0002, Lingzhe Zhang, Chiming Duan, Pei Xiao 0005, Minghua He, Xixuan Yang, Ying Li 0012
ISSRE3
2025 LogAction: Consistent Cross-system Anomaly Detection through Logs via Active Domain Adaptation
abstract
Log-based anomaly detection is a essential task for ensuring the reliability and performance of software systems. However, the performance of existing anomaly detection methods heavily relies on labeling, while labeling a large volume of logs is highly challenging. To address this issue, many approaches based on transfer learning and active learning have been proposed. Nevertheless, their effectiveness is hindered by issues such as the gap between source and target system data distributions and cold-start problems. In this paper, we propose LogAction, a novel log-based anomaly detection model based on active domain adaptation. LogAction integrates transfer learning and active learning techniques. On one hand, it uses labeled data from a mature system to train a base model, mitigating the cold-start issue in active learning. On the other hand, LogAction utilize free energy-based sampling and uncertainty-based sampling to select logs located at the distribution boundaries for manual labeling, thus addresses the data distribution gap in transfer learning with minimal human labeling efforts. Experimental results on six different combinations of datasets demonstrate that LogAction achieves an average 93.01% F1 score with only 2% of manual labels, outperforming some state-of-the-art methods by 26.28%. Website: https://logaction.github.io
Chiming Duan, Minghua He, Pei Xiao 0005, Zhewei Zhong, Yan Niu, Lingzhe Zhang, Siyu Yu, Yifan Wu 0002, Weijie Hong, Ying Li 0012, Gang Huang 0001
ASE9
2025 United We Stand: Towards End-to-End Log-based Fault Diagnosis via Interactive Multi-Task Learning
abstract
Log-based fault diagnosis is essential for maintaining software system availability. However, existing fault diagnosis methods are built using a task-independent manner, which fails to bridge the gap between anomaly detection and root cause localization in terms of data form and diagnostic objectives, resulting in three major issues: 1) Diagnostic bias accumulates in the system; 2) System deployment relies on expensive monitoring data; 3) The collaborative relationship between diagnostic tasks is overlooked. Facing this problems, we propose a novel end-to-end log-based fault diagnosis method, Chimera, whose key idea is to achieve end-to-end fault diagnosis through bidirectional interaction and knowledge transfer between anomaly detection and root cause localization. Chimera is based on interactive multitask learning, carefully designing interaction strategies between anomaly detection and root cause localization at the data, feature, and diagnostic result levels, thereby achieving both sub-tasks interactively within a unified end-to-end framework. Evaluation on two public datasets and one industrial dataset shows that Chimera outperforms existing methods in both anomaly detection and root cause localization, achieving improvements of over 2.92%~5.00% and 19.01% ~ 37.09%, respectively. It has been successfully deployed in production, serving an industrial cloud platform.
Minghua He, Chiming Duan, Pei Xiao 0005, Siyu Yu, Lingzhe Zhang, Weijie Hong, Yifan Wu 0002, Ying Li 0012, Gang Huang 0001
ASE6
2025 Walk the Talk: Is Your Log-based Software Reliability Maintenance System Really Reliable?
abstract
Log-based software reliability maintenance systems are crucial for sustaining stable customer experience. However, existing deep learning-based methods represent a black box for service providers, making it impossible for providers to understand how these methods detect anomalies, thereby hindering trust and deployment in real production environments. To address this issue, this paper defines a trustworthiness metric—diagnostic faithfulness—for models to gain service providers’ trust, based on surveys of SREs at a major cloud provider. We design two evaluation tasks: attention-based root cause localization and event perturbation. Empirical studies demonstrate that existing methods perform poorly in diagnostic faithfulness. Consequently, we propose FaithLog, a faithful log-based anomaly detection system, which achieves faithfulness through a carefully designed causality-guided attention mechanism and adversarial consistency learning. Evaluation results on two public datasets and one industrial dataset demonstrate that the proposed method achieves state-of-the-art performance in diagnostic faithfulness.
Minghua He, Chiming Duan, Pei Xiao 0005, Lingzhe Zhang, Kangjin Wang, Yifan Wu 0002, Ying Li 0012, Gang Huang 0001
ASE5
2025 CoorLog: Efficient-Generalizable Log Anomaly Detection via Adaptive Coordinator in Software Evolution
abstract
Frequent software updates lead to log evolution, posing generalization challenges for current log anomaly detection. Traditional log anomaly detection research focuses on using small deep learning models (SMs), but these models inherently lack generalization due to their closed-world assumption. Large language models (LLMs) exhibit strong semantic understanding and generalization capabilities, making them promising for log anomaly detection. However, they suffer from computational inefficiencies. To balance efficiency and generalization, we propose a collaborative log anomaly detection scheme (CoorLog) that uses an adaptive coordinator to integrate SM and LLM. The coordinator determines if incoming logs have evolved. Non-evolved logs are routed to the SM, while evolved logs are directed to the LLM for detailed inference using the constructed Evol-CoT. To gradually adapt to evolution, we introduce the adaptive evolution mechanism (AEM), which updates the coordinator to redirect evolved logs identified by the LLM to the SM. Simultaneously, the SM is fine-tuned to inherit the LLM’s judgment on these logs. Extensive experiments on real-world datasets demonstrate that CoorLog achieves superior F1-scores in both intra-version and inter-version anomaly detection. Additionally, CoorLog reduces processing time by 91.63% and token consumption by 85.59% compared to using an LLM alone.
Pei Xiao 0005, Chiming Duan, Minghua He, Yifan Wu 0002, Gege Gao, Lingzhe Zhang, Weijie Hong, Ying Li 0012, Gang Huang 0001
ASE8
2025 Hybrid physics-based and data-driven method for the rotor angle prediction
Lingzhe Zhang, Huaiyuan Wang
Eng. Appl. Artif. Intell.1
2025 Towards Close-to-Zero Runtime Collection Overhead: Raft-Based Anomaly Diagnosis on System Faults for Distributed Storage System
abstract
Distributed storage systems are fundamental infrastructures of today’s large-scale software systems such as cloud systems. Diagnosing anomalies in distributed storage systems is essential for maintaining software availability. Existing anomaly diagnosis approaches mainly rely on the run-time data including monitoring data and application logs. However, collecting and analyzing the run-time data requires huge computing, storage, and management costs. Typically, more fine-grained run-time data can reveal more symptoms of anomalies, but on the contrary, requires more computing, storage, and management costs. As a result, solving the anomaly diagnosis problem is a balancing between the quality of run-time data and system overhead or cost. In this paper, we take into account both data quality and system overhead or cost by introducing a new type of run-time data-Raft logs. Raft logs are naturally produced by distributed storage systems and collecting raft logs will not bring any extra system overhead. To verify the ability of Raft logs in reflecting anomalies, we conduct a comprehensive study on the interconnection between the anomalies and Raft logs. Based on the study, we propose an effectiveRaft-BasedAnomalyDiagnosis approach namedRBAD. For evaluation, we expose the first open-sourced comprehensive dataset with multiple runtime data containing both Raft logs, application logs and monitoring data. Experiments based on this dataset demonstrate RBAD’s superiority, outperforming monitoring-based methods by 15.38% and log-based methods by 53.10%.
Lingzhe Zhang, Mengxi Jia, Yong Yang 0011, Zhonghai Wu, Ying Li 0012
IEEE Trans. Serv. Comput.1
2025 E-Log: Fine-Grained Elastic Log-Based Anomaly Detection and Diagnosis for Databases
abstract
Database Management Systems (DBMS) form the backbone of modern large-scale software systems, where reliable anomaly detection and diagnosis are essential for ensuring system availability. However, existing log-based methods often impose significant performance overhead by collecting large volumes of logs, which is impractical for DBMS requiring high read/write throughput. This paper addresses a critical yet underexplored challenge: how to balance logging granularity with runtime efficiency for effective anomaly management in databases. We presentE-Log, a novel fine-grained elastic log-based framework for anomaly detection and diagnosis. E-Log intelligently adjusts the amount and detail of logging based on system state—maintaining lightweight logging during normal operation for efficient anomaly detection, and triggering rich, informative logging only upon anomaly suspicion for accurate diagnosis. This adaptive strategy significantly reduces runtime overhead while preserving diagnostic precision. We implement E-Log on Apache IoTDB and evaluate it using benchmarks including TSBS, TPCx-IoT, and IoT-Bench. Experimental results show that E-Log improves anomaly detection accuracy by 3.15% and diagnosis performance by 9.32% compared to state-of-the-art methods. Moreover, it reduces log storage size by 43.53% and increases average write throughput by 26.22%. These results highlight E-Log's potential to enable efficient, accurate, and scalable anomaly management in high-performance database systems.
Lingzhe Zhang, Mengxi Jia, Zhonghai Wu, Ying Li 0012
IEEE Trans. Serv. Comput.1
2024 Reducing Events to Augment Log-based Anomaly Detection Models: An Empirical Study
abstract
As software systems grow increasingly intricate, the precise detection of anomalies have become both essential and challenging. Current log-based anomaly detection methods depend heavily on vast amounts of log data leading to inefficient inference and potential misguidance by noise logs. However, the quantitative effects of log reduction on the effectiveness of anomaly detection remain unexplored. Therefore, we first conduct a comprehensive study on six distinct models spanning three datasets. Through the study, the impact of log quantity and their effectiveness in representing anomalies is qualifies, uncovering three distinctive log event types that differently influence model performance. Drawing from these insights, we propose LogCleaner: an efficient methodology for the automatic reduction of log events in the context of anomaly detection. Serving as middleware between software systems and models, LogCleaner continuously updates and filters anti-events and duplicative-events in the raw generated logs. Experimental outcomes highlight LogCleaner’s capability to reduce over 70% of log events in anomaly detection, accelerating the model’s inference speed by approximately 300%, and universally improving the performance of models for anomaly detection.
Lingzhe Zhang, Kangjin Wang, Mengxi Jia, Yong Yang 0011, Ying Li 0012
ESEM1
2024 Multivariate Log-based Anomaly Detection for Distributed Database
abstract
Distributed databases are fundamental infrastructures of today's large-scale software systems such as cloud systems. Detecting anomalies in distributed databases is essential for maintaining software availability. Existing approaches, predominantly developed using Loghub-a comprehensive collection of log datasets from various systems-lack datasets specifically tailored to distributed databases, which exhibit unique anomalies. Additionally, there's a notable absence of datasets encompassing multi-anomaly, multi-node logs. Consequently, models built upon these datasets, primarily designed for standalone systems, are inadequate for distributed databases, and the prevalent method of deeming an entire cluster anomalous based on irregularities in a single node leads to a high false-positive rate. This paper addresses the unique anomalies and multivariate nature of logs in distributed databases. We expose the first open-sourced, comprehensive dataset with multivariate logs from distributed databases. Utilizing this dataset, we conduct an extensive study to identify multiple database anomalies and to assess the effectiveness of state-of-the-art anomaly detection using multivariate log data. Our findings reveal that relying solely on logs from a single node is insufficient for accurate anomaly detection on distributed database. Leveraging these insights, we propose MultiLog, an innovative multivariate log-based anomaly detection approach tailored for distributed databases. Our experiments, based on this novel dataset, demonstrate MultiLog's superiority, outperforming existing state-of-the-art methods by approximately 12%.
Lingzhe Zhang, Mengxi Jia, Ying Li 0012, Yong Yang 0011, Zhonghai Wu
KDD1
2024 Time-tired compaction: An elastic compaction scheme for LSM-tree based time-series database
Lingzhe Zhang, Xiangdong Huang 0001, Yan-Kai Wang, Jialin Qiao, Shaoxu Song, Jianmin Wang 0001
Adv. Eng. Informatics1
2022 Separation or Not: On Handing Out-of-Order Time-Series Data in Leveled LSM-Tree
abstract
LSM-Tree is widely adopted for storing time-series data in Internet of Things. According to conventional policy (denoted by$\pi_{c}$), when writing, the data will first be buffered in MemTable in memory. When it is full, the data will be written to the disk to form SSTables. Compaction is triggered to sort the data in each layer of the LSM-Tree on the disk. However, the arrival of data can be unordered due to reasons such as transition delay. Apache IoTDB uses in-order and out-of-order MemTables to separately buffer the in-order and out-of-order data to accelerate queries, namely the separation policy (denoted by$\pi_{s}$). However, given a specific space of memory budget to buffer the data, write amplification (WA) of the leveled LSM-Tree will be influenced by$\pi_{s}$. Whether the influence by separation is positive or negative, and how intense WA is influenced, depend on the properties of workloads and the capacity of the in-order and out-of-order MemTables. It is highly demanded to build robust models for estimating the expected amount of data rewritten in each compaction, and predicting the WA under$\pi_{c}$and$\pi_{s}$. Note that as an industrial paper, rather than proposing novel techniques for research problems, we focus on the practice of whether separating or not for lower write amplification. Experiments on synthetic and real-world datasets show that the models for estimating WA are accurate under various delay distributions. In addition, based on the estimation models, we implement an analyzer module in the open-source Apache IoTDB, for choosing the policy with lower WA. We apply the method in the use case of our industrial partner, a service provider of engineering machinery. The use case verifies the effectiveness of deciding whether separation or not by WA estimation.
Yuyuan Kang, Xiangdong Huang 0001, Shaoxu Song, Lingzhe Zhang, Jialin Qiao, Chen Wang 0018, Jianmin Wang 0001, Julian Feinauer
ICDE4