Wanhao Zhang

dblp:272/6189 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2024
0009-0007-9446-5819ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2024 LogRAG: Semi-Supervised Log-based Anomaly Detection with Retrieval-Augmented Generation
abstract
Log-based anomaly detection is critical in monitoring the operation of microservice systems and in the realtime reporting of system failures. Utilizing deep learning-based log anomaly detection methods facilitates effective detection of anomalies within logs. However, existing methods are greatly dependent on log parsers, and parsing errors can considerably affect downstream anomaly detection tasks. Additionally, methods that predict the next log event in a sequence are susceptible to the instability of sequences and the emergence of unseen logs as systems evolve, resulting in a higher false positive rate. In this paper, we propose a semi-supervised log anomaly detection framework based on retrieval-augmented generation (RAG). This framework conducts phased detection using both Log Tokens and Log Templates to mitigate the impact of log parsing errors. It also utilizes a single-class classifier to model the normal behavior of the system, thereby circumventing the effects of unstable sequences. Finally, it employs large language model (LLM) empowered by RAG to reevaluate detected anomalous logs.
Wanhao Zhang, Qianli Zhang, Enyu Yu, Yuxiang Ren, Yeqing Meng, Mingxi Qiu, Jilong Wang 0001
ICWS1
2024 Leveraging RAG-Enhanced Large Language Model for Semi-Supervised Log Anomaly Detection
abstract
Log-based anomaly detection is critical in monitoring the operations of information systems and in the real-time reporting of system failures. Utilizing deep learning-based log anomaly detection methods facilitates effective detection of anomalies within logs. However, existing methods are greatly dependent on log parsers, and parsing errors can considerably affect downstream anomaly detection tasks. Additionally, methods that predict the next log event in a sequence are susceptible to the instability of sequences and the emergence of unseen logs as systems evolve, resulting in a higher false positive rate. In this paper, we put forward LogRAG, a semi-supervised log anomaly detection framework based on retrieval-augmented generation (RAG). This framework conducts phased detection using both Log Tokens and Log Templates to mitigate the impact of log parsing errors. It also utilizes a single-class classifier to model the normal behavior of the system, thereby circumventing the effects of unstable sequences. Finally, it employs large language model (LLM) empowered by RAG to reevaluate detected anomalous logs, thereby improving accuracy. LogRAG demonstrates a 15% improvement in F1 Score on the BGL dataset and a 60% improvement on the Spirit dataset when compared to the previous best semi-supervised learning algorithm.
Wanhao Zhang, Qianli Zhang, Enyu Yu, Yuxiang Ren, Yeqing Meng, Mingxi Qiu, Jilong Wang 0001
ISSRE1
2023 FTM-RCA: A Fast Two-Stage Multi-dimensional Root-Cause Analysis of Network Anomalies
abstract
Multi-dimensional Root Cause Analysis (RCA) is often applied to identify abnormal traffic patterns, i.e., localizing the abnormal combination of traffic header fields. Several techniques have been proposed recently, but they were mainly designed for smaller-scaled datasets and were not feasible in the real network due to the high computational overhead. To overcome the aforementioned limitations, we propose FTM-RCA, which accelerates RCA by breaking the analysis procedure into two stages: coarse-grained rules filtering and fine-grained localization. In the first stage, an optimized frequent itemset mining (FIM) technique called CUSC is proposed, which can detect high-volume combinations faster based on the mutual exclusion of dimension values. Experiments on CUSC show that it can speed up by 44.87% and reduce memory consumption by 21.89% compared to the best previous FIM algorithms. In the second stage, a dimension-based search method is proposed to identify the root cause combinations, which consists of two key components: 1) drill-down strategy, which utilizes Contributive Power to measure the correlation between the combination and anomaly. 2) pruning strategy, which adopts the Shannon entropy to avoid generating trivial results. As a result, the overall diagnostic time of FTM-RCA is at least 25 times faster than the previous best research while improving accuracy by an average of 21.6%. Also, our practical application in real network also illustrates the applicability of FTM-RCA.
Yeqing Meng, Qianli Zhang, Xiangyu Tang, Wanhao Zhang, Jilong Wang 0001
IWQoS4
2023 Neural-Hidden-CRF: A Robust Weakly-Supervised Sequence Labeler
abstract
We propose a neuralized undirected graphical model called Neural-Hidden-CRF to solve the weakly-supervised sequence labeling problem. Under the umbrella of undirected graphical theory, the proposed Neural-Hidden-CRF embedded with a hidden CRF layer models the variables of word sequence, latent ground truth sequence, and weak label sequence with the global perspective that undirected graphical models particularly enjoy. In Neural-Hidden-CRF, we can capitalize on the powerful language model BERT or other deep models to provide rich contextual semantic knowledge to the latent ground truth sequence, and use the hidden CRF layer to capture the internal label dependencies. Neural-Hidden-CRF is conceptually simple and empirically powerful. It obtains new state-of-the-art results on one crowdsourcing benchmark and three weak-supervision benchmarks, including outperforming the recent advanced model CHMM by 2.80 F1 points and 2.23 F1 points in average generalization and inference performance, respectively.
Hailong Sun 0001, Wanhao Zhang, Chunyi Xu, Qianren Mao
KDD3