Qianli Zhang

dblp:30/1560 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
3since 2021 · last 2024
0000-0003-2084-7762ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 3 · 1 first-author · 1 since 2021Security and privacy · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 LogRAG: Semi-Supervised Log-based Anomaly Detection with Retrieval-Augmented Generation
abstract
Log-based anomaly detection is critical in monitoring the operation of microservice systems and in the realtime reporting of system failures. Utilizing deep learning-based log anomaly detection methods facilitates effective detection of anomalies within logs. However, existing methods are greatly dependent on log parsers, and parsing errors can considerably affect downstream anomaly detection tasks. Additionally, methods that predict the next log event in a sequence are susceptible to the instability of sequences and the emergence of unseen logs as systems evolve, resulting in a higher false positive rate. In this paper, we propose a semi-supervised log anomaly detection framework based on retrieval-augmented generation (RAG). This framework conducts phased detection using both Log Tokens and Log Templates to mitigate the impact of log parsing errors. It also utilizes a single-class classifier to model the normal behavior of the system, thereby circumventing the effects of unstable sequences. Finally, it employs large language model (LLM) empowered by RAG to reevaluate detected anomalous logs.
Wanhao Zhang, Qianli Zhang, Enyu Yu, Yuxiang Ren, Yeqing Meng, Mingxi Qiu, Jilong Wang 0001
ICWS2
2024 Leveraging RAG-Enhanced Large Language Model for Semi-Supervised Log Anomaly Detection
abstract
Log-based anomaly detection is critical in monitoring the operations of information systems and in the real-time reporting of system failures. Utilizing deep learning-based log anomaly detection methods facilitates effective detection of anomalies within logs. However, existing methods are greatly dependent on log parsers, and parsing errors can considerably affect downstream anomaly detection tasks. Additionally, methods that predict the next log event in a sequence are susceptible to the instability of sequences and the emergence of unseen logs as systems evolve, resulting in a higher false positive rate. In this paper, we put forward LogRAG, a semi-supervised log anomaly detection framework based on retrieval-augmented generation (RAG). This framework conducts phased detection using both Log Tokens and Log Templates to mitigate the impact of log parsing errors. It also utilizes a single-class classifier to model the normal behavior of the system, thereby circumventing the effects of unstable sequences. Finally, it employs large language model (LLM) empowered by RAG to reevaluate detected anomalous logs, thereby improving accuracy. LogRAG demonstrates a 15% improvement in F1 Score on the BGL dataset and a 60% improvement on the Spirit dataset when compared to the previous best semi-supervised learning algorithm.
Wanhao Zhang, Qianli Zhang, Enyu Yu, Yuxiang Ren, Yeqing Meng, Mingxi Qiu, Jilong Wang 0001
ISSRE2
2023 FTM-RCA: A Fast Two-Stage Multi-dimensional Root-Cause Analysis of Network Anomalies
abstract
Multi-dimensional Root Cause Analysis (RCA) is often applied to identify abnormal traffic patterns, i.e., localizing the abnormal combination of traffic header fields. Several techniques have been proposed recently, but they were mainly designed for smaller-scaled datasets and were not feasible in the real network due to the high computational overhead. To overcome the aforementioned limitations, we propose FTM-RCA, which accelerates RCA by breaking the analysis procedure into two stages: coarse-grained rules filtering and fine-grained localization. In the first stage, an optimized frequent itemset mining (FIM) technique called CUSC is proposed, which can detect high-volume combinations faster based on the mutual exclusion of dimension values. Experiments on CUSC show that it can speed up by 44.87% and reduce memory consumption by 21.89% compared to the best previous FIM algorithms. In the second stage, a dimension-based search method is proposed to identify the root cause combinations, which consists of two key components: 1) drill-down strategy, which utilizes Contributive Power to measure the correlation between the combination and anomaly. 2) pruning strategy, which adopts the Shannon entropy to avoid generating trivial results. As a result, the overall diagnostic time of FTM-RCA is at least 25 times faster than the previous best research while improving accuracy by an average of 21.6%. Also, our practical application in real network also illustrates the applicability of FTM-RCA.
Yeqing Meng, Qianli Zhang, Xiangyu Tang, Wanhao Zhang, Jilong Wang 0001
IWQoS2
2019 A survey on resource scheduling for data transfers in inter-datacenter WANs
Hui Wang 0011, Jilong Wang 0001, Changqing An, Qianli Zhang
Comput. Networks4
2014 UDP traffic classification using most distinguished port
abstract
Comparing to TCP traffic, the composition of UDP traffic is still unclear. Although it is observed that a large fraction of UDP traffic appears to be P2P applications, application level classification of UDP traffic is still very hard since most of these applications are private protocols based. In this paper, a novel method is proposed to classify UDP traffic. Based on the assumption that traffic from two communicating half-tuples identified by theis from the same application, all half-tuples can be grouped into several connected subgraphs. The port numbers which are adopted by most links or half-tuples in each subgroup can thus be used to characterize the application types of the whole subgroup. Experiment results show that this approach is feasible and can classify UDP traffic only using flow level information. The port numbers adopted by most links or half-tuples are surprisingly stable among different time periods, for example, for Youku application remain the same for more than 90% of periods in all the 1429 periods.
Qianli Zhang, Jilong Wang 0001, Xing Li 0001
APNOMS1
2008 DRAGON-Lab - Next generation internet technology experiment platform
Jilong Wang 0001, ZhongHui Li, Guohan Lu, Caiping Jiang, Xing Li 0001, Qianli Zhang
Sci. China Ser. F Inf. Sci.6
2007 On the Design of Fast Prefix-Preserving IP Address Anonymization Scheme
Qianli Zhang, Jilong Wang 0001, Xing Li 0001
ICICS1
2006 SANTT: Sharing Anonymized Network Traffic Traces among Researchers
abstract
Current Internet research suffers from limited information available from public network traffic traces due to the privacy concern of ISP and lack of effective trace distribution systems. In this paper, we present our SANTT (sharing anonymized network traces) system to share valuable network traffic traces safely and widely. SANTT employs a novel prefix-preserving anonymization method to sanitize privacy information in packets and takes advantage of specific high-speed traffic capturing hardware. After removing privacy information, SANTT distributes the traces by multicast, which makes it easy for researchers to access. The experimental results show that SANTT implemented on IXP2400 platform can process the fastest traffic rate (1,1953,125 PPS) offered by Gigabit links with high accuracy and stability
Xiaoxin Shao, Qianli Zhang, Shijin Kong, Changqing An, Xing Li 0001
NOMS2
2003 Evolving training model method for one-class SVM
abstract
This paper proposes and analyzes an evolving training model method for selecting the best training parameters of one-class support vector machines (SVM). The method: 1) presents and computes effectively the generalization performance of one-class SVM, including using fraction of support vectors and /spl xi//spl alpha//spl rho/-estimate of recall to evaluate the size of region and the generalization fraction of data points in the region, respectively; and 2) uses genetic algorithms to evolve the training model, the evolution is supervised by the generalization performance of one-class SVM. Experiments on an artificial data illustrate the adaptation of the region to the distribution. Experiments on a standard intrusion detection dataset demonstrate that our method not only improves the false positive rate and detection rate, but also is able to control the tradeoff between these measures.
Quang-Anh Tran, Qianli Zhang, Xing Li 0001
SMC2
1999 Session State Transition Based Large Network IDS
Qianli Zhang
Recent Advances in Intrusion Detection1