Pei Xiao 0005

dblp:83/3968-5 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0002-1674-0308ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 RuAG: Learned-rule-augmented Generation for Large Language Models
abstract
In-context learning (ICL) and Retrieval-Augmented Generation (RAG) have gained attention for their ability to enhance LLMs' reasoning by incorporating external knowledge but suffer from limited contextual window size, leading to insufficient information injection. To this end, we propose a novel framework to automatically distill large volumes of offline data into interpretable first-order logic rules, which are injected into LLMs to boost their reasoning capabilities. Our method begins by formulating the search process relying on LLMs' commonsense, where LLMs automatically define head and body predicates. Then, we apply Monte Carlo Tree Search (MCTS) to address the combinational searching space and efficiently discover logic rules from data. The resulting logic rules are translated into natural language, allowing targeted knowledge injection and seamless integration into LLM prompts for LLM's downstream task reasoning. We evaluate our framework on public and private industrial tasks, including Natural Language Processing (NLP), time-series, decision-making, and industrial tasks, demonstrating its effectiveness in enhancing LLM's capability over diverse tasks.
Yudi Zhang 0006, Pei Xiao 0005, Lu Wang 0029, Chaoyun Zhang, Yali Du 0001, Yevgeniy Puzyrev, Randolph Yao, Si Qin, Qingwei Lin, Mykola Pechenizkiy, Dongmei Zhang 0001, Saravan Rajmohan, Qi Zhang 0066
ICLR2
2025 CSLParser: A Collaborative Framework Using Small and Large Language Models for Log Parsing
abstract
Log parsing is a prerequisite for log analysis. Recently, large language models (LLMs) have demonstrated high accuracy in log parsing. However, their frequent invocations incur substantial costs. To address this issue, some methods have turned to small language models (SLMs), which offer improved efficiency but suffer from reduced accuracy due to limited model capacity. To achieve both high accuracy and efficiency, we propose CSLParser, a collaborative log parsing framework using SLMs and LLMs. CSLParser delegates most log parsing tasks to SLMs and selectively invokes LLMs to correct parsing results generated by SLMs, thereby effectively reducing the invocation cost of LLMs while maintaining high accuracy. Specifically, to enhance the accuracy of SLMs, we propose a diversified sampling strategy to select diverse samples for training, enabling SLMs to effectively handle diverse log patterns. To efficiently invoke LLMs, we design a rule-based selection strategy to identify hard cases that are challenging for SLMs to correctly parse, which are subsequently corrected by LLMs. Additionally, we propose a dynamic template updating mechanism that merges similar templates based on structural and semantic information to further enhance parsing accuracy. Extensive experiments on public large-scale log datasets show that CSLParser outperforms state-of-the-art baselines in both accuracy and efficiency.
Weijie Hong, Yifan Wu 0002, Lingzhe Zhang, Chiming Duan, Pei Xiao 0005, Minghua He, Xixuan Yang, Ying Li 0012
ISSRE5
2025 LogAction: Consistent Cross-system Anomaly Detection through Logs via Active Domain Adaptation
abstract
Log-based anomaly detection is a essential task for ensuring the reliability and performance of software systems. However, the performance of existing anomaly detection methods heavily relies on labeling, while labeling a large volume of logs is highly challenging. To address this issue, many approaches based on transfer learning and active learning have been proposed. Nevertheless, their effectiveness is hindered by issues such as the gap between source and target system data distributions and cold-start problems. In this paper, we propose LogAction, a novel log-based anomaly detection model based on active domain adaptation. LogAction integrates transfer learning and active learning techniques. On one hand, it uses labeled data from a mature system to train a base model, mitigating the cold-start issue in active learning. On the other hand, LogAction utilize free energy-based sampling and uncertainty-based sampling to select logs located at the distribution boundaries for manual labeling, thus addresses the data distribution gap in transfer learning with minimal human labeling efforts. Experimental results on six different combinations of datasets demonstrate that LogAction achieves an average 93.01% F1 score with only 2% of manual labels, outperforming some state-of-the-art methods by 26.28%. Website: https://logaction.github.io
Chiming Duan, Minghua He, Pei Xiao 0005, Zhewei Zhong, Yan Niu, Lingzhe Zhang, Siyu Yu, Yifan Wu 0002, Weijie Hong, Ying Li 0012, Gang Huang 0001
ASE3
2025 United We Stand: Towards End-to-End Log-based Fault Diagnosis via Interactive Multi-Task Learning
abstract
Log-based fault diagnosis is essential for maintaining software system availability. However, existing fault diagnosis methods are built using a task-independent manner, which fails to bridge the gap between anomaly detection and root cause localization in terms of data form and diagnostic objectives, resulting in three major issues: 1) Diagnostic bias accumulates in the system; 2) System deployment relies on expensive monitoring data; 3) The collaborative relationship between diagnostic tasks is overlooked. Facing this problems, we propose a novel end-to-end log-based fault diagnosis method, Chimera, whose key idea is to achieve end-to-end fault diagnosis through bidirectional interaction and knowledge transfer between anomaly detection and root cause localization. Chimera is based on interactive multitask learning, carefully designing interaction strategies between anomaly detection and root cause localization at the data, feature, and diagnostic result levels, thereby achieving both sub-tasks interactively within a unified end-to-end framework. Evaluation on two public datasets and one industrial dataset shows that Chimera outperforms existing methods in both anomaly detection and root cause localization, achieving improvements of over 2.92%~5.00% and 19.01% ~ 37.09%, respectively. It has been successfully deployed in production, serving an industrial cloud platform.
Minghua He, Chiming Duan, Pei Xiao 0005, Siyu Yu, Lingzhe Zhang, Weijie Hong, Yifan Wu 0002, Ying Li 0012, Gang Huang 0001
ASE3
2025 Walk the Talk: Is Your Log-based Software Reliability Maintenance System Really Reliable?
abstract
Log-based software reliability maintenance systems are crucial for sustaining stable customer experience. However, existing deep learning-based methods represent a black box for service providers, making it impossible for providers to understand how these methods detect anomalies, thereby hindering trust and deployment in real production environments. To address this issue, this paper defines a trustworthiness metric—diagnostic faithfulness—for models to gain service providers’ trust, based on surveys of SREs at a major cloud provider. We design two evaluation tasks: attention-based root cause localization and event perturbation. Empirical studies demonstrate that existing methods perform poorly in diagnostic faithfulness. Consequently, we propose FaithLog, a faithful log-based anomaly detection system, which achieves faithfulness through a carefully designed causality-guided attention mechanism and adversarial consistency learning. Evaluation results on two public datasets and one industrial dataset demonstrate that the proposed method achieves state-of-the-art performance in diagnostic faithfulness.
Minghua He, Chiming Duan, Pei Xiao 0005, Lingzhe Zhang, Kangjin Wang, Yifan Wu 0002, Ying Li 0012, Gang Huang 0001
ASE4
2025 CoorLog: Efficient-Generalizable Log Anomaly Detection via Adaptive Coordinator in Software Evolution
abstract
Frequent software updates lead to log evolution, posing generalization challenges for current log anomaly detection. Traditional log anomaly detection research focuses on using small deep learning models (SMs), but these models inherently lack generalization due to their closed-world assumption. Large language models (LLMs) exhibit strong semantic understanding and generalization capabilities, making them promising for log anomaly detection. However, they suffer from computational inefficiencies. To balance efficiency and generalization, we propose a collaborative log anomaly detection scheme (CoorLog) that uses an adaptive coordinator to integrate SM and LLM. The coordinator determines if incoming logs have evolved. Non-evolved logs are routed to the SM, while evolved logs are directed to the LLM for detailed inference using the constructed Evol-CoT. To gradually adapt to evolution, we introduce the adaptive evolution mechanism (AEM), which updates the coordinator to redirect evolved logs identified by the LLM to the SM. Simultaneously, the SM is fine-tuned to inherit the LLM’s judgment on these logs. Extensive experiments on real-world datasets demonstrate that CoorLog achieves superior F1-scores in both intra-version and inter-version anomaly detection. Additionally, CoorLog reduces processing time by 91.63% and token consumption by 85.59% compared to using an LLM alone.
Pei Xiao 0005, Chiming Duan, Minghua He, Yifan Wu 0002, Gege Gao, Lingzhe Zhang, Weijie Hong, Ying Li 0012, Gang Huang 0001
ASE1
2024 LogCAE: An Approach for Log-based Anomaly Detection with Active Learning and Contrastive Learning
abstract
Log-based anomaly detection plays a crucial role in maintaining the reliability of software systems. Unsupervised models are more suitable for real-world usage because they do not rely on huge data labeling efforts. However, their effectiveness is limited because of the lack of supervision of data labels. To balance model effectiveness and labeling efforts, existing approaches enhance model capabilities by incorporating relatively few but key human labels as a golden signal, thereby improving the model ability with acceptable labeling efforts. However, these methods still face limitations of complex human labels and insufficient utilization of human knowledge. In this paper, we introduce LogCAE, a two-stage log anomaly detection approach based on active learning and contrastive learning. It utilizes an unsupervised model to learn from unlabeled log data without human labels and incorporates human knowledge through active learning during online optimization. We employ contrastive learning to optimize the representation of log samples in feature space for more efficient usage of human labels. We conducted experiments on three distinct public log datasets (Thunderbird, BGL, and Zookeeper). The results show that our method improves 12.93% F1-score on average with 6.06% labeled data samples. Besides, our approach is more effective in utilizing human labels than state-of-the-art approaches.
Pei Xiao 0005, Chiming Duan, Huaqian Cai, Ying Li 0012, Gang Huang 0001
ISSRE1
2022 An Artificial Intelligence Multiprocessing Scheme for the Diagnosis of Osteosarcoma MRI Images
abstract
Osteosarcoma is the most common malignant osteosarcoma, and most developing countries face great challenges in the diagnosis due to the lack of medical resources. Magnetic resonance imaging (MRI) has always been an important tool for the detection of osteosarcoma, but it is a time-consuming and labor-intensive task for doctors to manually identify MRI images. It is highly subjective and prone to misdiagnosis. Existing computer-aided diagnosis methods of osteosarcoma MRI images focus only on accuracy, ignoring the lack of computing resources in developing countries. In addition, the large amount of redundant and noisy data generated during imaging should also be considered. To alleviate the inefficiency of osteosarcoma diagnosis faced by developing countries, this paper proposed an artificial intelligence multiprocessing scheme for pre-screening, noise reduction, and segmentation of osteosarcoma MRI images. For pre-screening, we propose the Slide Block Filter to remove useless images. Next, we introduced a fast non-local means algorithm using integral images to denoise noisy images. We then segmented the filtered and denoised MRI images using a U-shaped network (ETUNet) embedded with a transformer layer, which enhances the functionality and robustness of the traditional U-shaped architecture. Finally, we further optimized the segmented tumor boundaries using conditional random fields. This paper conducted experiments on more than 70,000 MRI images of osteosarcoma from three hospitals in China. The experimental results show that our proposed methods have good results and better performance in pre-screening, noise reduction, and segmentation.
Jia Wu 0002, Pei Xiao 0005, Haojie Huang 0003, Fangfang Gou, Zhixun Zhou, Zhehao Dai
IEEE J. Biomed. Health Informatics2