Xiaolin Chai

dblp:305/2325 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0002-7447-3367ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 RELog: Robust and Efficient Anomaly Detection Based on Complete Logs for Large-Scale Systems
abstract
As large-scale systems continue to grow in complexity, effectively leveraging the multi-dimensional information contained in semi-structured logs has become a critical challenge for ensuring system stability. Existing methods rely on limited log content, which restricts semantic modeling capability and leads to degraded performance when abnormal samples are scarce or data distributions are imbalanced. To address these challenges, we propose RELog, an efficient and robust anomaly detection framework that comprehensively utilizes multi-dimensional log features. RELog adopts a hierarchical process for efficient log parsing and anomaly detection. During log parsing, Large Language Models are introduced to enhance the semantic understanding of low-confidence templates generated by the heuristic method. For anomaly detection, we design dedicated encoders for templates, parameters, and time features, tailored to the low complexity and high redundancy of logs, enabling discriminative embeddings. Furthermore, we propose an efficient hybrid sequence encoder integrating state-space modeling and attention for capturing continuous log patterns and critical dependencies, complemented by a replacement classification task for imbalanced data. Experimental results demonstrate that RELog achieves high-accuracy log parsing and anomaly detection while maintaining efficiency. Moreover, RELog demonstrates robustness across varying anomaly ratios and dynamic environments, effectively detecting diverse types of anomalies reflected by parameter and time features.
Xiaolin Chai, Zhaoyang Lou, Yan Sun 0004, Mohsen Guizani
IEEE Trans. Computers1
2025 An Effective Log Sequence Anomaly Detection Method Guided by Large Language Models
abstract
With the rapid advancement of information technology, the volume of log data generated by information systems has expanded significantly. The effective detection and processing of anomalies in log sequences has emerged as a crucial research topic. Recent studies have successfully applied pre-trained models for anomaly detection in log sequences. However, due to the semi-structured nature of logs and the presence of domain-specific vocabulary, along with the diverse logging styles of different vendors, those models struggled to provide a comprehensive understanding and unified representation of logs. Additionally, current research primarily focused on log template sequences, neglecting the impact of variable values and contextual relationships among logs. In this paper, we propose LlmBertLog, a BERT-based log sequence anomaly detection model guided by large language models. We first utilize a large language model to transform a limited set of raw logs into natural language descriptions as sample data. Subsequently, we induct three novel pre-training tasks designed to enhance BERT's comprehension and representation capabilities through methods such as semantic augmentation and semantic masking. We then integrate a Transformer with positional encoding and fine-tune LlmBertLog for log sequence anomily detection task, thereby improving its performance in log sequence content mining. Extensive experiments demonstrate that LlmBertLog achieves F1 scores of 0.994 on the BGL dataset and 0.998 on the HDFS dataset for log sequence anomaly detection, surpassing the current best benchmarks by 2.3% and 5.5%, respectively.
Xiaolin Chai, Hong Luo 0001, Yan Sun 0004
CSCWD2
2025 Log Sequence Anomaly Detection Based on Template and Parameter Parsing via BERT
abstract
Logs record various operations and events during system running in text format, which is an essential basis for detecting and identifying potential security threats or system failures, and is widely used in system management to ensure security and reliability. Existing log sequence anomaly detection is limited by log parsing and does not consider all key features of logs, which may cause false or missed detection. In this article, we propose a fast and accurate log parsing method and feed the entire log content into the deep learning network for analysis. To avoid semantic loss during parsing, we replace some variables with tokens containing semantic information and divide logs with appropriate granularity. To ensure the speed and accuracy of parsing, we propose a similarity-based fast merging method to deal with redundant templates. For anomaly detection, we use the complete log content features as input to the model. We use Bidirectional Encoder Representation from Transformers (BERT) to output anomaly detection results directly after considering both the global and local information of log sequences. Experiments show that our log parsing method achieves the best average parsing quality on 16 datasets, and the anomaly detection method achieves optimal results on different datasets.
Xiaolin Chai, Yan Sun 0004, Sajal K. Das 0001
IEEE Trans. Dependable Secur. Comput.1
2024 TopoRCA: A Lightweight Root Cause Analysis System Based on Application Topology
abstract
The advent of the cloud-native era, driven by technologies like cloud computing, big data, and virtualization, has promoted the widespread adoption of microservices architecture in application development. The complex network environment and componentized distributed deployments have introduced new challenges in application performance monitoring. This paper introduces TopoRCA, a system that leverages application topology to identify the root causes of issues in microservices applications. TopoRCA utilizes lightweight machine learning models for anomaly detection, locating faults within the system. It also efficiently and intelligently infers the root cause behind system failures by correlating instance node metrics and service call metrics. We use a microservices dataset consisting of 10 service instances in our experiment. Our experimental results demonstrate that 42% of the root causes are ranked first, 61% in the top three, and 80% in the top five by TopoRCA.
Youliang Huang, Xiaolin Chai, Yan Sun 0004
CSCWD3
2024 GTformer: Graph-Based Temporal-Order-Aware Transformer for Long-Term Series Forecasting
abstract
In the production environment of the Internet of Things (IoT), sensors of various qualities generate a large amount of multivariate time series (MTS) data. The long-term prediction of time series data generated by various IoT devices provides longer foresight and helps execute necessary resource scheduling or fault alarms in advance, thus improving the efficiency of system operation and ensuring system security. In recent years, deep learning models like Transformers have achieved advanced performance in multivariate long-term time series forecasting (MLTSF) tasks. However, many previous research attempts either overlooked the interseries dependencies or ignored the need to model the strict temporal order of MTS data. In this article, we introduce GTformer, a graph-based temporal-order-aware transformer model. We propose an adaptive graph learning method specifically designed for MTS data to capture both uni-directional and bi-directional relations. In addition, we generate positional encoding in a sequential way to emphasize the strict temporal order of time series. By adopting these two components, our model can have a better understanding of the interseries and intraseries dependencies of MTS data. We conducted extensive experiments on eight real-world data sets, and the results show that our model achieves better predictions compared with state-of-the-art methods.
Aobo Liang, Xiaolin Chai, Yan Sun 0004, Mohsen Guizani
IEEE Internet Things J.2
2023 ZAlert: A Real Time Prediction Framework For Network Alert
abstract
With the rapid development of enterprise digital transformation and the widespread use of cloud-native architecture applications, the complexity of enterprise-level systems is getting higher and higher. This requires the engineers to be able to locate and repair alert incidents accurately and quickly to ensure the stability of complex network equipment. In this paper, we propose a novel alert prediction framework, ZAlert, which can predict the occurrence of future alerts and locate the equipment that may cause alerts based on real-time alarm data. First, ZAlert extracts text and statistical features from alarm data to build a high-performance comprehensive learning model to predict the category of alert that may occur in the future with an average performance of 0.81 on F1-score. Then, we use the knowledge graph based on the CMDB(Configuration Management Database) to search for alarm devices. In this way, engineers can quickly find devices that may cause alarms, and improve the the overall efficiency of engineers’ operation and maintenance.
ShuYang Zuo, Xiaolin Chai, Hong Luo 0001, Yan Sun 0004
CSCWD2
2022 S2FCM: A Two-Stage Fine Classification Model for Students' Attention Analysis
abstract
As the school’s surveillance cameras generate a large amount of video data every day, analyzing these video images to promote the development of intelligent campuses has become a research hot spot. By classifying the body postures, we can analyze the attention of students during the learning process. Many researchers have proposed methods to identify the posture of the human body in images. However, due to the occlusion and angle of the video data, these methods are not suitable for classroom images. In this paper, we propose S2FCM, a two-stage fine classification model to analyze the attention of students. We first propose a two-branch network to classify body posture. To deal with the occlusion in the image, we use the skeleton heat map as the input of the graph convolutional structure. In the second stage, we propose a position-robust network to classify the head posture. To solve the problems caused by the position difference, we design a position correction module. Experiments show that the macro-f1-score of our classification model is at least 20% higher than that of the other two general methods.
Zhuhou Zhang, Xiaolin Chai, Hong Luo 0001
CSCWD2
2022 Research and Implementation of Host Behavior Anomaly Detection Technology Based on Deep Learning
abstract
In recent years, with the growth of network hosts and applications, Abnormal host behavior brings great challenges to the protection of user information privacy. Effective detection of abnormal host behavior is of great significance. Host anomaly detection technology can identify abnormal behaviors of the host by establishing the normal behavior benchmark, which is universal and capable of detecting new network attack behavior. However, the existing host behavior anomaly detection technology confronts with three problems(i.e., a large amount of data, difficulty in labeling, and strong manual dependence). To solve these problems, this paper proposes an unsupervised learning model based on BiGAN(Bidirectional Generative Adversarial Network) to detect abnormal host behavior. Our model introduces several loss functions and modifies them to apply to BiGAN. At the same time, we discuss which loss function is more suitable for host behavior anomaly detection. Experimental results show that the BiGAN model using the Hinge loss function is the most stable and has excellent anomaly detection performance.
Yonghao Gu, Xiaolin Chai, Peng Qi 0006
CSCWD3
2022 An artistic analysis model based on sequence cartoon images for scratch
abstract
With the development of visual programming languages, researchers pay attention to the automatic evaluation of visual projects. Previous work focus on the code evaluation but ignored another essential part—the visualization results. Scratch is a widely used programming platform, and projects created on it are displayed in the form of cartoon clips. It is valuable to explore the visual aesthetics embodied in these clips to fill the gap in the assessment system. We propose a model that predicts the human view scores of cartoon clips created on Scratch. Our method is divided into two steps to evaluate the aesthetic of the sequence images that compose cartoon clips. First, we train an image classification network to predict the relative aesthetics of individual images. Then we construct an aesthetic space for the sequence image and improve the rating within a specific range. We put forward ScratchGAN to generate a Scratch-cartoon-style aesthetic analysis data set for training the classification network. Experimental results show that our Generative Adversarial Network framework can well transform photos into a Scratch-cartoon style. The single image assessment network can generate predictions that fit human cartoon aesthetic opinions. Our method achieves satisfactory results in the aesthetic evaluation of sequence cartoon images.
Xiaolin Chai, Yan Sun 0004, Hong Luo 0001, Mohsen Guizani
Int. J. Intell. Syst.1