VLDB 2026 Research / reviewers in the wild / expert
Guojun Chu
dblp:295/6967
· DBLP profile ↗
4ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0003-2352-4731ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LogNotion: Highlighting Massive Logs to Assist Human Reading and Decision MakingabstractMassive logs contain crucial information about the working status of software systems, which contributes to anomaly detection and troubleshooting. For engineers, it is a laborious task to manually inspect raw logs to know the system running status, and therefore an automated log summarization tool can be helpful. However, due to the specificity of logs in terms of grammar, vocabulary and semantics, existing natural language-based methods cannot perform well in log analysis. To address these issues, we propose LogNotion, a general log summarization framework that highlights the log messages to assist human reading and decision making. We first explore the role played by triplets in log analysis, and propose a triplet extraction method based on sequence tagging and component alignment, in which the specificity of logs is fully taken into account. Then, we propose an unsupervised log summarization method to extract both regular and noteworthy information based on triplets. Comprehensive experiments are conducted on seven real-world log datasets and the results show that LogNotion improves the average ROUGE-1 by 0.26, recall by 0.12, and compression ratio by 2.13%, compared to state-of-the-art log summarization tools. The helpfulness, readability and generalizability are also verified through human evaluation and cross-dataset tests. Guojun Chu, Jingyu Wang 0001, Tao Sun 0010, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Jianxin Liao |
IEEE Trans. Serv. Comput. | 1 |
| 2025 | Anomaly Detection on Interleaved Log Data With Semantic Association Mining on Log-Entity GraphabstractLogs record crucial information about runtime status of software system, which can be utilized for anomaly detection and fault diagnosis. However, techniques struggle to perform effectively when dealing with interleaved logs and entities that influence each other. Although manually specifying a grouping field for each dataset can handle the single grouping scenario, the problems of multiple and heterogeneous grouping still remain unsolved. To break through these limitations, we first design a log semantic association mining approach to convert log sequences into Log-Entity Graph, and then propose a novel log anomaly detection model named Lograph. The semantic association can be utilized to implicitly group the logs and sort out complex dependencies between entities, which have been overlooked in existing literature. Also, a Heterogeneous Graph Attention Network is utilized to effectively capture anomalous patterns of both logs and entities, where Log-Entity Graph serves as a data management and feature engineering module. We evaluate our model on real-world log datasets, comparing with nine baseline models. The experimental results demonstrate that Lograph can improve the accuracy of anomaly detection, especially on the datasets where entity relationships are intricate and grouping strategies are not applicable. Guojun Chu, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Bo He 0003, Yuhan Jing, Lei Zhang 0094, Jianxin Liao |
IEEE Trans. Software Eng. | 1 |
| 2023 | Exploiting Spatial-Temporal Behavior Patterns for Fraud Detection in Telecom NetworksabstractFraud detection in telecom network is a crucial problem that threatens users’ privacy and property security. In recent years, fraudsters adopt more advanced camouflage strategies to avoid being detected by traditional algorithms. To deal with these new types of fraud, it is necessary to analyze the integrated spatial-temporal features, which are rarely involved in existing literature. In this article, we propose a novel fraud detection model based on the intertwined spatial-temporal patterns of user behaviors. Specifically, we first introduce the extension of statistical and interactive features to dynamic call patterns, and build a probabilistic model to simulate users’ call behaviors. Then the sequential patterns reflecting users’ own behaviors are obtained by the mixture Hidden Markov Models, and the structural patterns reflecting the collaboration between users in the telecom network are obtained by the attention-based Graph-SAGE model. Finally, our model outputs a fraud score for each user to detect potential fraudsters. We conduct extensive experiments on a real-world telecom dataset. The experimental results demonstrate that our intertwined spatial-temporal call patterns can effectively represent user behavior and improve the accuracy of fraud detection compared with state-of-the-art methods. The results also validate the efficiency and the interpretability of our model. Guojun Chu, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Shimin Tao, Hao Yang 0006, Jianxin Liao, Zhu Han 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2021 | Prefix-Graph: A Versatile Log Parsing Approach Merging Prefix Tree with Probabilistic GraphabstractLogs play an important part in analyzing system behavior and diagnosing system failures. As the basic step of log analysis, log parsing converts raw log messages into structured log templates. However, existing log parsing approaches are not adaptive and versatile enough to ensure their high accuracy on all types of datasets. In particular, it is required to design regular expressions or fine-tune the hyper-parameters manually for the best performance. In this paper, we propose Prefix-Graph, an online versatile log parsing approach. Prefix-Graph is a probabilistic graph structure extended from prefix tree. It iteratively merges together two branches which have high similarity in probability distribution, and represents log templates as the combination of cut-edges in root-to-leaf paths of the graph. Since no domain knowledge is used and all the parameters are fixed, Prefix-Graph can be easily applied to different log datasets without any additional manual work. We evaluate our approach on 10 real-world datasets and 117GB log messages obtained from Huawei. The experimental results demonstrate that Prefix-Graph achieves the highest average accuracy of 0.975 and the smallest standard deviation of 0.037. Our approach is superior to baseline methods in terms of adaptability and versatility. Guojun Chu, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Shimin Tao, Jianxin Liao |
ICDE | 1 |